AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

679
open jobs
26
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 637 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

621 to 640 of 679 jobs
  • Business, Consulting & Operations
  • Platform: Mercor
  • Location: United States
  • Level: Expert

This role involves supporting the development of advanced language models as a marketing subject matter expert. You will design relevant training tasks, evaluate AI model outputs against structured rubrics, and refine scoring criteria to ensure the quality of training data. This role is suited to experienced marketers who have already contributed to AI model evaluation.

Posted July 10, 2026

OpenRecently verified$60-80/hr

Referral link: we may earn a fee. Apply without it

  • Business, Consulting & Operations
  • Platform: Mercor
  • Location: United States
  • Level: Expert

Retail specialist role with a GenAI team at a leading AI laboratory. You will bring your deep domain expertise in merchandising, category management, and retail operations to evaluate AI model outputs and enrich training data. The role combines strategic advisory to research and engineering teams, design of complex retail tasks, and rigorous evaluation of outputs against structured evaluation frameworks.

Posted July 10, 2026

OpenRecently verified$60-80/hr

Referral link: we may earn a fee. Apply without it

  • Writing, Creative & Design
  • Platform: Mercor
  • Location: United States
  • Language: Russian
  • Level: Intermediate

Mercor is recruiting experienced music producers and sound engineers to evaluate generative AI music models. You will analyze music generated by AI according to detailed quality criteria, in Russian and English, rating musicality, creativity, adherence to instructions, vocal quality, and mixing.

Posted July 6, 2026

OpenRecently verified$35-49/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States

An AI research organization is seeking advanced users of language models to evaluate the quality of AI responses to complex, contextualized personal tasks. Your role involves judging whether AI systems provide useful, personalized, and realistic advice in areas such as nutrition, health, productivity, or career. This position is for candidates who intensively use AI tools in their daily lives and can identify what makes an AI response relevant or failing.

Posted June 13, 2026

OpenRecently verified$50-200/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Assamese

You join a red teaming team to test the robustness and security of conversational AI models. Your mission is to generate adversarial inputs (jailbreaks, prompt injections, misuse cases) to uncover vulnerabilities in AI systems before their deployment. You annotate failures, classify risks, and produce documented datasets that clients can leverage to improve their models.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Bangla

This role involves testing conversational AI models as an AI safety expert by generating adversarial training data to identify vulnerabilities. You will produce reports and structured datasets that enable clients to improve the robustness and security of their AI systems. The ideal profile masters critical content analysis, methodological rigor, and technical communication.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Gujarati

You join a red teaming team to test and secure conversational AI models. Your job consists of generating adversarial inputs, identifying vulnerabilities (biases, disinformation, harmful behaviors) and producing high-quality annotated data that makes AI safer. This role requires fluent mastery of English and Gujarati as well as refined judgment on language and content.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Kannada

This role involves evaluating the safety of conversational AI models by testing them with adversarial inputs and sophisticated attack techniques. You will generate high-quality annotation data, document discovered vulnerabilities, and contribute to strengthening the robustness of AI systems. This job is suited for someone capable of carefully analyzing AI responses, detecting biases and subtle flaws, and communicating findings in a structured manner.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Malayalam

You join a specialized AI safety team to test conversational language models by exposing their flaws and vulnerabilities through adversarial attacks. The job consists of generating high-quality training data through error annotation, risk classification, and documentation of reproducible attacks. This role is for rigorous language experts with critical judgment about the quality and accuracy of AI responses.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Odia

Join a red teaming team dedicated to AI safety: you will probe conversational models with adversarial inputs to discover vulnerabilities and generate robust test data. This role requires native mastery of English and Odia, as well as fine judgment to assess the relevance, accuracy, and appropriateness of AI responses when facing sensitive topics. You will document your findings in a reproducible manner so that clients can strengthen the security of their systems.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Punjabi

Join a red teaming team specialized in identifying vulnerabilities of conversational AI models. You will generate high-quality training data by testing AI systems with adversarial inputs, documenting flaws, and producing reproducible reports that clients can use to strengthen the security of their models.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Tamil

This role involves testing and evaluating the robustness of conversational AI models by subjecting them to adversarial attacks, bypasses, and malicious use cases. You will generate high-quality annotation data to identify vulnerabilities, biases, and systemic risks, following established taxonomies and benchmarks. This role is suited for AI safety experts with advanced proficiency in both English and Tamil.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Telugu

You join an adversarial testing team tasked with detecting vulnerabilities in conversational AI models by generating high-quality human-generated data. Your job consists of exploring security flaws, biases and harmful behaviors through malicious inputs, then documenting systemic risks so that clients can correct them. You must be bilingual in English and Telugu and demonstrate rigor in analyzing AI responses on sensitive topics.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Marathi

This role involves testing conversational AI models as a safety expert by generating adversarial attacks, security bypasses, and misuse scenarios to identify vulnerabilities. You will produce high-quality annotation data and reproducible reports that help clients strengthen the robustness of their AI systems. The ideal profile must possess sound judgment about language and content, be rigorous in identifying subtle anomalies, and capable of clearly communicating observations.

Posted June 3, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Vetto

This remote project asks contributors to role-play users holding assigned political views in multi-turn conversations with an AI system, then judge whether the AI stays neutral, balanced and responsible under pushback. It suits people who already take part in political debate in everyday life and can keep a persona separate from their own opinions while applying detailed written guidelines.

Posted May 22, 2026

OpenRecently verifiedPay not disclosed
  • Health & Medicine
  • Platform: Vetto

This project asks nutrition professionals to review and improve clinical and educational nutrition scenarios that will train an AI planning assistant to act as a nutrition tutor. The work involves breaking down dietary problems, mapping decision paths, and justifying each choice with concrete data. It is aimed at experienced nutritionists across specialties, and at final-year nutrition or dietetics students with clinical practice.

Posted May 11, 2026

OpenRecently verifiedPay not disclosed
  • Health & Medicine
  • Platform: Vetto

This freelance role asks mental health professionals to take part in multi-turn simulated conversations with an AI system, playing a patient persona in scenarios involving emotional distress, suicidal ideation, and depression. Contributors then review the AI's replies against clinical criteria such as risk recognition, empathy, and crisis-response practice. It is aimed at licensed or formally trained clinicians and other mental health practitioners.

Posted May 8, 2026

OpenRecently verifiedPay not disclosed
  • Health & Medicine
  • Platform: Vetto

This project asks nutrition professionals to review and improve real-world dietary scenarios used to train an AI planning assistant that works as a nutrition tutor. Contributors break down clinical or dietary challenges, compare alternative explanations, and justify each decision with concrete data so the reasoning is both rigorous and teachable. It is aimed at experienced nutritionists across clinical, sports, pediatric, oncological, and community specialties, and final-year nutrition or dietetics students with clinical experience may also apply.

Posted April 28, 2026

OpenRecently verifiedPay not disclosed
  • Science & Research
  • Platform: Handshake AI
  • Location: United States
  • Level: Intermediate

Handshake AI is seeking energy professionals with hands-on expertise to evaluate AI-generated content and create training data for renewable energy and broader energy industry workflows. Contributors will leverage their professional experience in areas such as renewable energy modeling, microgrid design, and asset performance analysis to provide feedback that improves AI understanding of energy engineering practices.

Posted April 25, 2026

  • Data & Machine Learning
  • Platform: Vetto

Vetto is hiring quality assurance reviewers to audit responses and tasks produced by human annotators for AI projects. The work checks that delivered data meets high quality standards and flags content that appears AI-generated or departs from the project guidelines. It suits people who can judge data accuracy and consistency against detailed evaluation criteria.

Posted March 4, 2026

OpenRecently verifiedPay not disclosed

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

679 AI Evaluation AI training jobs are open on SideHustler today. 26 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 637 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (139), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.