AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

682
open jobs
29
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 640 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 9, 2026.

281 to 300 of 682 jobs
  • Science & Research
  • Platform: DataAnnotation

This role involves evaluating and improving how AI models handle rigorous mathematical reasoning. You will create challenging math problems, review AI solutions for logical errors and unjustified steps, and write correct solutions that serve as training examples for the models to learn from.

Talent poolRecently verified$40-125/hr
  • Engineering
  • Platform: DataAnnotation
  • Level: Intermediate

This role involves creating rigorous mechanical engineering problems and solutions to train AI models to reason like practicing engineers. You will write verifiable problems from your specialty, derive correct answers, test them against AI systems, and review AI-generated engineering work for real-world applicability.

Talent poolRecently verified$40-125/hr
  • Engineering
  • Platform: DataAnnotation
  • Level: Intermediate

A mechanical designer writes rigorous engineering problems and solutions to train AI models to reason like practicing engineers. You'll create challenging, closed-ended technical problems from your specialty, verify correct answers, test them against AI systems, and review AI-generated engineering work for real-world applicability.

Talent poolRecently verified$40-125/hr
  • Engineering
  • Platform: DataAnnotation
  • Level: Intermediate

A mechanical engineer will create challenging, closed-ended engineering problems based on their specialty, derive correct solutions, and evaluate how well AI models perform on real-world engineering tasks. The role involves writing problems that test genuine engineering reasoning, component sizing, failure analysis, CAD interpretation, design impact prediction, and assessing AI-generated analysis and designs for correctness and practical applicability.

Talent poolRecently verified$40-125/hr
  • Health & Medicine
  • Platform: DataAnnotation

Medical coders evaluate AI-generated medical coding outputs against real clinical documentation and coding guidelines, identifying errors in ICD-10-CM, CPT, HCPCS codes, modifiers, and E&M levels. The role requires writing test prompts to assess coding accuracy and providing corrected assignments with clear documentation-based explanations to serve as training signals for AI models.

Talent poolRecently verified$40-125/hr
  • Health & Medicine
  • Platform: DataAnnotation

This role involves evaluating how AI models handle nursing judgment in clinical scenarios such as patient assessment, medication safety, and triage protocols. Nurses will identify unsafe reasoning, incorrect clinical priorities, and protocol violations, then provide corrections grounded in real bedside practice.

Talent poolRecently verified$40-125/hr
  • Software Engineering
  • Platform: DataAnnotation

This role involves evaluating how AI models reason about offensive security concepts and identifying flaws in their exploit chain reasoning. Penetration testers with hands-on red team experience will test model outputs across reconnaissance, exploitation, privilege escalation, and lateral movement, then write accurate attack paths that reflect real-world tradecraft when models fall short.

Talent poolRecently verified$40-125/hr
  • Science & Research
  • Platform: DataAnnotation

This role involves evaluating AI models' physics reasoning across mechanics, electromagnetism, quantum mechanics, and thermodynamics. You will create problems to test AI understanding, identify errors in AI solutions, and write gold-standard solutions that clearly explain the physics at each step.

Talent poolRecently verified$40-125/hr
  • Finance & Accounting
  • Platform: DataAnnotation

This role involves stress-testing AI models' financial reasoning by writing prompts that probe quantitative analysis and evaluating their outputs for errors in pricing, risk assessment, portfolio construction, and statistical modeling. You will review AI-generated financial analysis for flaws in derivations and model application, and write correct analyses to serve as training data.

Talent poolRecently verified$40-125/hr
  • Health & Medicine
  • Platform: DataAnnotation

This role asks radiologists and imaging technologists to read de-identified scans and judge how accurately frontier AI models interpret them. Contributors write structured reports and impressions, then compare them with model output to find errors and gaps. It suits clinicians with hands-on imaging experience who can explain why a machine-generated report is wrong.

Talent poolRecently verified$40-125/hr
  • Health & Medicine
  • Platform: DataAnnotation

This role involves evaluating AI models' clinical reasoning on nursing topics including assessment, medication administration, and patient care protocols. Registered nurses will identify gaps between textbook knowledge and real bedside practice, write test prompts to probe nursing judgment, and provide corrected guidance when AI outputs contain unsafe or incorrect clinical information.

Talent poolRecently verified$40-125/hr
  • Business, Consulting & Operations
  • Platform: DataAnnotation

This role involves evaluating AI-generated answers on healthcare revenue cycle topics including claims, denials, and payer operations. You will write test prompts to probe the AI's understanding of real-world billing mechanics and provide expert corrections and guidance based on hands-on revenue cycle experience.

Talent poolRecently verified$40-125/hr
  • Engineering
  • Platform: DataAnnotation
  • Level: Intermediate

A Senior Mechanical Design Engineer will create rigorous engineering problems and solutions to evaluate and improve AI model reasoning across mechanical specialties. You will write closed-ended technical challenges from your expertise, derive correct answers, test them against AI systems, and review AI-generated engineering work for real-world accuracy and applicability.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Slovak

This role involves judging AI output on Slovak translation and localization, including whether the text reads naturally and fits its cultural context. Reviewers pinpoint errors, explain why a model went wrong, and draft stronger answers when needed. It suits professionals with Slovak linguistic expertise and clear written English, since their explanations shape how models learn.

Talent poolRecently verified$25-40/hr
  • Software Engineering
  • Platform: DataAnnotation

You will evaluate AI-generated code by running models on real engineering tasks, analyzing their output against production standards, and stress-testing them to find failures. This role is ideal for experienced software engineers who want to contribute to AI model improvement through rigorous code review and red-teaming.

Talent poolRecently verified$40-150/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Telugu

Evaluate AI model outputs on Telugu translation, localization, and cultural adaptation tasks. Review and rate the quality of AI-generated Telugu content, identify errors, and provide corrected versions to help train improved AI systems.

Talent poolRecently verified$25-40/hr
  • Finance & Accounting
  • Platform: DataAnnotation

Evaluate AI models' reasoning about financial markets, trading strategies, and execution mechanics. You'll assess whether AI-generated trading approaches account for real-world factors like liquidity, spreads, and slippage, then write detailed explanations to improve model training.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Ukrainian

As a Ukrainian Specialist, you will evaluate and improve AI-generated Ukrainian text by assessing translations, rating model outputs, and rewriting substandard content to ensure it reflects natural, native Ukrainian. The role focuses on identifying and correcting grammatical issues like case errors, aspect misuse, and Russian-influenced phrasing that models commonly introduce.

Talent poolRecently verified$25-40/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Urdu

Evaluate AI model outputs on Urdu translation and localization tasks, assessing translation quality and cultural appropriateness. You will identify errors in AI-generated content, explain model failures, and provide corrected translations to improve AI training data.

Talent poolRecently verified$25-40/hr
  • Engineering
  • Platform: DataAnnotation
  • Level: Intermediate

A Validation Engineer writes challenging mechanical engineering problems from their specialty, solves them, and evaluates how well AI systems perform on real engineering reasoning tasks. The role focuses on creating benchmarks that measure whether AI can do actual engineering work like component sizing, failure mode identification, and design change prediction.

Talent poolRecently verified$40-125/hr

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

682 AI Evaluation AI training jobs are open on SideHustler today. 29 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 9, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 640 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (142), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.