AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

682
open jobs
29
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 640 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 9, 2026.

261 to 280 of 682 jobs
  • Data & Machine Learning
  • Platform: DataAnnotation
  • Level: Intermediate

This role involves evaluating how well AI models can perform computational biology tasks. Computational biologists and bioinformaticians will design realistic analysis scenarios from their own work, run them through AI systems, and grade the results against professional standards to help benchmark model capabilities.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Croatian

This role involves judging how well frontier AI models handle Croatian translation, localization, and cultural nuance on real language projects. The contributor finds errors, explains why they occur, and writes better answers when a model falls short. It is aimed at linguists and bilingual reviewers with professional language experience who want flexible, remote, project-based work.

Talent poolRecently verified$25-40/hr
  • Software Engineering
  • Platform: DataAnnotation

This role involves evaluating AI models' ability to generate secure code by identifying vulnerabilities, testing their security reasoning, and writing corrections when they fail. It suits security engineers and cybersecurity professionals who can spot weaknesses in authentication, cryptography, input handling, and secret management that automated tools miss.

Talent poolRecently verified$40-125/hr
  • Data & Machine Learning
  • Platform: DataAnnotation

As a Data Scientist, you will evaluate analyses generated by AI models on real datasets, checking whether the statistical methods and conclusions are sound. You will identify flaws in the reasoning, stress-test the logic, and write detailed assessments that serve as training data for future model improvements.

Talent poolRecently verified$40-150/hr
  • Software Engineering
  • Platform: DataAnnotation

Evaluate how AI models handle DevOps tasks such as CI/CD pipelines, infrastructure-as-code, and cloud operations. Identify failures in automation, security configurations, and deployment processes, then write correct solutions that a platform engineer would trust.

Talent poolRecently verified$40-150/hr
  • Science & Research
  • Platform: DataAnnotation
  • Level: Intermediate

A drug discovery scientist will design realistic preclinical tasks and evaluate how well frontier AI models perform them, grading responses against professional standards. The role involves creating scenarios based on real workflows, running them through AI systems, and providing detailed feedback on whether the models correctly conduct SAR analysis, DMPK interpretation, screening triage, and other core discovery work.

Talent poolRecently verified$40-125/hr
  • Engineering
  • Platform: DataAnnotation
  • Level: Intermediate

This role involves evaluating how advanced AI models handle electronics and electrical engineering problems. You will design realistic engineering challenges, test them against frontier AI systems, and judge the output quality using your professional expertise, providing detailed critiques and reasoning to help train better AI tools.

Talent poolRecently verified$40-125/hr
  • Engineering
  • Platform: DataAnnotation
  • Level: Intermediate

DataAnnotation seeks an experienced embedded systems developer to evaluate how AI models handle real electrical and electronics engineering problems. You will design realistic engineering challenges, run them through frontier AI models, judge the output against professional standards, and provide detailed written critiques that help train the next generation of AI engineering tools.

Talent poolRecently verified$40-125/hr
  • Science & Research
  • Platform: DataAnnotation
  • Level: Intermediate

This role involves designing realistic experimental biology tasks and evaluating how frontier AI models perform on benchmark scientific challenges. You'll create scenarios from your own laboratory practice, run them through AI systems, and grade the AI's output against professional standards to help train models on hands-on experimental work.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Gujarati

Evaluate AI models on Gujarati translation and localization tasks, assessing the quality of machine-generated outputs and providing corrections. The role involves analyzing errors in AI-generated Gujarati text, explaining failure modes, and writing better translations to help train improved language models.

Talent poolRecently verified$25-40/hr
  • Engineering
  • Platform: DataAnnotation
  • Level: Intermediate

DataAnnotation seeks an experienced hardware engineer to evaluate how AI models handle real electrical and electronics engineering problems. You will design realistic engineering challenges, run them through frontier AI models, judge the outputs against professional standards, and provide detailed written critiques that help train the next generation of AI engineering tools.

Talent poolRecently verified$40-125/hr
  • Health & Medicine
  • Platform: DataAnnotation

This role involves evaluating AI-generated responses to health and medical questions for clinical accuracy and safety. The expert identifies errors in reasoning, missed contraindications, unsafe advice, and writes correct answers grounded in current clinical practice, ensuring AI models provide guidance that real clinicians would trust.

Talent poolRecently verified$40-125/hr
  • Business, Consulting & Operations
  • Platform: DataAnnotation

This role involves evaluating AI-generated responses about healthcare administration and operations, including billing, coding, insurance workflows, and compliance. You will write test prompts, review AI outputs for accuracy and feasibility, and provide correct guidance based on practical operational experience.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Icelandic

This role involves judging how frontier AI models handle Icelandic translation, localization, and cultural nuance on real projects. The reviewer rates model outputs, pinpoints errors, and explains why they happen. It suits people with professional language expertise who can write clear English explanations.

Talent poolRecently verified$25-40/hr
  • Finance & Accounting
  • Platform: DataAnnotation

This role involves evaluating AI systems that price insurance risks and make claims decisions by reviewing underwriting logic, policy language interpretation, and risk assessment. Insurance professionals will verify whether AI models correctly apply underwriting principles, pricing models, and policy terms, then provide written feedback to improve model training.

Talent poolRecently verified$40-125/hr
  • Legal
  • Platform: DataAnnotation

As an IP Attorney, you will evaluate how AI models handle patent prosecution, trademark analysis, and IP strategy work. You'll test the model's reasoning on patentability, claim construction, and prior art analysis, then write rigorous legal opinions to correct errors and train the system. This role suits experienced IP practitioners who can identify and remedy flawed AI reasoning across intellectual property domains.

Talent poolRecently verified$40-125/hr
  • Data & Machine Learning
  • Platform: DataAnnotation

As a Machine Learning Engineer, you will evaluate and improve how AI models reason about machine learning systems, including training dynamics, evaluation design, and deployment strategies. You will write prompts to test model reasoning, identify subtle errors in AI outputs, and provide correct solutions based on real practitioner experience.

Talent poolRecently verified$40-150/hr
  • Business, Consulting & Operations
  • Platform: DataAnnotation

This role involves evaluating AI systems on their understanding of managed care operations, including benefit determination, prior authorization, and payer contracts. You'll identify errors in AI-generated coverage decisions and write correct guidance based on real-world managed care practices to improve AI training.

Talent poolRecently verified$40-125/hr
  • Business, Consulting & Operations
  • Platform: DataAnnotation

This role asks experienced consultants to test cutting-edge AI models on business problems drawn from their own practice. Contributors write client-style briefs, review each response as they would an associate's work, and flag where recommendations are sharp and where they are generic. It suits people with consulting, strategy, PMO, or deal advisory backgrounds.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Marathi

Evaluate AI model outputs on Marathi translation and localization tasks, identifying errors and cultural nuances. You will rate the quality of frontier models' Marathi work, explain where they fail, and provide corrected answers to help train future AI systems. This is a flexible, remote position for Marathi language specialists.

Talent poolRecently verified$25-40/hr

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

682 AI Evaluation AI training jobs are open on SideHustler today. 29 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 9, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 640 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (142), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.