AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

679
open jobs
26
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 637 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

521 to 540 of 679 jobs
  • Engineering
  • Platform: micro1
  • Location: 58 countries

This role involves training AI models by leveraging your expertise in CAD. You will evaluate content generated by AI in manufacturing workflows, provide detailed feedback on manufacturing processes, and develop evaluation frameworks to measure accuracy and efficiency. The role requires strong programming skills, particularly with Python and coding agents.

Posted September 3, 2026

OpenRecently verified$80-150/hr

Referral link to micro1's job list: search for this role there. Open this exact job

  • Data & Machine Learning
  • Platform: Terac
  • Location: United States

This study invites machine learning engineers and AI researchers to build test scenarios on a remote reinforcement learning platform. Participants set environmental parameters, define spatial constraints and agent interaction rules, run preliminary agent tests, and then share usability feedback in an interview. It is aimed at people with hands-on experience in simulation design and RL environments.

Posted September 3, 2026

  • Writing, Creative & Design
  • Platform: xAI

This role asks a design specialist to evaluate and improve design work across graphic, interface, web and game formats, so that an AI model learns to tell clear, functional design from attractive but weak work. The person writes precise critiques, labels datasets, and creates strong reference examples when existing work falls short. It suits designers with a proven portfolio who can judge quality beyond their own specialty and explain their reasoning clearly in English.

Posted September 3, 2026

OpenRecently verifiedPay not disclosed
  • Science & Research
  • Platform: xAI

This role asks a humanities scholar to judge and improve how an AI model handles questions in linguistics, history, classical studies, literature, philosophy, ethics, and the arts. The work centers on checking the model's answers against primary sources, flagging errors and unsupported claims, and writing reference answers that separate evidence from interpretation and speculation. It suits candidates with deep scholarly training who write precise, carefully calibrated analysis.

Posted September 3, 2026

OpenRecently verifiedPay not disclosed
  • Health & Medicine
  • Platform: Mercor
  • Location: United States
  • Language: Spanish, Chinese, Filipino, Russian, Haitian Creole
  • Level: Intermediate

This role is for practicing inpatient hospitalist physicians to evaluate and annotate clinical documentation generated by medical AI systems. You will review patient records, identify gaps and inaccuracies in notes produced by AI, and provide structured feedback to engineering teams. The role requires advanced proficiency in at least one language other than English.

Posted September 2, 2026

OpenRecently verified$170/hr

Referral link: we may earn a fee. Apply without it

  • Health & Medicine
  • Platform: Mercor
  • Location: United States
  • Language: Spanish, Chinese, Filipino, Russian, Haitian Creole
  • Level: Intermediate

You are a general practitioner or internist currently practicing in an outpatient setting and proficient in English plus another language at an advanced level. You annotate and evaluate clinical notes generated by AI, flagging omissions and inaccuracies to improve clinical decision support tools. This remote and part-time role allows you to leverage your expertise in medical documentation with a specialized AI health product team.

Posted September 2, 2026

OpenRecently verified$170-190/hr

Referral link: we may earn a fee. Apply without it

  • Writing, Creative & Design
  • Platform: micro1
  • Location: 58 countries
  • Level: Intermediate

This role involves annotating and analyzing short audio recordings by documenting precisely the instruments, tempo, tonality and ambient sounds perceived. It is aimed at musicians, producers or sound engineers with proven practical experience and excellent analytical listening ability to identify details in complex sound environments.

Posted September 2, 2026

OpenRecently verified$30-45/hr

Referral link to micro1's job list: search for this role there. Open this exact job

  • Data & Machine Learning
  • Platform: Terac
  • Location: United States

This project asks evaluators to examine how a digital shopping assistant answers real e-commerce questions and to locate where its reasoning or product suggestions fall short. The work is remote and ongoing, and it is aimed at people with quality assurance, data evaluation, or e-commerce backgrounds who can build structured evaluation frameworks.

Posted September 2, 2026

  • Health & Medicine
  • Platform: Mercor
  • Location: United States

This role is for experienced pediatric nurses with acute care expertise. You will evaluate outputs generated by AI systems from real clinical documentation, annotate pediatric nursing assessment data, and provide expert feedback to improve the reliability and safety of medical AI tools. The role combines your field experience in pediatrics with a direct contribution to developing clinically robust algorithms.

Posted September 1, 2026

OpenRecently verified$55-65/hr

Referral link: we may earn a fee. Apply without it

  • Health & Medicine
  • Platform: micro1
  • Location: 58 countries
  • Level: Intermediate

This role involves assessing risks of exploitation and sexual abuse in adolescents by analyzing complex cases to improve AI systems training. You will apply your expertise in child protection to produce clear recommendations grounded in recognized safeguarding frameworks, taking into account developmental factors and trauma. This position is aimed at experienced clinical professionals working in child protection or adolescent mental health.

Posted September 1, 2026

OpenRecently verified$45-95/hr

Referral link to micro1's job list: search for this role there. Open this exact job

  • Science & Research
  • Platform: micro1
  • Location: 58 countries

You will contribute to an AI training project by solving, validating, and commenting on complex problems in your scientific field of expertise. The role combines rigorous scientific analysis, Python programming for calculations and simulations, and critical evaluation of solutions generated by AI.

Posted September 1, 2026

OpenRecently verified$80-150/hr

Referral link to micro1's job list: search for this role there. Open this exact job

  • Science & Research
  • Platform: Mercor
  • Location: United States

This role involves designing realistic evaluation challenges for AI models applied to drug discovery. As an expert in chemistry and molecular biology, you will build complex research scenarios based on real and ambiguous data, where AI must demonstrate genuine reasoning capability rather than simple pattern association. You will be responsible for the entire process: challenge design, data corpus assembly, response evaluation, and iterative revision.

Posted August 31, 2026

OpenRecently verified$60-90/hr

Referral link: we may earn a fee. Apply without it

  • Engineering
  • Platform: Mercor
  • Location: United States

This evaluation role involves designing simulations and technical specifications to test the ability of advanced AI models to reason from first principles in your area of expertise in engineering. You will define complex design problems (controllers, circuits, geometries) that the AI must solve by demonstrating reasoning capability rather than simple data retrieval.

Posted August 31, 2026

OpenRecently verified$60-90/hr

Referral link: we may earn a fee. Apply without it

  • Software Engineering
  • Platform: Mercor
  • Location: United States
  • Level: Intermediate

You will evaluate the quality and correctness of AI-assisted software development traces used to train and assess models at a leading AI laboratory. You must judge the accuracy of coding sessions, the consistency of workflows, and the quality of reasoning, then provide structured written feedback according to an evaluation grid.

Posted August 28, 2026

OpenRecently verified$70-90/hr

Referral link: we may earn a fee. Apply without it

  • Software Engineering
  • Platform: Mercor
  • Location: United States
  • Level: Intermediate

This job involves evaluating the quality and compliance of Kubernetes scenarios for AI model training. You will analyze manifests, diagnose cluster failures, and provide structured feedback according to an evaluation grid, as a DevOps expert or platform engineer.

Posted August 28, 2026

OpenRecently verified$70-90/hr

Referral link: we may earn a fee. Apply without it

  • Finance & Accounting
  • Platform: Mercor
  • Location: United States
  • Level: Intermediate

This role involves designing realistic scenarios for life insurance and annuity advice and underwriting, then evaluating an AI model's responses against industry standards. You must have solid professional experience in life insurance or investment advisory, and be able to provide detailed feedback to improve the AI's reasoning on mortality risk and product suitability.

Posted August 28, 2026

OpenRecently verified$80/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Level: Intermediate

You will evaluate the methodological quality and scientific rigor of machine learning tasks designed to train and test models at a cutting-edge AI laboratory. You will provide structured feedback on experimental design, model selection and evaluation protocols, critiquing claims against observed results.

Posted August 28, 2026

OpenRecently verified$70-90/hr

Referral link: we may earn a fee. Apply without it

  • Software Engineering
  • Platform: Mercor
  • Location: United States
  • Level: Intermediate

This role involves evaluating the quality, correctness, and reproducibility of software engineering tasks used to train and test the AI models of a leading laboratory. You audit tasks at the repository level, baseline patches, test harnesses, and scoring system integrity, providing structured feedback based on an evaluation grid.

Posted August 28, 2026

OpenRecently verified$70-90/hr

Referral link: we may earn a fee. Apply without it

  • Finance & Accounting
  • Platform: micro1
  • Location: 58 countries
  • Level: Intermediate

This job involves contributing to the creation of high-quality datasets for AI model training, leveraging expertise in investment banking. You will design and refine tasks simulating real investment banking activities such as pitch book construction, valuation analysis, and comparable transaction review. Your financial expertise domain is essential for AI systems to learn to reason and perform on authentic, professionally-quality content.

Posted August 27, 2026

OpenRecently verified$45-100/hr

Referral link to micro1's job list: search for this role there. Open this exact job

  • Finance & Accounting
  • Platform: micro1
  • Location: 58 countries
  • Level: Entry level

You will evaluate and improve the ability of AI models to reason about complex financial questions, particularly in valuation, financial modeling, and real-world investment scenarios. You will create prompts and reference responses based on your operational experience, while flagging errors and reasoning weaknesses in the model.

Posted August 27, 2026

OpenRecently verified$245-280/hr

Referral link to micro1's job list: search for this role there. Open this exact job

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

679 AI Evaluation AI training jobs are open on SideHustler today. 26 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 637 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (139), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.