AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

679
open jobs
26
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 637 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

121 to 140 of 679 jobs
  • Science & Research
  • Platform: Alignerr

This contract role asks mathematicians to turn dense informal proofs into structured, machine-checkable Lean 4 formalizations. The work also involves helping AI researchers understand where automated provers fail and how to improve formal verification pipelines. It is aimed at people with strong proof-writing backgrounds who want to work at the intersection of mathematics and AI.

OpenRecently verified$170-200/hr
  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: French

This remote contract asks you to review AI-generated French translations and original French text for accuracy, naturalness, and cultural fit. You would flag errors and awkward phrasing, then suggest clearer, more idiomatic alternatives with structured feedback. It suits native or near-native French speakers who write well in both French and English and can work independently on their own schedule.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This is a remote, hourly contract role where you review JavaScript code produced by AI models and judge its correctness, efficiency and style. You will also point out flaws and suggest better alternatives, which helps improve how AI tools write frontend code. It suits experienced frontend developers who can explain code quality clearly and work independently on their own schedule.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This remote contract role asks JavaScript engineers to review code produced by AI models, checking correctness, efficiency and modern ES6+ practices. Work includes spotting bugs and anti-patterns in frontend logic and proposing cleaner alternatives, along with written explanations of the code. It suits experienced frontend developers who can write clearly, and training is provided so prior AI experience is not required.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This remote contract role asks experienced TypeScript frontend developers to review and assess code written by AI models, and to show those models what type-safe, modern web development looks like. It suits engineers who can explain architectural and typing decisions clearly and who prefer to work on their own schedule on task-based assignments.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This remote contract role asks senior full-stack engineers to assess AI-generated code across frontend and backend, judging correctness, quality, and security. It also involves building internal tooling for data annotation and reviewing system designs. It suits experienced engineers who can give precise written feedback that helps AI models write better code.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: German

This remote contract role asks German-speaking voice actors to record clean audio from scripts and to judge AI-generated speech for naturalness, accuracy and expressiveness. It suits performers with prior narration, dubbing or audiobook experience who can work from a quiet, well-equipped recording space. No AI background is required.

OpenRecently verified$250-280/hr
  • Business, Consulting & Operations
  • Platform: Alignerr
  • Level: Intermediate

This remote, freelance role asks compliance and risk specialists to help build and assess AI systems that reason about security governance. The work centers on checking security policies, controls, and audit documentation against established frameworks and writing precise feedback that teaches models the nuances of compliance logic. It suits experienced GRC, compliance, or information security professionals who can work independently on flexible hours.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Greek

This remote, flexible contract role asks native or near-native Greek speakers to assess text produced by AI systems for accuracy, naturalness, and cultural fit. The work involves flagging errors and awkward phrasing and proposing more idiomatic alternatives, with structured feedback. It suits linguists or language professionals with a methodical eye for quality, and no prior AI experience is needed.

  • Health & Medicine
  • Platform: Alignerr

This remote, hourly contract asks you to review and grade AI-generated content about healthcare data, electronic health record systems, and clinical workflows. The work helps AI labs make their models more accurate on medical information. It suits health informatics professionals who can judge clinical and technical accuracy and explain their feedback in writing.

  • Health & Medicine
  • Platform: Alignerr

This remote contract role asks public health specialists to test and improve how AI systems reason about population health. The work involves building complex scenarios such as outbreak simulations and trial interpretation, writing reference answers based on research and official guidelines, and reviewing AI health advice for accuracy and bias. No prior AI experience is required, since subject matter expertise is the main qualification.

  • Health & Medicine
  • Platform: Alignerr

This remote contract role asks public health graduates to bring their domain knowledge to AI training. The work involves building realistic population health scenarios, writing evidence-based reference answers, and reviewing AI-generated health advice for accuracy, logic, bias, and ethics. It suits independent, self-directed specialists who can explain health data clearly, and no prior AI experience is required.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Spanish

This remote contract role involves reviewing transcripts of AI-assisted healthcare phone calls held in Central American Spanish. Labelers judge how the AI confirms patient details, handles health concerns and poor audio, and explains scheduling and next steps. It suits native or near-native Spanish speakers with basic healthcare knowledge who can follow detailed guidelines.

  • Data & Machine Learning
  • Platform: Alignerr

This contract role involves labeling and checking clinical and medical research datasets so they can train medical AI systems. It suits people with a healthcare, clinical research, or life sciences background who can apply detailed annotation guidelines consistently and document their reasoning on ambiguous cases.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Hebrew

This remote freelance role asks native or near-native Hebrew speakers to review AI-generated Hebrew text and translations for accuracy, naturalness, register and cultural fit. Reviewers also write structured feedback and suggest idiomatic alternatives. It suits linguists and editors who can work independently on a flexible schedule.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Hungarian

This role asks you to check AI-generated Hungarian text and translations for accuracy, naturalness, and cultural fit, then give structured feedback so the systems write and translate better. It suits native or near-native Hungarian speakers with a careful eye for language quality, and no prior AI experience is required. The work is remote and asynchronous, done on a flexible schedule.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This remote contract role asks an IAM specialist to help train and evaluate AI systems on identity and access security problems. The work covers reviewing access control scenarios, spotting misconfigurations and attack paths, and writing structured feedback. It suits experienced identity security professionals who can explain how identity failures happen in real enterprise environments.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This is a remote, hourly contract to help train AI systems on security incident work. The contractor reviews realistic alerts and incident cases and checks whether the AI's analysis matches how experienced security teams think and act. It suits working security practitioners with hands-on SOC or incident response experience who can explain their reasoning in writing.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Indonesian

This remote contract role asks native or near-native Indonesian speakers to check text produced by AI systems, judging whether translations and original passages are accurate, natural and suited to context. The work is done asynchronously and results in structured feedback that helps models use Indonesian more faithfully. It suits people who know contemporary Indonesian across registers and are comfortable reviewing written content at volume.

  • Software Engineering
  • Platform: Alignerr

This contract role asks software engineers to write, review and evaluate code produced by AI systems, giving expert feedback that helps coding assistants become more accurate, secure and reliable. It suits independent developers who can read code closely, explain their technical reasoning in plain writing, and work asynchronously on their own schedule.

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

679 AI Evaluation AI training jobs are open on SideHustler today. 26 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 637 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (139), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.