AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

679
open jobs
26
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 637 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

141 to 160 of 679 jobs
  • Software Engineering
  • Platform: Alignerr

This contract role asks engineers to write, debug and review code, then judge code produced by AI models so that those models learn to reason about software like experienced developers. It suits software engineers who write clean code, explain their technical choices clearly in writing, and work independently on their own schedule.

  • Finance & Accounting
  • Platform: Alignerr

This remote, task-based contract asks finance professionals to review and improve investment insights produced by AI systems. The work covers company analysis, macro and sector assessment, and structured write-ups of opportunities and risks. It is aimed at people with investment research, strategy, or consulting backgrounds and needs no prior AI experience.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Italian

This remote, part-time freelance role is for native or near-native Italian speakers who can judge the quality of Italian text produced by AI systems. The work involves checking translations and original content for accuracy and naturalness, then writing structured feedback. No prior AI experience is required.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Japanese

This role consists of reviewing Japanese text produced by AI systems, including translations and original content, and judging whether it is accurate, natural and culturally appropriate. Reviewers then give structured feedback and suggest more idiomatic alternatives. It suits people with native or near-native Japanese who can also write clearly in English, and no prior AI experience is needed.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Japanese

This role is for native or near-native Japanese speakers who can judge the quality of Japanese text produced by AI systems. The work is to check outputs for accuracy, naturalness and cultural fit, then write structured feedback that helps improve them. It is a remote hourly freelance contract with flexible, self-managed hours.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This contract role asks JavaScript developers to judge code written by AI systems for correctness, efficiency and modern ES6+ practice. It also involves writing reference solutions and explaining the reasoning behind them so that AI models learn to produce better JavaScript. It suits experienced developers who can work remotely on their own schedule.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

You will check code written by AI models for correctness, speed, and modern JavaScript conventions, and write your own reference solutions to compare against. The role suits experienced JavaScript developers who can explain technical reasoning clearly in writing. It is a remote, hourly, self-scheduled contract with no prior AI experience required.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Korean

This remote, task-based role asks Korean speakers to review text produced by AI systems for fluency, accuracy, naturalness, and cultural fit. It suits people with native or near-native Korean and strong written English who can give clear, structured feedback on language quality.

  • Science & Research
  • Platform: Alignerr

This contract role involves turning mathematical theorems, lemmas, and proofs from textbooks and articles into precise, machine-checked Lean 4 code. The work also includes checking existing formalizations for logical soundness and flagging gaps in informal arguments, with the output used to build datasets that teach AI systems to reason about mathematics. It suits mathematicians and formal verification experts who can work independently and on their own schedule.

OpenRecently verified$170-200/hr
  • Legal
  • Platform: Alignerr

This remote freelance role asks you to review and assess AI-generated content on law, government policy, and regulatory compliance. You check accuracy, citations, terminology, and how well the AI handles nuance across legislation, case law, and public policy. It suits people with a background in law, government, or policy who can write clear, structured feedback and work on their own schedule.

  • Legal
  • Platform: Alignerr

This remote contract role asks legal and policy professionals to check AI-generated content on legislation, case law, and government matters for accuracy and nuance. The work is independent and asynchronous, and it suits people with strong legal or public policy knowledge who want to help shape how AI handles consequential outputs. No prior AI experience is needed.

  • Software Engineering
  • Platform: Alignerr

This is a remote, hourly freelance contract to review and assess Lua code written by AI models. The work involves judging correctness, efficiency and adherence to best practices, then writing structured feedback that helps those systems improve. It is aimed at developers with deep Lua knowledge, such as those who have built systems in Roblox Studio or scripted game engines.

  • Data & Machine Learning
  • Platform: Alignerr

This is a remote, hourly contract role in which experienced machine learning and software engineers write step-by-step reasoning logs and training datasets that teach frontier AI models how to solve complex technical problems. The work also involves auditing and refining AI-generated reasoning chains for errors and logical gaps. It is aimed at technical professionals who can work independently and on their own schedule.

  • Data & Machine Learning
  • Platform: Alignerr

This remote contract role asks domain experts with graduate-level training to write hard machine learning evaluation problems that test how advanced AI systems reason. Specialists also judge AI-generated solutions for correctness, creativity and methodological rigor, and they document difficulty and likely failure modes. It suits researchers whose deep, niche expertise goes beyond standard ML knowledge.

OpenRecently verified$200-400/hr
  • Business, Consulting & Operations
  • Platform: Alignerr
  • Level: Intermediate

The role involves reviewing AI-generated management and leadership scenarios and judging the reasoning behind them, so that AI systems handle organizational behavior and decision-making more accurately. It is aimed at experienced management educators who can spot flawed logic and give structured, constructive feedback. The work is remote, asynchronous, and done on a flexible schedule.

  • Data & Machine Learning
  • Platform: Alignerr

This remote, hourly contract involves checking and assessing geographic information produced by AI systems, such as addresses, place names, and location search results. It suits people who know their own city or region well and can spot errors in digital maps, with no GIS or geography degree needed.

  • Business, Consulting & Operations
  • Platform: Alignerr
  • Level: Intermediate

This remote contract role invites experienced marketing educators to review AI-generated marketing questions, scenarios and explanations used to train models. The work checks whether the reasoning on branding, consumer behavior and go-to-market planning is accurate and useful. It suits instructors who have taught marketing at university level and can give structured feedback at their own pace.

  • Business, Consulting & Operations
  • Platform: Alignerr

This flexible, task-based contract asks experienced marketing professionals to design realistic enterprise marketing scenarios that train and evaluate advanced AI agents. The work centers on drafting on-brand announcements, building competitive comparisons and tracing campaign attribution, each paired with a precise scoring rubric. It suits marketers who are comfortable with campaign data and want to shape how AI handles real marketing work.

  • Science & Research
  • Platform: Alignerr

This remote, asynchronous contract role asks holders of a Master's or PhD in a STEM or highly technical field to build hard, multi-disciplinary problems and review AI-generated solutions to strengthen how AI systems reason. It suits self-directed researchers with publications or substantial original projects who can explain complex ideas precisely in writing. Hours are flexible and the work may extend into further projects.

  • Science & Research
  • Platform: Alignerr

This remote contract role asks materials science specialists to write and check demanding problems that test how AI models reason about semiconductor materials and molecular modeling. It is aimed at Master's and PhD holders, including working researchers, academics and industry experts. Candidates review AI-generated scientific answers and help close their gaps, and no prior AI experience is required.

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

679 AI Evaluation AI training jobs are open on SideHustler today. 26 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 637 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (139), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.