AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

679
open jobs
26
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 637 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

161 to 180 of 679 jobs
  • Science & Research
  • Platform: Alignerr

This remote contract role asks materials science experts to help train and evaluate AI models on semiconductor materials, molecular modeling, and other advanced materials problems. It is aimed at researchers with a master's or doctoral background who can write and review scientifically rigorous problems. No prior AI experience is needed because training is provided.

  • Science & Research
  • Platform: Alignerr

This contract role asks mathematics experts to label, organize and check math content that trains AI models. The work is remote and asynchronous, with a minimum weekly commitment. It suits people with strong math backgrounds in research, teaching or applied problem solving.

  • Science & Research
  • Platform: Alignerr

This freelance role asks mathematicians to design demanding problems and reference solutions that test how AI systems reason. The work also includes checking model outputs for accuracy and giving structured feedback, with benchmarks spanning undergraduate to PhD-level material. It suits people with advanced mathematical training who can work independently and remotely.

  • Science & Research
  • Platform: Alignerr

This freelance contract is for PhD-level mathematicians who will evaluate and improve how AI systems handle mathematical reasoning. The work covers writing and reviewing advanced problems, checking AI-generated proofs and calculations, and organizing mathematical content for training datasets. It suits holders of a doctorate in mathematics or a closely related field who can work independently and remotely on their own schedule.

  • Science & Research
  • Platform: Alignerr

This freelance role asks mathematicians to turn informal proofs into machine-checkable formal proofs in Lean 4, helping shape how advanced AI systems reason. The work covers theorems from many mathematical fields, contributions to large formal libraries, and auditing existing proofs. It suits holders of advanced mathematics degrees who can work independently on flexible, remote schedules.

OpenRecently verified$170-200/hr
  • Science & Research
  • Platform: Alignerr

This contract role asks mathematicians to write, solve and check advanced problems that help AI models reason more reliably in mathematics. It suits people with a Master's or PhD in mathematics or a closely related field who can work remotely and set their own schedule.

  • Science & Research
  • Platform: Alignerr

This remote, hourly contract role asks mathematicians with a Master's or PhD to write, solve and review advanced problems, and to evaluate the mathematical reasoning produced by AI models. It suits researchers who are strong in algebra, calculus, statistics or discrete mathematics and who can work independently on their own schedule.

  • Science & Research
  • Platform: Alignerr

The job consists of creating, solving and reviewing advanced mathematical problems so that AI models can reason more accurately. It is aimed at people with a Master's or PhD in mathematics or a closely related field who can work remotely and set their own schedule. No prior AI experience is required, but strong mathematical depth and rigor are essential.

  • Engineering
  • Platform: Alignerr

This remote contract role asks mechanical engineers to build and check hard engineering problems that test how AI models reason about physical systems. The work includes writing reference solutions and reviewing AI-generated engineering outputs for accuracy and physical validity. It suits engineers with a graduate degree and solid technical expertise who can explain their reasoning clearly in writing.

  • Engineering
  • Platform: Alignerr

This remote freelance role asks mechanical engineers to help train and evaluate AI models on complex engineering problems. Contributors design multi-domain challenges, write reference solutions, and check AI reasoning for accuracy, safety, and compliance with engineering standards. It suits engineers with a graduate degree who want to apply their expertise to AI development, and no prior AI experience is required.

  • Health & Medicine
  • Platform: Alignerr

This remote contract role asks medical science professionals to review and validate AI-generated biomedical content for accuracy, clarity, and context. It suits people with a background in medical affairs, clinical research, or scientific communication who want to shape how AI handles medical information.

  • Science & Research
  • Platform: Alignerr
  • Level: Intermediate

This freelance remote role asks conservation and environmental science professionals to check how AI systems handle questions about ecosystems, land use, soil, water and biodiversity. The reviewer judges whether the scientific content and recommendations match sound practice, then gives structured, actionable feedback to improve the model's reasoning. It suits experienced practitioners who want to shape AI reliability while working on their own schedule.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This contract role asks security practitioners to build and check realistic scenarios about networks, endpoints, and email threats for training AI systems. The work centers on reviewing firewall settings, EDR telemetry, and infrastructure cases to explain what failed and why. It suits people with several years of hands-on enterprise security operations experience.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Norwegian

This remote contract role asks you to check Norwegian text and translations produced by AI systems for accuracy, naturalness and cultural fit. It is aimed at native or near-native Norwegian speakers with a careful eye for language quality, and no earlier AI background is needed. The work is task-based and flexible in hours.

  • Health & Medicine
  • Platform: Alignerr

This remote freelance role asks registered nurses and clinical informatics professionals to review AI-generated clinical content, EHR workflows and health information outputs. The work is to judge their accuracy, safety and fit with real nursing practice, then give structured feedback that helps healthcare AI systems behave more reliably. It suits clinicians with hands-on health IT experience who can work independently on task-based assignments.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This remote contract role asks analysts to bring offensive security knowledge to AI development without writing exploits or taking part in live intrusions. The work centers on analyzing attack paths and adversary behavior in structured scenarios, then producing labeled reasoning data and written assessments that help AI systems reason about cybersecurity. It suits experienced security professionals who can explain how real attacks unfold in production environments.

  • Health & Medicine
  • Platform: Alignerr

This remote contract role asks oncology clinical research experts to review AI-generated cancer content for scientific accuracy, clinical validity, and regulatory alignment. The work covers trial protocols, safety and efficacy data, biomarker findings, and FDA or EMA-style reporting. It is aimed at professionals with hands-on experience designing and analyzing oncology trials who want to work asynchronously on their own schedule.

  • Data & Machine Learning
  • Platform: Alignerr

This contract role asks security-minded people to probe AI systems built on the OpenClaw platform by running red-teaming exercises and writing adversarial prompts. The work suits cybersecurity practitioners who enjoy finding weaknesses and who can document their findings clearly for technical and non-technical readers.

  • Science & Research
  • Platform: Alignerr

This contract role asks physicists with a PhD to judge and improve the physics material used to train advanced AI systems. The work covers writing hard problems, checking AI-generated solutions for errors, and labeling concepts so models learn the structure of the field. It suits remote experts who can work independently and do not need prior AI experience.

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

679 AI Evaluation AI training jobs are open on SideHustler today. 26 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 637 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (139), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.