AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

679
open jobs
26
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 637 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

461 to 480 of 679 jobs
  • Health & Medicine
  • Platform: Mercor
  • Location: United States

Mercor is recruiting physicians (M.D./D.O.) to evaluate and train AI models on clinical reasoning. You will be called upon to perform annotation, scoring, and algorithm validation tasks on an as-needed basis according to your areas of specialty, while explaining your medical expertise.

Posted September 15, 2026

OpenRecently verified$115-150/hr

Referral link: we may earn a fee. Apply without it

  • Business, Consulting & Operations
  • Platform: Mercor
  • Location: United States

This job involves evaluating and testing how AI models use office and productivity software. You put your practical knowledge to work improving these tools by writing realistic tasks, rating the results generated by AI, and reporting compatibility or robustness issues. This job is for professionals with solid expertise in these software programs who want to participate in AI evaluation.

Posted September 15, 2026

OpenRecently verified$60-70/hr

Referral link: we may earn a fee. Apply without it

  • Science & Research
  • Platform: Mercor
  • Location: United States

Mercor is recruiting physicists to evaluate and rate responses generated by advanced AI models. You will write physics problems, verify model derivations and reasoning, and provide critical expertise based on your research experience.

Posted September 15, 2026

OpenRecently verified$70-100/hr

Referral link: we may earn a fee. Apply without it

  • Software Engineering
  • Platform: Mercor
  • Location: United States

Mercor is looking for experienced software engineers to evaluate and rate code generated by AI models. You will be asked to write engineering problem statements, examine AI solutions for production robustness, and identify hidden defects. This role is aimed at developers with solid industrial experience and excellent ability to document their technical analysis.

Posted September 15, 2026

OpenRecently verified$70-150/hr

Referral link: we may earn a fee. Apply without it

  • Business, Consulting & Operations
  • Platform: micro1
  • Location: 58 countries

You analyze textual, audio and video data to identify and categorize subtle behavioral indicators such as tone, gestures and speech patterns, in order to support the development of AI systems. This role is aimed at professionals from behavioral consulting, coaching or theater backgrounds who can interpret the nuances of human interactions.

Posted September 15, 2026

OpenRecently verified$20-70/hr

Referral link to micro1's job list: search for this role there. Open this exact job

  • Science & Research
  • Platform: Mercor
  • Location: United States

You will write test prompts to evaluate whether AI models can distinguish legitimate chemical requests from potentially dangerous misuses. As an analytical chemist, you will develop scenarios along the dual-use line, assess model responses against a defined policy, and document your analyses with accessible technical reasoning.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Finance & Accounting
  • Platform: Mercor
  • Location: United States
  • Level: Expert

You will analyze official communications from a non-G7 central bank (statements, minutes, speeches) to evaluate and improve AI models in macroeconomics. Based on fixed time cutoffs and without hindsight knowledge, you will interpret monetary policy signals, compare them to market expectations, and provide calibrated analyses on the central bank's stance. This role requires deep expertise in a specific central bank and mastery of its communication language.

Posted September 14, 2026

OpenRecently verified$150-250/hr

Referral link: we may earn a fee. Apply without it

  • Science & Research
  • Platform: Mercor
  • Location: United States

You will work on evaluating frontier AI models to test their ability to distinguish legitimate chemical requests from malicious uses. As a chemical defense researcher, you will write targeted questions, evaluate model responses against a defined policy framework, and document your judgments with rigorous technical justifications. This role requires deep expertise in chemical defense to draw the line between routine professional questions and those with high potential for misuse.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Engineering
  • Platform: Mercor
  • Location: United States

You will contribute to evaluating AI models as an expert in explosives and energetic materials. Your role involves writing sharp technical questions in your field, evaluating model responses against a safety policy, and documenting correct answers with technical justification. This work requires the ability to distinguish legitimate professional questions from suspicious requests that could serve malicious purposes.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Science & Research
  • Platform: Mercor
  • Location: United States

You will contribute to evaluating AI models for radiological safety by writing test requests at different difficulty levels and judging whether the model's responses meet safety standards. This role requires recognized expertise in radiological crisis management, capable of distinguishing legitimate professional questions from potentially dangerous requests.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Science & Research
  • Platform: Mercor
  • Location: United States

You will write and evaluate requests designed to test the ability of AI models to distinguish legitimate technical questions from dangerous requests in the field of energetic materials. This role requires deep expertise in post-incident analysis and regulatory inspection to draw the line between routine professional uses and potential misuse.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Science & Research
  • Platform: Mercor
  • Location: United States

You will contribute to the evaluation of frontier AI models as an expert in nuclear materials and non-proliferation measures. Your job is to write sophisticated test questions in your field and evaluate whether models correctly refuse dangerous requests while answering legitimate questions, drawing a subtle line between civil and military applications.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Health & Medicine
  • Platform: Mercor
  • Location: United States

You will join a team of chemistry and chemical safety experts tasked with testing the capabilities of frontier AI models to assess risks of misuse of legitimate purposes. Drawing on your expertise in forensic chemistry or toxicology, you will write sophisticated test questions, evaluate model responses against defined policy criteria, and develop reference answers with technical justifications.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Science & Research
  • Platform: Mercor
  • Location: United States

This role involves assessing the ability of AI models to correctly judge the potential for misuse of technical requests concerning energetic materials and propulsion. As a specialist, you will write progressively complex questions, evaluate model responses against a defined policy, and document your analyses to guide the behavior of AI systems when facing real-world risks.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Health & Medicine
  • Platform: Mercor
  • Location: United States

You will help test the robustness of AI models against malicious risks as a radiation protection specialist. Your role consists of designing deliberately ambiguous technical requests, evaluating the model's responses according to a defined security policy, and documenting best practices, while knowing how to distinguish legitimate professional questions from malicious intentions.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Health & Medicine
  • Platform: Mercor
  • Location: United States

You will evaluate how AI models handle sensitive requests in industrial hygiene and process safety, distinguishing legitimate professional questions from malicious bypass attempts. You will write graduated test requests, assess model responses against a defined policy standard, and document your judgments with technical reasoning accessible to non-specialists. This role is for experienced practitioners who have evaluated real exposures and process risks in professional settings.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Legal
  • Platform: Mercor
  • Location: United States

You will write and evaluate test cases to verify that an AI model knows how to refuse dangerous requests while answering legitimate questions in the field of nuclear nonproliferation. This role requires professional expertise in export controls, treaty implementation, or nuclear program analysis, so you can draw the line between routine compliance questions and circumvention attempts.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Engineering
  • Platform: Mercor
  • Location: United States

You will contribute to evaluating frontier AI models as a nuclear fuel cycle expert. Your role will consist of writing nuanced technical questions and judging whether the model's answers comply with a safety policy, distinguishing between legitimate requests and potentially dangerous queries. This work requires a deep understanding of the parameters that differentiate a civilian application from a diversion for proliferation purposes.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Health & Medicine
  • Platform: Mercor
  • Location: United States

This job involves evaluating the capacity of AI models to distinguish legitimate requests from malicious uses in the field of nuclear medicine and medical isotopes. You will write nuanced test prompts and judge whether the model's responses comply with a security policy, while explaining your decisions in writing. This role requires specialized expertise to draw the line between routine professional questions and those concealing dangerous intentions.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Science & Research
  • Platform: Mercor
  • Location: United States

You will participate in evaluating cutting-edge AI models as a nuclear security specialist. You will write calibrated test prompts across three risk levels to judge whether the model correctly distinguishes legitimate professional questions from dangerous requests, then you will evaluate and document its responses according to a defined policy standard.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

679 AI Evaluation AI training jobs are open on SideHustler today. 26 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 637 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (139), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.