AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

679
open jobs
26
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 637 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

661 to 679 of 679 jobs
  • Software Engineering
  • Platform: Mercor
  • Location: United States

Join a network of full-stack engineering experts to collaborate with leading AI laboratories and companies. You will be matched with varied roles consisting of training and evaluating AI models, designing tasks based on real-world scenarios, and providing specialized feedback to advance AI research. Typical engagements range from 15 to 30 hours per week.

Posted February 20, 2026

Talent poolRecently verified$70-150/hr

Referral link: we may earn a fee. Apply without it

  • Health & Medicine
  • Platform: Vetto

This project asks wellness professionals to examine realistic wellness scenarios that will be used to train AI wellness assistants. The work is to spot risks, contraindications, context, and trade-offs, then rewrite weak cases so they test sound professional judgment. It is aimed at practitioners such as psychologists, therapists, dietitians, personal trainers, and strength and conditioning coaches.

Posted February 20, 2026

OpenRecently verifiedPay not disclosed
  • Business, Consulting & Operations
  • Platform: Vetto

This remote task asks experienced teachers to review the scenarios used to train AI teaching assistants and judge which ones truly test how a model teaches. Reviewers then rework weak scenarios so they better reflect real learning interactions. It is aimed at tutors and educators with strong pedagogy, especially in STEM and the Humanities.

Posted February 5, 2026

OpenRecently verifiedPay not disclosed
  • Health & Medicine
  • Platform: Vetto

This project-based role involves reviewing pre-generated health scenarios used to train AI health assistants. Reviewers rate how realistic each scenario is and how much safety risk it carries, then strengthen weak or missing constraints. It is aimed at generalist doctors and experienced nurses with strong patient-facing experience who can tell authentic care-seeking situations from generic ones.

Posted February 5, 2026

OpenRecently verifiedPay not disclosed
  • Finance & Accounting
  • Platform: Vetto
  • Level: Intermediate

This project asks experienced finance professionals to review realistic client scenarios that will be used to train AI financial assistants. The work involves rating existing tasks and rewriting weak ones so that one decisive constraint would change the correct advice. It is aimed at financial planners, investment advisors and bank relationship managers with at least three years of professional practice.

Posted January 29, 2026

OpenRecently verifiedPay not disclosed
  • Software Engineering
  • Platform: Vetto

This job involves checking whether code produced by AI models correctly solves a given software engineering prompt, and whether the accompanying tests really verify that behavior. It is aimed at experienced software engineers who can judge problem definitions and test suites with strict technical standards.

Posted January 27, 2026

OpenRecently verifiedPay not disclosed
  • Data & Machine Learning
  • Platform: Vetto

This role is a hands-on research position focused on post-training large language models, where expert annotations are converted into training data and experimental reward signals. It suits a researcher with a PhD or equivalent experience in machine learning or a related quantitative field who can work independently and iterate fast on messy experiments.

Posted January 23, 2026

OpenRecently verifiedPay not disclosed
  • Software Engineering
  • Platform: Mercor
  • Location: 40 countries
  • Level: Intermediate

This remote-style contract role asks experienced SOC analysts to review and validate security alerts and investigations, separating true positives from false positives across SIEM, endpoint, cloud and identity data. Work includes building ground-truth investigations and judging whether automated or human conclusions are valid, incomplete or incorrect. It suits seasoned Tier 2 or higher analysts who are fluent in Splunk and comfortable making decisive evaluations.

Posted January 18, 2026

Talent poolRecently verified$70-95/hr

Referral link: we may earn a fee. Apply without it

  • Writing, Creative & Design
  • Platform: xAI

This role asks you to review, improve and write polished text across genres and formats, labeling and annotating AI-generated writing so that the system produces clearer, better structured and more accurate prose. It is intended for experienced writers who can prove expertise in at least one writing specialty, such as fiction, technical writing, journalism or copywriting.

Posted January 7, 2026

OpenRecently verifiedPay not disclosed
  • Software Engineering
  • Platform: Mercor
  • Location: United States
  • Level: Expert

A senior-level role at a fundamental AI models laboratory: you will participate in enhancing the capabilities of a model on complex software engineering tasks. This role is suited for senior software engineers with solid experience from leading technology companies.

Posted July 19, 2025

OpenRecently verified$150-210/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: United States
  • Language: German

You will evaluate the quality of German texts generated by a large language model and curate high-level German linguistic data to improve its capabilities. This role is suited for experts with native or near-native fluency in German and professional experience in literary editing or linguistic annotation.

Posted June 18, 2025

OpenRecently verified$50/hr

Referral link: we may earn a fee. Apply without it

  • Software Engineering
  • Platform: Mercor
  • Location: United States

You will build and refine evaluations to test the performance of frontier code models by transforming pull requests into engineering tasks. You will use a proprietary evaluation framework to validate how coding agents solve complex problems, identifying real weaknesses in their implementations. This role is for software engineers with solid fundamentals, capable of navigating substantial codebases and producing rigorous evaluations.

Posted June 2, 2025

OpenRecently verified$100-120/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: United States

This role involves evaluating and improving the performance of a language model on English language tasks. You will annotate texts, edit content generated by AI, and provide detailed feedback to refine its understanding and generation capabilities. The ideal profile is a native or near-native English speaker with professional experience in editing, proofreading, linguistic annotation, or data evaluation for natural language processing.

Posted June 2, 2025

OpenRecently verified$50/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Handshake AI
  • Location: United States

Test and evaluate Large Language Models through remote, part-time projects in collaboration with AI research labs. Apply your analytical skills and reasoning abilities to help improve AI model performance.

Posted March 10, 2025

  • Finance & Accounting
  • Platform: Handshake AI
  • Location: United States
  • Level: Intermediate

This position involves evaluating AI-generated content related to retail finance and financial planning practices. Professionals with hands-on FP&A experience will provide feedback to help AI systems better understand retail finance workflows, budget management, and performance analysis.

Posted March 10, 2025

  • Software Engineering
  • Platform: Handshake AI
  • Location: United States
  • Level: Entry level

A remote position for game developers and designers with experience in game engines and development tools to contribute to AI research projects. Participants will develop domain-specific prompts and evaluate large language model responses to improve how AI understands professional tasks in game development.

Posted March 10, 2025

OpenRecently verifiedUp to $125/hr
  • Data & Machine Learning
  • Platform: Handshake AI
  • Location: United States
  • Level: Entry level

Handshake AI seeks skilled PCB and EDA tool users to evaluate AI-generated content and provide feedback on electronics design tasks. Contributors will work flexibly on project-based contract work that involves assessing large language model responses related to PCB layout, schematic capture, and embedded hardware design.

Posted March 10, 2025

OpenRecently verifiedUp to $125/hr
  • Science & Research
  • Platform: Handshake AI
  • Location: United States

Philosophy PhDs evaluate AI-generated content and provide feedback to improve AI's understanding of philosophy reasoning and theoretical analysis. This is flexible, remote, hourly contract work that can be combined with research, teaching, or other employment.

Posted March 10, 2025

OpenRecently verifiedUp to $120/hr
  • Software Engineering
  • Platform: Handshake AI
  • Location: United States
  • Level: Entry level

Experienced software engineers are needed to support AI research through flexible, part-time contract work. You will evaluate AI-generated content, create expert-level software engineering problems, and provide structured feedback to help AI models improve their reasoning about production systems and architecture.

Posted March 10, 2025

OpenRecently verifiedUp to $120/hr

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

679 AI Evaluation AI training jobs are open on SideHustler today. 26 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 637 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (139), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.