AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

679
open jobs
26
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 637 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

221 to 240 of 679 jobs
  • Finance & Accounting
  • Platform: Alignerr

This is a remote, hourly contract role for experienced traders and market followers who want to help train AI systems on retail trading behavior and prediction markets. The work involves applying real market intuition to structured tasks such as estimating event probabilities, interpreting sentiment signals, and reviewing AI-generated market analysis.

  • Software Engineering
  • Platform: Alignerr
  • Level: Expert

This contract role centers on building and tuning high-performance Rust systems that power AI data pipelines, annotation tools, and model evaluation workflows. It suits senior engineers with at least five years of production Rust experience who care about performance, concurrency, and correctness in large-scale systems.

  • Software Engineering
  • Platform: Alignerr
  • Level: Expert

This is a remote, hourly contract role for senior software engineers who will help build, test and improve autonomous AI coding agents. The work centers on judging whether AI-written code is correct, efficient, secure and well designed, and on finding where these agents fail. It suits experienced engineers who can read unfamiliar codebases and explain clearly why code works or breaks.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This remote contract role involves designing and building the software that measures how well AI models perform, including evaluation pipelines, automated testing harnesses, and internal dashboards and APIs. It suits senior engineers who have shipped production systems and want to work closely with AI research teams on reliable, repeatable evaluation tooling.

  • Software Engineering
  • Platform: Alignerr
  • Level: Expert

This contract role is for senior engineers who build the backend systems behind AI model development. The work covers high-performance pipelines, annotation tooling, and evaluation infrastructure that support model training and benchmarking. It is aimed at engineers with several years of production experience in Go, Rust, Python, or C++ who are fully remote and available 20 to 40 hours per week.

  • Writing, Creative & Design
  • Platform: Alignerr

This remote contract role asks you to judge posts, captions and comments written by AI, checking whether they sound natural, fit the culture of each platform and hold attention. It suits people who use social media actively and have a good instinct for what works online, with no marketing or content creation background needed.

  • Software Engineering
  • Platform: Alignerr

This freelance role asks you to write code and to evaluate code produced by AI models, helping make AI coding assistants more accurate and reliable. You will review AI outputs for correctness, efficiency, and security, then explain your judgments in clear written feedback. It suits experienced programmers who can work independently on a flexible, remote schedule.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This remote contract role asks experienced software engineers to assess and rank code produced by AI models. The work involves judging correctness, efficiency and maintainability, then writing structured feedback that helps coding assistants learn what good code looks like. It suits developers who prefer flexible, independent, task-based work.

  • Software Engineering
  • Platform: Alignerr

This contract role asks you to review code produced by AI models and to write backend solutions yourself, covering server-side logic and API design at several difficulty levels. It suits backend engineers who can judge code on correctness, efficiency and clarity and explain their reasoning in clear English. The work is fully remote and asynchronous, with a minimum of 15 hours per week.

  • Software Engineering
  • Platform: Alignerr

This contract asks frontend engineers to review and improve JavaScript, HTML, and CSS code written by AI models, checking correctness, efficiency, and accessibility. It is fully remote and asynchronous, suited to developers who can explain their reasoning clearly in English.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This is a remote contract role building internal C tooling that supports data pipelines, annotation systems, and evaluation workflows for AI training. It is aimed at senior engineers with several years of production C experience who can also work across full-stack and interoperability tasks.

  • Software Engineering
  • Platform: Alignerr

This contract role asks machine learning engineers to review and evaluate machine learning code generated by AI models, checking correctness, efficiency, scalability and clarity. Contractors also write their own ML solutions and explain modeling choices so that frontier AI systems improve at this work. It suits experienced ML practitioners who prefer fully remote, asynchronous, self-directed project work.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This remote contract role asks Ruby engineers to review, evaluate and improve backend coding scenarios that are used to train AI systems. It suits experienced developers who can judge code quality, explain their reasoning clearly and work independently on task-based assignments.

  • Software Engineering
  • Platform: Alignerr

This remote, task-based contract involves designing realistic engineering challenges that evaluate advanced AI agents on enterprise work such as diagnosing failing integrations, tracing bugs, and writing fixes. It is aimed at working software engineers who debug real systems, write infrastructure and configuration code, and can verify their fixes in live or simulated environments.

  • Software Engineering
  • Platform: Alignerr

This remote contract role involves running code written by AI, checking whether it behaves as intended, and recording bugs and quality issues in structured reports. It suits people who can read and run simple programs in at least one language and think logically, without needing a formal QA or computer science background.

  • Science & Research
  • Platform: Alignerr
  • Level: Intermediate

This remote freelance role asks soil and water specialists to check AI-generated scenarios on erosion control, watershed management, and conservation planning. Reviewers flag flawed assumptions and outdated methods, then write structured feedback so the AI's answers become more accurate and practical. It is aimed at scientists with several years of hands-on experience who can work independently on task-based assignments.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Spanish

This job asks native or heritage speakers of Peninsular Spanish to help train and check AI speech models. The work involves listening to generated audio, transcribing and labeling samples, and writing English feedback on accuracy and naturalness. It suits people with a trained ear for European Spanish who want flexible remote work and no prior AI experience.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Spanish

This role consists of reviewing Spanish text produced by AI, including translations and original content, to judge its accuracy, naturalness and cultural fit for Spanish speakers. It is aimed at native or near-native Spanish speakers with a careful, methodical approach, and the work is done remotely on a flexible, task-based schedule. No prior AI or technology experience is required.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Spanish

This remote freelance role asks Spanish speakers to review AI-generated Spanish text and judge its tone, register and regional fit. The work includes proposing more natural, culturally suited wording and giving structured feedback that helps AI systems communicate in Spanish. It suits people with native-level Spanish and a solid grasp of regional variants, and no prior AI experience is required.

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

679 AI Evaluation AI training jobs are open on SideHustler today. 26 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 637 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (139), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.