AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

679
open jobs
26
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 637 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

241 to 260 of 679 jobs
  • Science & Research
  • Platform: Alignerr

This remote, asynchronous role involves labeling and organizing statistical content, such as equations, datasets and problems, to build training data for AI systems. It also includes reviewing AI-generated statistics solutions, writing instructional material and testing models for reasoning errors and bias. It is aimed at people with strong probability and statistics backgrounds from research, teaching or applied work.

  • Engineering
  • Platform: Alignerr

This remote contract role asks structural and civil engineers to turn their practical know-how into AI training tasks. Each task is a focused engineering problem with a clear problem statement, a deterministic checker, and a verified reference solution. It suits practicing engineers who use tools such as OpenSees and know design codes, and no AI background is needed.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This is a remote hourly contract for a senior Rust engineer who will design and maintain high-performance data pipelines, annotation tooling, and evaluation systems used to train AI models. The work is production-focused, involves backend and tooling services, and suits engineers who can work independently with research and data teams.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Tamil

This contract role asks Tamil speakers to review AI-generated Tamil text and original content for accuracy, naturalness and cultural fit. The work is remote, hourly and done asynchronously, with structured feedback on mistranslations, awkward phrasing and register mismatches. It suits people with native or near-native Tamil who are detail-oriented and need no prior AI experience.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This freelance remote contract asks threat intelligence and security professionals to help train and assess AI systems on cybersecurity topics. The work centers on checking adversary campaigns, attack chains and indicators of compromise, and on judging whether AI-generated security content is accurate and realistic.

  • Finance & Accounting
  • Platform: Alignerr

This contract role involves reviewing AI-generated content on trade processing, market structure, and operational workflows to check accuracy and logical consistency. It is aimed at professionals with hands-on experience in trading, brokerage, or financial operations who can work independently on flexible, task-based assignments. No prior AI experience is required.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This remote contract role asks senior TypeScript developers to review AI-generated code for type safety, scalability and modern practices. The work centers on writing structured feedback that helps AI models understand architectural choices and type logic. It is aimed at experienced developers who can work independently on flexible hourly commitments.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This contract role asks TypeScript developers to review and improve AI-generated code so it is type-safe, idiomatic, scalable and ready for production. It is aimed at senior developers with deep TypeScript knowledge who can clearly explain why a piece of code is well designed. No prior AI background is required.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Urdu

This contract role asks Urdu speakers to assess and refine text generated by AI systems, checking that it is accurate, reads naturally, and suits real Urdu readers. It is aimed at people with native or near-native Urdu and a careful, methodical eye for written quality, and no prior AI background is needed. The work is remote and performed asynchronously on the contractor's own schedule.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Vietnamese

This remote contract role asks you to check text produced by AI systems in Vietnamese and judge whether it reads naturally, is accurate, and fits Vietnamese cultural expectations. It suits native or near-native Vietnamese speakers with a careful, methodical approach to written language, and no prior AI experience is needed.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: German

This remote, contract-based role asks German-speaking voice professionals to record scripted German audio and judge AI-generated speech for naturalness, pronunciation, and expressiveness. It suits people with voice acting or narration experience who want to help AI systems sound more human, and no prior AI background is needed.

OpenRecently verified$250-280/hr
  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This remote contract role asks security practitioners to assess and improve AI systems on cybersecurity reasoning. The work centers on judging vulnerability scenarios, severity and remediation choices so the AI reflects how security teams actually operate. It suits people with hands-on vulnerability management experience who can explain security decisions clearly in writing.

  • Science & Research
  • Platform: Alignerr
  • Level: Intermediate

This remote, task-based role asks scientists to check AI-generated material on wildlife and habitat conservation for accuracy and ecological soundness. Reviewers judge the quality of the reasoning behind proposed conservation strategies and give structured feedback so the AI describes these topics correctly. It suits experienced ecologists and wildlife specialists who can work independently on a flexible schedule.

  • Finance & Accounting
  • Platform: DataAnnotation

This role involves evaluating and improving AI models' accounting capabilities by grading their responses against real-world accounting standards and practices. Accountants will write test prompts, review AI outputs for errors in bookkeeping and tax scenarios, and provide correct answers with detailed explanations to help train the models.

Talent poolRecently verified$40-125/hr
  • Legal
  • Platform: DataAnnotation

This role involves evaluating and improving AI legal reasoning by grading model outputs against professional standards, identifying errors like invented citations and misapplied precedent, and writing correct legal analysis when needed. It is suited for attorneys and legal professionals who can assess AI-generated legal content and provide high-quality written corrections.

Talent poolRecently verified$40-125/hr
  • Legal
  • Platform: DataAnnotation

This role evaluates how AI models handle audit procedures, internal controls, and assurance judgments by testing their reasoning, identifying gaps in professional skepticism, and writing correct audit approaches. It is for experienced auditors who can assess whether AI-generated evidence actually supports audit opinions.

Talent poolRecently verified$40-125/hr
  • Health & Medicine
  • Platform: DataAnnotation

This role involves evaluating how well AI models reason through cardiology cases, identifying errors in clinical judgment that could harm patients. Cardiologists will write prompts to test AI reasoning, review AI outputs for accuracy against current guidelines, and provide correct answers when the model errs, serving as the training signal for AI improvement.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Catalan

Evaluate AI model performance on Catalan translation and localization tasks, assessing translation quality and cultural appropriateness. The role involves reviewing AI-generated outputs, identifying errors, and providing corrections and explanations to improve model training.

Talent poolRecently verified$25-40/hr
  • Science & Research
  • Platform: DataAnnotation

This role involves evaluating AI-generated chemical reasoning and responses to identify flawed mechanisms, unsafe procedures, and analytical errors. You will write prompts to test chemical knowledge, review AI outputs for errors, and provide correct expert-level answers with proper attention to safety and experimental conditions.

Talent poolRecently verified$40-125/hr
  • Health & Medicine
  • Platform: DataAnnotation

A Clinical Documentation Specialist evaluates AI-generated medical documentation for integrity, compliance, and adherence to coding and privacy standards. The role requires identifying documentation gaps, missed query opportunities, and HIPAA violations, then providing corrections based on CDI best practices.

Talent poolRecently verified$40-125/hr

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

679 AI Evaluation AI training jobs are open on SideHustler today. 26 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 637 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (139), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.