1234 open jobs

Remote AI training jobs

Open roles in AI training, model evaluation, RLHF and data labeling for domain experts. Search, filter, then apply directly on Mercor, micro1, Turing, Ethos, Mindrift, DataAnnotation, Alignerr, Handshake AI, Surge AI, xAI, Vetto or Terac.

DataAnnotation×Clear all
61 to 80 of 113 jobs
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Korean

This role involves evaluating and correcting AI-generated Korean language output, focusing on speech levels, honorifics, particles, and translation quality. The position is designed for native or near-native Korean speakers who can identify subtle linguistic errors that AI models commonly make and provide corrected versions that sound natural.

Talent poolRecently verified$25-40/hr
  • Legal
  • Platform: DataAnnotation

As a legal expert, you will evaluate AI-generated legal analysis for accuracy, identifying errors such as invented citations, misapplied precedent, and jurisdictional confusion. You will draft test prompts covering contracts, statutes, procedure, and case analysis, then write corrected legal analysis when the model's reasoning falls short.

Talent poolRecently verified$40-125/hr
  • Legal
  • Platform: DataAnnotation

Evaluate AI-generated legal work on litigation matters including pleadings, discovery, and motions by checking citations, holdings, procedural rules, and standards of review. Create test prompts for litigation scenarios and rewrite AI outputs to meet court-quality standards when necessary.

Talent poolRecently verified$40-125/hr
  • Data & Machine Learning
  • Platform: DataAnnotation

As a Machine Learning Engineer, you will evaluate and improve how AI models reason about machine learning systems, including training dynamics, evaluation design, and deployment strategies. You will write prompts to test model reasoning, identify subtle errors in AI outputs, and provide correct solutions based on real practitioner experience.

Talent poolRecently verified$40-150/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Malay

This role involves evaluating and correcting AI-generated Malay text, focusing on register (formal versus colloquial), regional accuracy (Malaysian versus Indonesian), and proper use of affixes. You will review translations, score model outputs, and write improved versions to train the model to produce natural, native-quality Malay.

Talent poolRecently verified$25-40/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Malayalam

Evaluate AI model outputs for Malayalam translation, localization, and cultural accuracy on real-world projects. Review frontier models, identify errors, rate quality, and write corrections to help train the next generation of AI systems.

Talent poolRecently verified$25-40/hr
  • Business, Consulting & Operations
  • Platform: DataAnnotation

This role involves evaluating AI systems on their understanding of managed care operations, including benefit determination, prior authorization, and payer contracts. You'll identify errors in AI-generated coverage decisions and write correct guidance based on real-world managed care practices to improve AI training.

Talent poolRecently verified$40-125/hr
  • Business, Consulting & Operations
  • Platform: DataAnnotation

This role asks experienced consultants to test cutting-edge AI models on business problems drawn from their own practice. Contributors write client-style briefs, review each response as they would an associate's work, and flag where recommendations are sharp and where they are generic. It suits people with consulting, strategy, PMO, or deal advisory backgrounds.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Marathi

Evaluate AI model outputs on Marathi translation and localization tasks, identifying errors and cultural nuances. You will rate the quality of frontier models' Marathi work, explain where they fail, and provide corrected answers to help train future AI systems. This is a flexible, remote position for Marathi language specialists.

Talent poolRecently verified$25-40/hr
  • Science & Research
  • Platform: DataAnnotation

This role involves evaluating and improving how AI models handle rigorous mathematical reasoning. You will create challenging math problems, review AI solutions for logical errors and unjustified steps, and write correct solutions that serve as training examples for the models to learn from.

Talent poolRecently verified$40-125/hr
  • Engineering
  • Platform: DataAnnotation
  • Level: Intermediate

This role involves creating rigorous mechanical engineering problems and solutions to train AI models to reason like practicing engineers. You will write verifiable problems from your specialty, derive correct answers, test them against AI systems, and review AI-generated engineering work for real-world applicability.

Talent poolRecently verified$40-125/hr
  • Engineering
  • Platform: DataAnnotation
  • Level: Intermediate

A mechanical designer writes rigorous engineering problems and solutions to train AI models to reason like practicing engineers. You'll create challenging, closed-ended technical problems from your specialty, verify correct answers, test them against AI systems, and review AI-generated engineering work for real-world applicability.

Talent poolRecently verified$40-125/hr
  • Engineering
  • Platform: DataAnnotation
  • Level: Intermediate

A mechanical engineer will create challenging, closed-ended engineering problems based on their specialty, derive correct solutions, and evaluate how well AI models perform on real-world engineering tasks. The role involves writing problems that test genuine engineering reasoning, component sizing, failure analysis, CAD interpretation, design impact prediction, and assessing AI-generated analysis and designs for correctness and practical applicability.

Talent poolRecently verified$40-125/hr
  • Health & Medicine
  • Platform: DataAnnotation

Medical coders evaluate AI-generated medical coding outputs against real clinical documentation and coding guidelines, identifying errors in ICD-10-CM, CPT, HCPCS codes, modifiers, and E&M levels. The role requires writing test prompts to assess coding accuracy and providing corrected assignments with clear documentation-based explanations to serve as training signals for AI models.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Norwegian

Evaluate how well AI models perform on Norwegian translation, localization, and cultural adaptation tasks. The role involves rating the quality of AI-generated translations, identifying errors, providing corrections, and offering detailed explanations to improve model training.

Talent poolRecently verified$25-40/hr
  • Health & Medicine
  • Platform: DataAnnotation

This role involves evaluating how AI models handle nursing judgment in clinical scenarios such as patient assessment, medication safety, and triage protocols. Nurses will identify unsafe reasoning, incorrect clinical priorities, and protocol violations, then provide corrections grounded in real bedside practice.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Odia

Evaluate AI models on Odia translation and localization tasks, assessing quality and cultural appropriateness of machine-generated content. You will review model outputs, identify errors, provide corrections, and generate training data to improve future AI systems.

Talent poolRecently verified$25-40/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation

This role involves evaluating and correcting AI-generated content in languages other than English. As a Bilingual Specialist, you will review translations, score model outputs, and write improved versions to teach AI systems to handle idioms, tone, and grammar the way native speakers actually use the language.

Talent poolRecently verified$25-40/hr
  • Legal
  • Platform: DataAnnotation

This role involves evaluating AI systems' performance on legal research, document review, and litigation filings by checking for errors like invented citations and overruled authorities. As a paralegal, you'll draft test prompts, review AI outputs for accuracy, and produce correct work product when the model falls short.

Talent poolRecently verified$40-125/hr
  • Software Engineering
  • Platform: DataAnnotation

This role involves evaluating how AI models reason about offensive security concepts and identifying flaws in their exploit chain reasoning. Penetration testers with hands-on red team experience will test model outputs across reconnaissance, exploitation, privilege escalation, and lateral movement, then write accurate attack paths that reflect real-world tradecraft when models fall short.

Talent poolRecently verified$40-125/hr

FAQ

Questions about AI training jobs

What kinds of jobs are listed here?

AI training, model evaluation, RLHF, red teaming, data labeling and expert review roles, in fields from software and data science to medicine, law, finance, languages and science.

Are these jobs remote?

They are done online. Some are limited to residents of certain countries: when the platform publishes that restriction, the listing shows it.

Do I need prior experience in AI?

Usually not. Most roles ask for professional experience in your own field. Each listing states the experience the platform requires.

How is the work paid?

By the hour or by the task, by the platform that hires you. We show the pay only when the platform publishes it. It is not a guarantee of income or hours.

How do I apply?

Apply now opens the official posting on Mercor, micro1, Turing, Ethos, Mindrift, DataAnnotation, Alignerr, Handshake AI, Surge AI, xAI, Vetto or Terac, where you complete the application. You need no account here. See how it works.