1234 open jobs

Remote AI training jobs

Open roles in AI training, model evaluation, RLHF and data labeling for domain experts. Search, filter, then apply directly on Mercor, micro1, Turing, Ethos, Mindrift, DataAnnotation, Alignerr, Handshake AI, Surge AI, xAI, Vetto or Terac.

DataAnnotation×Clear all
81 to 100 of 113 jobs
  • Health & Medicine
  • Platform: DataAnnotation

A physician evaluates AI model responses on clinical questions against current standards of care, identifying unsafe or outdated guidance and writing correct clinical answers when needed. This role requires assessing AI clinical reasoning and creating prompts that test differential diagnosis, treatment planning, and patient communication.

Talent poolRecently verified$40-125/hr
  • Science & Research
  • Platform: DataAnnotation

This role involves evaluating AI models' physics reasoning across mechanics, electromagnetism, quantum mechanics, and thermodynamics. You will create problems to test AI understanding, identify errors in AI solutions, and write gold-standard solutions that clearly explain the physics at each step.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Polish

As a Polish Specialist, you will evaluate and correct AI-generated Polish text, assessing translations for accuracy and naturalness while teaching models to produce native-sounding output. You'll work with a focus on Polish grammar complexities like cases, verb aspect, and formality registers, drawing on native speaker intuition to improve model behavior.

Talent poolRecently verified$25-40/hr
  • Finance & Accounting
  • Platform: DataAnnotation

Evaluate investment theses and portfolio allocations built by AI models, identifying risks and construction flaws that algorithms miss. Write detailed assessments that teach AI systems how professional capital allocation actually works. This role is for experienced investment professionals who can catch what models overlook.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Portuguese

This role involves evaluating and correcting AI-generated Portuguese translations and text outputs to improve model quality. As a Portuguese specialist, you will assess translations for accuracy, fluency, and cultural appropriateness, and write improved versions when needed, helping train AI models to produce natural-sounding Portuguese.

Talent poolRecently verified$25-40/hr
  • Finance & Accounting
  • Platform: DataAnnotation

Evaluate and critique AI-generated private equity LBO models and financial diligence by underwriting deals as if personal capital were at stake. Assess entry and exit multiples, debt structures, and return drivers, then provide written feedback that teaches AI systems how real private-market transactions are analyzed. This role is for finance professionals with transaction experience who can identify model shortcomings and articulate investment-committee-level judgments.

Talent poolRecently verified$40-125/hr
  • Finance & Accounting
  • Platform: DataAnnotation

A public accountant evaluates AI-generated audit and assurance work against accounting standards to catch flawed reasoning, unsupported conclusions, and misapplied procedures. You'll design audit scenarios, review AI responses for compliance with GAAP and GAAS, and write correct approaches when the model's work falls short.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Punjabi

Evaluate artificial intelligence systems on Punjabi translation, localization, and cultural accuracy by reviewing model outputs, identifying errors, and providing corrected translations. This role is ideal for Punjabi linguists who want to contribute to training next-generation language models while working flexibly on real-world projects.

Talent poolRecently verified$25-40/hr
  • Finance & Accounting
  • Platform: DataAnnotation

This role involves stress-testing AI models' financial reasoning by writing prompts that probe quantitative analysis and evaluating their outputs for errors in pricing, risk assessment, portfolio construction, and statistical modeling. You will review AI-generated financial analysis for flaws in derivations and model application, and write correct analyses to serve as training data.

Talent poolRecently verified$40-125/hr
  • Health & Medicine
  • Platform: DataAnnotation

This role asks radiologists and imaging technologists to read de-identified scans and judge how accurately frontier AI models interpret them. Contributors write structured reports and impressions, then compare them with model output to find errors and gaps. It suits clinicians with hands-on imaging experience who can explain why a machine-generated report is wrong.

Talent poolRecently verified$40-125/hr
  • Health & Medicine
  • Platform: DataAnnotation

This role involves evaluating AI models' clinical reasoning on nursing topics including assessment, medication administration, and patient care protocols. Registered nurses will identify gaps between textbook knowledge and real bedside practice, write test prompts to probe nursing judgment, and provide corrected guidance when AI outputs contain unsafe or incorrect clinical information.

Talent poolRecently verified$40-125/hr
  • Business, Consulting & Operations
  • Platform: DataAnnotation

This role involves evaluating AI-generated answers on healthcare revenue cycle topics including claims, denials, and payer operations. You will write test prompts to probe the AI's understanding of real-world billing mechanics and provide expert corrections and guidance based on hands-on revenue cycle experience.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Romanian

This role involves evaluating AI model outputs for Romanian language tasks, including translation and localization. Specialists review AI-generated content, identify errors, explain failures, and provide corrected versions to help train AI systems.

Talent poolRecently verified$25-40/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Russian

Evaluate and correct AI-generated Russian text, focusing on grammar, verb aspect, formality levels, and cultural appropriateness. You will assess machine translations and model outputs to help AI systems produce natural-sounding Russian that native speakers recognize as authentic, not machine-translated.

Talent poolRecently verified$25-40/hr
  • Engineering
  • Platform: DataAnnotation
  • Level: Intermediate

A Senior Mechanical Design Engineer will create rigorous engineering problems and solutions to evaluate and improve AI model reasoning across mechanical specialties. You will write closed-ended technical challenges from your expertise, derive correct answers, test them against AI systems, and review AI-generated engineering work for real-world accuracy and applicability.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Slovak

This role involves judging AI output on Slovak translation and localization, including whether the text reads naturally and fits its cultural context. Reviewers pinpoint errors, explain why a model went wrong, and draft stronger answers when needed. It suits professionals with Slovak linguistic expertise and clear written English, since their explanations shape how models learn.

Talent poolRecently verified$25-40/hr
  • Software Engineering
  • Platform: DataAnnotation

You will evaluate AI-generated code by running models on real engineering tasks, analyzing their output against production standards, and stress-testing them to find failures. This role is ideal for experienced software engineers who want to contribute to AI model improvement through rigorous code review and red-teaming.

Talent poolRecently verified$40-150/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Spanish

As a Spanish Specialist, you will evaluate and correct AI-generated Spanish translations and text for quality, cultural appropriateness, and naturalness. You will assess translations between English and Spanish, rank model outputs, and write improved versions when needed, serving as a training signal to teach AI models to produce authentic Spanish that sounds like native speakers wrote it.

Talent poolRecently verified$25-40/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Swedish

Evaluate and improve AI model outputs on Swedish translation and localization tasks. You will assess the quality of language models' work, identify errors and cultural nuances, and provide corrected examples to help train better AI systems.

Talent poolRecently verified$25-40/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Filipino

This role involves evaluating and improving AI-generated Tagalog text, with focus on linguistic authenticity, natural verb forms, and respect markers. The position is designed for native or near-native Tagalog speakers who can identify and correct issues like unnatural formality, incorrect politeness markers, and awkward code-switching between Tagalog and English.

Talent poolRecently verified$25-40/hr

FAQ

Questions about AI training jobs

What kinds of jobs are listed here?

AI training, model evaluation, RLHF, red teaming, data labeling and expert review roles, in fields from software and data science to medicine, law, finance, languages and science.

Are these jobs remote?

They are done online. Some are limited to residents of certain countries: when the platform publishes that restriction, the listing shows it.

Do I need prior experience in AI?

Usually not. Most roles ask for professional experience in your own field. Each listing states the experience the platform requires.

How is the work paid?

By the hour or by the task, by the platform that hires you. We show the pay only when the platform publishes it. It is not a guarantee of income or hours.

How do I apply?

Apply now opens the official posting on Mercor, micro1, Turing, Ethos, Mindrift, DataAnnotation, Alignerr, Handshake AI, Surge AI, xAI, Vetto or Terac, where you complete the application. You need no account here. See how it works.