1234 open jobs

Remote AI training jobs

Open roles in AI training, model evaluation, RLHF and data labeling for domain experts. Search, filter, then apply directly on Mercor, micro1, Turing, Ethos, Mindrift, DataAnnotation, Alignerr, Handshake AI, Surge AI, xAI, Vetto or Terac.

DataAnnotation×Clear all
1 to 20 of 113 jobs
  • Finance & Accounting
  • Platform: DataAnnotation

This role involves evaluating and improving AI models' accounting capabilities by grading their responses against real-world accounting standards and practices. Accountants will write test prompts, review AI outputs for errors in bookkeeping and tax scenarios, and provide correct answers with detailed explanations to help train the models.

Talent poolRecently verified$40-125/hr
  • Finance & Accounting
  • Platform: DataAnnotation

This role involves evaluating and validating actuarial models built by AI systems, assessing their reserves, pricing methodologies, and underlying assumptions. You will verify calculation methods, identify shortcomings, and provide documented analysis that ensures actuarial rigor meets professional standards.

Talent poolRecently verified$40-125/hr
  • Software Engineering
  • Platform: DataAnnotation

As an Application Security Engineer, you will evaluate how AI models reason about secure design and identify security vulnerabilities they miss, including architectural flaws, unsafe deserialization, and overly permissive access controls. You will threat-model systems, conduct security code reviews, and write clear guidance on secure design and vulnerability fixes.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Arabic

This role involves evaluating and improving Arabic translations and AI-generated text by assessing fluency, tone, and cultural authenticity. Native or near-native Arabic speakers will teach language models to produce more natural, human-like Arabic output by rating translations, identifying gaps that automated systems miss, and rewriting model outputs to match native speaker standards.

Talent poolRecently verified$25-40/hr
  • Legal
  • Platform: DataAnnotation

This role involves evaluating and improving AI legal reasoning by grading model outputs against professional standards, identifying errors like invented citations and misapplied precedent, and writing correct legal analysis when needed. It is suited for attorneys and legal professionals who can assess AI-generated legal content and provide high-quality written corrections.

Talent poolRecently verified$40-125/hr
  • Legal
  • Platform: DataAnnotation

This role evaluates how AI models handle audit procedures, internal controls, and assurance judgments by testing their reasoning, identifying gaps in professional skepticism, and writing correct audit approaches. It is for experienced auditors who can assess whether AI-generated evidence actually supports audit opinions.

Talent poolRecently verified$40-125/hr
  • Engineering
  • Platform: DataAnnotation
  • Level: Intermediate

This role involves creating challenging engineering problems and solutions that test AI model capabilities in mechanical reasoning and design. You'll write problems from your specialty, solve them yourself, verify the answers, and evaluate how well AI systems perform on real engineering tasks like component sizing, failure analysis, and design validation.

Talent poolRecently verified$40-125/hr
  • Software Engineering
  • Platform: DataAnnotation

Backend engineers will evaluate how AI models reason about system design, data modeling, and API contracts, identifying failures like race conditions and N+1 queries that only surface under real load. The role involves stress-testing model-generated backend code and writing correct solutions when models fall short, with explicit attention to scaling and consistency trade-offs.

Talent poolRecently verified$40-150/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Bangla

A Bengali specialist evaluates and corrects AI-generated Bengali text, focusing on accurate pronoun usage, regional variations, and natural phrasing. The role involves reviewing translations, scoring model outputs, and providing native-speaker feedback to improve AI language understanding.

Talent poolRecently verified$25-40/hr
  • Science & Research
  • Platform: DataAnnotation

This role involves evaluating AI models' biological reasoning against established scientific literature. You will write test prompts to probe biological understanding, review AI-generated responses for errors and fabricated citations, and provide scientifically accurate answers based on primary sources when the model's output is inadequate.

Talent poolRecently verified$40-125/hr
  • Health & Medicine
  • Platform: DataAnnotation

This role involves evaluating how well AI models reason through cardiology cases, identifying errors in clinical judgment that could harm patients. Cardiologists will write prompts to test AI reasoning, review AI outputs for accuracy against current guidelines, and provide correct answers when the model errs, serving as the training signal for AI improvement.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Catalan

Evaluate AI model performance on Catalan translation and localization tasks, assessing translation quality and cultural appropriateness. The role involves reviewing AI-generated outputs, identifying errors, and providing corrections and explanations to improve model training.

Talent poolRecently verified$25-40/hr
  • Science & Research
  • Platform: DataAnnotation

This role involves evaluating AI-generated chemical reasoning and responses to identify flawed mechanisms, unsafe procedures, and analytical errors. You will write prompts to test chemical knowledge, review AI outputs for errors, and provide correct expert-level answers with proper attention to safety and experimental conditions.

Talent poolRecently verified$40-125/hr
  • Languages, Translation & Voice
  • Platform: DataAnnotation
  • Language: Chinese

A Chinese language specialist who evaluates and corrects AI-generated Mandarin text for naturalness, cultural appropriateness, and native fluency. The role involves rating model outputs, identifying errors in measure words and idioms, and rewriting content to sound authentically Chinese rather than translated.

Talent poolRecently verified$25-40/hr
  • Health & Medicine
  • Platform: DataAnnotation

A Clinical Documentation Specialist evaluates AI-generated medical documentation for integrity, compliance, and adherence to coding and privacy standards. The role requires identifying documentation gaps, missed query opportunities, and HIPAA violations, then providing corrections based on CDI best practices.

Talent poolRecently verified$40-125/hr
  • Legal
  • Platform: DataAnnotation

A compliance officer evaluates whether AI models correctly analyze regulatory obligations, identify applicable rules, and provide accurate guidance on policy and reporting requirements. The role involves testing AI responses against current regulatory texts across multiple jurisdictions and writing correct analysis when models make errors.

Talent poolRecently verified$40-125/hr
  • Data & Machine Learning
  • Platform: DataAnnotation
  • Level: Intermediate

This role involves evaluating how well AI models can perform computational biology tasks. Computational biologists and bioinformaticians will design realistic analysis scenarios from their own work, run them through AI systems, and grade the results against professional standards to help benchmark model capabilities.

Talent poolRecently verified$40-125/hr
  • Legal
  • Platform: DataAnnotation

This role involves evaluating AI-generated contract drafting and redlining to identify logical errors, ambiguities, and one-sided terms that automated systems miss. Contract professionals will draft test prompts, review model outputs for issues like inconsistent definitions and broken cross-references, and rewrite clauses to ensure precise legal language.

Talent poolRecently verified$40-125/hr
  • Finance & Accounting
  • Platform: DataAnnotation

A corporate accountant evaluates AI model outputs on financial accounting tasks, checking for errors in GAAP treatment, consolidations, and month-end closing procedures. You will draft test prompts, review AI-generated entries for mistakes, and provide correct treatments with clear working papers for training purposes.

Talent poolRecently verified$40-125/hr
  • Legal
  • Platform: DataAnnotation

This role involves evaluating AI-generated legal work on mergers, acquisitions, financings, and corporate governance to identify logical gaps, unworkable terms, and misaligned risk allocation. You will draft test prompts for AI models handling complex transactional agreements and rewrite provisions where the AI output falls short, applying the discipline of carefully negotiated deals.

Talent poolRecently verified$40-125/hr

FAQ

Questions about AI training jobs

What kinds of jobs are listed here?

AI training, model evaluation, RLHF, red teaming, data labeling and expert review roles, in fields from software and data science to medicine, law, finance, languages and science.

Are these jobs remote?

They are done online. Some are limited to residents of certain countries: when the platform publishes that restriction, the listing shows it.

Do I need prior experience in AI?

Usually not. Most roles ask for professional experience in your own field. Each listing states the experience the platform requires.

How is the work paid?

By the hour or by the task, by the platform that hires you. We show the pay only when the platform publishes it. It is not a guarantee of income or hours.

How do I apply?

Apply now opens the official posting on Mercor, micro1, Turing, Ethos, Mindrift, DataAnnotation, Alignerr, Handshake AI, Surge AI, xAI, Vetto or Terac, where you complete the application. You need no account here. See how it works.