1234 open jobs

Remote AI training jobs

Open roles in AI training, model evaluation, RLHF and data labeling for domain experts. Search, filter, then apply directly on Mercor, micro1, Turing, Ethos, Mindrift, DataAnnotation, Alignerr, Handshake AI, Surge AI, xAI, Vetto or Terac.

"prompt"×Clear all
21 to 40 of 68 jobs
  • Business, Consulting & Operations
  • Platform: DataAnnotation

This role involves evaluating AI-generated responses about healthcare administration and operations, including billing, coding, insurance workflows, and compliance. You will write test prompts, review AI outputs for accuracy and feasibility, and provide correct guidance based on practical operational experience.

Talent poolRecently verified$40-125/hr
  • Legal
  • Platform: DataAnnotation

As a legal expert, you will evaluate AI-generated legal analysis for accuracy, identifying errors such as invented citations, misapplied precedent, and jurisdictional confusion. You will draft test prompts covering contracts, statutes, procedure, and case analysis, then write corrected legal analysis when the model's reasoning falls short.

Talent poolRecently verified$40-125/hr
  • Legal
  • Platform: DataAnnotation

Evaluate AI-generated legal work on litigation matters including pleadings, discovery, and motions by checking citations, holdings, procedural rules, and standards of review. Create test prompts for litigation scenarios and rewrite AI outputs to meet court-quality standards when necessary.

Talent poolRecently verified$40-125/hr
  • Data & Machine Learning
  • Platform: DataAnnotation

As a Machine Learning Engineer, you will evaluate and improve how AI models reason about machine learning systems, including training dynamics, evaluation design, and deployment strategies. You will write prompts to test model reasoning, identify subtle errors in AI outputs, and provide correct solutions based on real practitioner experience.

Talent poolRecently verified$40-150/hr
  • Health & Medicine
  • Platform: DataAnnotation

Medical coders evaluate AI-generated medical coding outputs against real clinical documentation and coding guidelines, identifying errors in ICD-10-CM, CPT, HCPCS codes, modifiers, and E&M levels. The role requires writing test prompts to assess coding accuracy and providing corrected assignments with clear documentation-based explanations to serve as training signals for AI models.

Talent poolRecently verified$40-125/hr
  • Legal
  • Platform: DataAnnotation

This role involves evaluating AI systems' performance on legal research, document review, and litigation filings by checking for errors like invented citations and overruled authorities. As a paralegal, you'll draft test prompts, review AI outputs for accuracy, and produce correct work product when the model falls short.

Talent poolRecently verified$40-125/hr
  • Health & Medicine
  • Platform: DataAnnotation

A physician evaluates AI model responses on clinical questions against current standards of care, identifying unsafe or outdated guidance and writing correct clinical answers when needed. This role requires assessing AI clinical reasoning and creating prompts that test differential diagnosis, treatment planning, and patient communication.

Talent poolRecently verified$40-125/hr
  • Finance & Accounting
  • Platform: DataAnnotation

This role involves stress-testing AI models' financial reasoning by writing prompts that probe quantitative analysis and evaluating their outputs for errors in pricing, risk assessment, portfolio construction, and statistical modeling. You will review AI-generated financial analysis for flaws in derivations and model application, and write correct analyses to serve as training data.

Talent poolRecently verified$40-125/hr
  • Health & Medicine
  • Platform: DataAnnotation

This role involves evaluating AI models' clinical reasoning on nursing topics including assessment, medication administration, and patient care protocols. Registered nurses will identify gaps between textbook knowledge and real bedside practice, write test prompts to probe nursing judgment, and provide corrected guidance when AI outputs contain unsafe or incorrect clinical information.

Talent poolRecently verified$40-125/hr
  • Business, Consulting & Operations
  • Platform: DataAnnotation

This role involves evaluating AI-generated answers on healthcare revenue cycle topics including claims, denials, and payer operations. You will write test prompts to probe the AI's understanding of real-world billing mechanics and provide expert corrections and guidance based on hands-on revenue cycle experience.

Talent poolRecently verified$40-125/hr
  • Writing, Creative & Design
  • Platform: Mercor
  • Location: United States

You will compare images generated by artificial intelligence based on the same prompt and designate the most successful image. This qualitative evaluation role requires careful analysis according to several criteria (faithfulness to the prompt, visual quality, anatomy, rendered text, artifacts) and the writing of a concise justification for your choice.

Posted October 3, 2026

OpenRecently verified$30/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: micro1
  • Location: 58 countries

You are invited to contribute as a senior AI trainer to a client project aimed at improving the quality of AI systems. You will evaluate AI-based chat and search tools, refine prompts to test model capabilities, and provide structured feedback to optimize their performance. This role is suited for professionals with proven experience in AI model training and data annotation.

Posted September 23, 2026

OpenRecently verified$14-36/hr

Referral link to micro1's job list: search for this role there. Open this exact job

  • Business, Consulting & Operations
  • Platform: micro1
  • Location: 58 countries
  • Level: Intermediate

This job involves designing educational resources and evaluation criteria to train AI systems to master writing tasks specific to the U.S. Department of Defense. As a DoD expert, you will create high-quality prompts, reference responses, and scoring rubrics based on this department's actual standards, notably for activity reports, tactical situation reports, and executive summaries. No prior AI experience is required; your domain knowledge and mastery of federal administrative writing are the essential assets.

Posted September 21, 2026

OpenRecently verified$40-80/hr

Referral link to micro1's job list: search for this role there. Open this exact job

  • Science & Research
  • Platform: Mercor
  • Location: United States

You will write test prompts to evaluate whether AI models can distinguish legitimate chemical requests from potentially dangerous misuses. As an analytical chemist, you will develop scenarios along the dual-use line, assess model responses against a defined policy, and document your analyses with accessible technical reasoning.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Health & Medicine
  • Platform: Mercor
  • Location: United States

This job involves evaluating the capacity of AI models to distinguish legitimate requests from malicious uses in the field of nuclear medicine and medical isotopes. You will write nuanced test prompts and judge whether the model's responses comply with a security policy, while explaining your decisions in writing. This role requires specialized expertise to draw the line between routine professional questions and those concealing dangerous intentions.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Science & Research
  • Platform: Mercor
  • Location: United States

You will participate in evaluating cutting-edge AI models as a nuclear security specialist. You will write calibrated test prompts across three risk levels to judge whether the model correctly distinguishes legitimate professional questions from dangerous requests, then you will evaluate and document its responses according to a defined policy standard.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Science & Research
  • Platform: Mercor
  • Location: United States

Mercor is looking for nuclear and safeguards experts to test the robustness of AI models against dual-use requests. You will write targeted prompts on the knife's edge between legitimate professional questions and dangerous requests, then evaluate the model's responses against a defined policy and write reference answers. This role is for practitioners with concrete experience in nuclear material verification and detecting diversion attempts.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Health & Medicine
  • Platform: Mercor
  • Location: United States

You will contribute to evaluating frontier AI models as a source security specialist. Your role is to write prompts that test the model's ability to distinguish legitimate professional questions from potentially dangerous requests, then evaluate its responses against a defined policy standard. This work requires practical expertise in Category 1 and 2 radioactive source security, as well as strong technical writing skills to justify your assessments.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Science & Research
  • Platform: Mercor
  • Location: United States

You are a chemistry expert with practical experience in synthesis, analytical chemistry, or chemical safety. You participate in evaluating AI models by writing calibrated test prompts across three risk levels, then assessing model responses against an established safety framework. This role requires both advanced technical expertise and the ability to document your judgments in an accessible way.

Posted September 9, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Science & Research
  • Platform: Mercor
  • Location: United States

This role involves evaluating an AI model's responses to requests involving energetic materials and explosives, distinguishing between legitimate questions and dangerous requests. You will write progressively more complex test prompts, assess the model's responses against a defined safety policy, and provide written justification for your technical judgments. This role is intended for experts in materials chemistry specializing in propulsion and initiation systems, with solid practical experience in the field.

Posted September 9, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

FAQ

Questions about AI training jobs

What kinds of jobs are listed here?

AI training, model evaluation, RLHF, red teaming, data labeling and expert review roles, in fields from software and data science to medicine, law, finance, languages and science.

Are these jobs remote?

They are done online. Some are limited to residents of certain countries: when the platform publishes that restriction, the listing shows it.

Do I need prior experience in AI?

Usually not. Most roles ask for professional experience in your own field. Each listing states the experience the platform requires.

How is the work paid?

By the hour or by the task, by the platform that hires you. We show the pay only when the platform publishes it. It is not a guarantee of income or hours.

How do I apply?

Apply now opens the official posting on Mercor, micro1, Turing, Ethos, Mindrift, DataAnnotation, Alignerr, Handshake AI, Surge AI, xAI, Vetto or Terac, where you complete the application. You need no account here. See how it works.