AI Safety AI training jobs

Open AI training and model evaluation roles that call for AI Safety expertise, gathered from 3 platforms. Part of Data & Machine Learning.

56
open jobs
$50/hr
median published rate
$120/hr
highest published rate
3
platforms hiring

Pay figures cover the 53 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

21 to 40 of 56 jobs
  • Engineering
  • Platform: Mercor
  • Location: United States

You will contribute to evaluating frontier AI models as a nuclear fuel cycle expert. Your role will consist of writing nuanced technical questions and judging whether the model's answers comply with a safety policy, distinguishing between legitimate requests and potentially dangerous queries. This work requires a deep understanding of the parameters that differentiate a civilian application from a diversion for proliferation purposes.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Health & Medicine
  • Platform: Mercor
  • Location: United States

This job involves evaluating the capacity of AI models to distinguish legitimate requests from malicious uses in the field of nuclear medicine and medical isotopes. You will write nuanced test prompts and judge whether the model's responses comply with a security policy, while explaining your decisions in writing. This role requires specialized expertise to draw the line between routine professional questions and those concealing dangerous intentions.

Posted September 14, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Science & Research
  • Platform: Mercor
  • Location: United States

This role involves evaluating an AI model's responses to requests involving energetic materials and explosives, distinguishing between legitimate questions and dangerous requests. You will write progressively more complex test prompts, assess the model's responses against a defined safety policy, and provide written justification for your technical judgments. This role is intended for experts in materials chemistry specializing in propulsion and initiation systems, with solid practical experience in the field.

Posted September 9, 2026

OpenRecently verified$65-75/hr

Referral link: we may earn a fee. Apply without it

  • Business, Consulting & Operations
  • Platform: Mercor
  • Location: United States
  • Language: Croatian

This role involves evaluating and improving the safety of AI models by analyzing their behavior on sensitive topics in Croatian. You will write expert prompts, classify content according to structured guidelines, and identify adversarial formulations, as an English-Croatian bilingual speaker with no prior AI experience required.

Posted September 4, 2026

OpenRecently verified$38-42/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: Belgium
  • Language: Dutch

This role asks bilingual English and Dutch speakers to test how AI models respond to sensitive topics in Dutch as used in Belgium. The work combines language fluency with cultural judgment, and no prior AI experience is needed because the workflow is taught on the job.

Posted September 4, 2026

OpenRecently verified$48-52/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: United States
  • Language: Finnish

This role involves evaluating and improving the safety of advanced AI models in Finnish as a bilingual English-Finnish speaker. You will apply your language mastery and judgment to examine how these models handle sensitive topics, without requiring prior experience in AI or machine learning.

Posted September 4, 2026

OpenRecently verified$48-52/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: United States
  • Language: French

This role consists of evaluating the safety of AI models in French by examining how they handle sensitive topics. You will write expert prompts, apply classification guidelines, and identify adversarial formulations, leveraging your bilingual mastery of French and English.

Posted September 4, 2026

OpenRecently verified$48-52/hr

Referral link: we may earn a fee. Apply without it

  • Science & Research
  • Platform: Mercor
  • Location: United States
  • Language: Indonesian

This remote-style role asks PhD scientists who are fluent in both English and Indonesian to help make AI models safer on technical subjects. Contributors write expert prompts in Indonesian, then judge model answers for accuracy, helpfulness, and handling of sensitive material. No prior AI experience is needed because the workflow is taught on the job.

Posted September 4, 2026

OpenRecently verified$19-23/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Norwegian

This role consists of evaluating and strengthening the safety of advanced AI models by analyzing their behavior on sensitive topics in Norwegian. You will use your bilingual proficiency (English-Norwegian) and cultural judgment to identify gaps and bypass attempts, with no prior AI experience required.

Posted September 4, 2026

OpenRecently verified$58-62/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Turkish

You will participate in improving AI model safety by evaluating how they handle sensitive subjects in Turkish. Your linguistic and cultural judgments will help identify and strengthen weaknesses in these systems when facing delicate content. No prior AI experience is required.

Posted September 4, 2026

OpenRecently verified$23-27/hr

Referral link: we may earn a fee. Apply without it

  • Generalist & Data Labeling
  • Platform: Mercor
  • Location: United States
  • Language: Ukrainian

This role involves contributing to AI model safety as a bilingual English-Ukrainian expert. You will evaluate how these models handle sensitive topics in Ukrainian, by writing expert prompts and classifying conversations according to structured guidelines. No prior experience in AI or machine learning is required.

Posted September 4, 2026

OpenRecently verified$38-42/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Danish

You join a red teaming team specialized in adversarial evaluation of conversational AI models. Your role consists of testing AI systems by exploring their vulnerabilities (jailbreaks, prompt injections, biases), generating high-quality data documenting these vulnerabilities, and producing reproducible reports to strengthen model safety. This position is for bilingual English-Danish experts with prior experience in red teaming or related fields (cybersecurity, adversarial ML, socio-technical analysis).

Posted July 30, 2026

OpenRecently verified$48-62/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Dutch

This role consists of testing the robustness and security of conversational AI models by subjecting them to sophisticated adversarial attacks. You will generate high-quality training data by identifying vulnerabilities, biases, and systemic risks that automated tests do not detect. The ideal profile has prior experience in red teaming, structured adversarial thinking, and the ability to clearly communicate security risks.

Posted July 30, 2026

OpenRecently verified$48-62/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Finnish

This role involves testing the limits and vulnerabilities of conversational AI models using adversarial techniques to identify security risks. You generate high-quality data that enables clients to improve the robustness and reliability of their AI systems. The ideal profile masters red teaming, structured adversarial thinking, and the ability to reproducibly document discovered flaws.

Posted July 30, 2026

OpenRecently verified$48-62/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Indonesian

You will join an adversarial testing team to evaluate the robustness and safety of conversational AI models. Your job is to identify vulnerabilities by generating malicious inputs, annotating failures, and documenting systemic risks to enable clients to improve their AI systems. This role suits experts with experience in security testing and a natural ability to explore the limits of systems.

Posted July 30, 2026

OpenRecently verified$17-25/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Malay

This role involves training conversational AI models by testing them adversarially to identify vulnerabilities and improve their safety. You will generate high-quality training data by annotating failures, classifying risks, and documenting reproducible attack cases. The ideal profile has prior experience in red teaming or cybersecurity, and masters a structured and methodical approach to testing the limits of AI systems.

Posted July 30, 2026

OpenRecently verified$17-25/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Norwegian Bokmål

You join a team dedicated to adversarial testing of conversational AI models, identifying their vulnerabilities to malicious inputs, manipulations, and biases. This role consists of generating high-quality training data to strengthen the robustness and security of AI systems. You are ideally a red teaming expert with prior experience in adversarial security testing and fluent proficiency in English and Norwegian.

Posted July 30, 2026

OpenRecently verified$48-62/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Portuguese

Join a red teaming team specialized in AI safety: you will probe conversational models with adversarial attacks, discover vulnerabilities, and generate data that makes AI systems safer. This role, entirely text-based and team-oriented, will allow you to deploy ethical hacking techniques, risk classification, and documentation so that clients can act on identified flaws.

Posted July 30, 2026

OpenRecently verified$29-45/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Swedish

You join a specialized team focused on adversarial evaluation of conversational AI models. Your role is to identify vulnerabilities in these systems by testing them with malicious inputs, then generate the data and reports that will enable clients to strengthen the security of their products. This position requires fluent mastery of English and Swedish, as well as prior experience in adversarial testing or security.

Posted July 30, 2026

OpenRecently verified$48-62/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Thai

You will join a red teaming team to test the robustness and security of conversational AI models by simulating adversarial attacks. The role involves generating high-quality training data by identifying vulnerabilities, biases, and abuse cases through structured and documented techniques. This position is for experts capable of thinking like an attacker to strengthen the defense of AI systems.

Posted July 30, 2026

OpenRecently verified$24-35/hr

Referral link: we may earn a fee. Apply without it

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Safety AI training jobs

How many AI Safety AI training jobs are open right now?

56 AI Safety AI training jobs are open on SideHustler today. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Safety AI training jobs pay?

Among the 53 open roles that publish an hourly rate in USD, pay runs from $15 to $120 per hour, with a median of $50. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Safety AI training jobs?

Mercor (37), Alignerr (16) and Vetto (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Safety AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Safety, Red Teaming, Adversarial Testing, AI Evaluation, Vulnerability Assessment. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.