AI Safety AI training jobs

Open AI training and model evaluation roles that call for AI Safety expertise, gathered from 3 platforms. Part of Data & Machine Learning.

56
open jobs
$50/hr
median published rate
$120/hr
highest published rate
3
platforms hiring

Pay figures cover the 53 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

41 to 56 of 56 jobs
  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Vietnamese

You will join a team of adversarial testers tasked with probing conversational AI models to identify security vulnerabilities, biases, and misinformation. The job involves generating high-quality training data by simulating attacks, classifying failures, and documenting systemic risks. This role is for experts who can think like an adversary, structure their testing according to established frameworks, and communicate results clearly.

Posted July 30, 2026

OpenRecently verified$17-25/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: 40 countries
  • Level: Expert

This role involves judging how frontier AI models respond to sensitive and ambiguous topics, checking answers for safety, accuracy and policy compliance. It suits experienced professionals from fields such as psychology, law, science or public policy who can apply safety rubrics and give structured feedback to model developers.

Posted July 16, 2026

OpenRecently verified$60-70/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: 40 countries
  • Level: Expert

This role involves stress-testing frontier AI models with adversarial prompts to uncover jailbreaks, unsafe behaviour and policy failures on sensitive topics. It is aimed at experienced safety, security, life sciences or policy professionals who can document weaknesses and help researchers strengthen model alignment.

Posted July 16, 2026

OpenRecently verified$70-84/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Assamese

You join a red teaming team to test the robustness and security of conversational AI models. Your mission is to generate adversarial inputs (jailbreaks, prompt injections, misuse cases) to uncover vulnerabilities in AI systems before their deployment. You annotate failures, classify risks, and produce documented datasets that clients can leverage to improve their models.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Bangla

This role involves testing conversational AI models as an AI safety expert by generating adversarial training data to identify vulnerabilities. You will produce reports and structured datasets that enable clients to improve the robustness and security of their AI systems. The ideal profile masters critical content analysis, methodological rigor, and technical communication.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Gujarati

You join a red teaming team to test and secure conversational AI models. Your job consists of generating adversarial inputs, identifying vulnerabilities (biases, disinformation, harmful behaviors) and producing high-quality annotated data that makes AI safer. This role requires fluent mastery of English and Gujarati as well as refined judgment on language and content.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Kannada

This role involves evaluating the safety of conversational AI models by testing them with adversarial inputs and sophisticated attack techniques. You will generate high-quality annotation data, document discovered vulnerabilities, and contribute to strengthening the robustness of AI systems. This job is suited for someone capable of carefully analyzing AI responses, detecting biases and subtle flaws, and communicating findings in a structured manner.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Malayalam

You join a specialized AI safety team to test conversational language models by exposing their flaws and vulnerabilities through adversarial attacks. The job consists of generating high-quality training data through error annotation, risk classification, and documentation of reproducible attacks. This role is for rigorous language experts with critical judgment about the quality and accuracy of AI responses.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Odia

Join a red teaming team dedicated to AI safety: you will probe conversational models with adversarial inputs to discover vulnerabilities and generate robust test data. This role requires native mastery of English and Odia, as well as fine judgment to assess the relevance, accuracy, and appropriateness of AI responses when facing sensitive topics. You will document your findings in a reproducible manner so that clients can strengthen the security of their systems.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Punjabi

Join a red teaming team specialized in identifying vulnerabilities of conversational AI models. You will generate high-quality training data by testing AI systems with adversarial inputs, documenting flaws, and producing reproducible reports that clients can use to strengthen the security of their models.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Tamil

This role involves testing and evaluating the robustness of conversational AI models by subjecting them to adversarial attacks, bypasses, and malicious use cases. You will generate high-quality annotation data to identify vulnerabilities, biases, and systemic risks, following established taxonomies and benchmarks. This role is suited for AI safety experts with advanced proficiency in both English and Tamil.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Telugu

You join an adversarial testing team tasked with detecting vulnerabilities in conversational AI models by generating high-quality human-generated data. Your job consists of exploring security flaws, biases and harmful behaviors through malicious inputs, then documenting systemic risks so that clients can correct them. You must be bilingual in English and Telugu and demonstrate rigor in analyzing AI responses on sensitive topics.

Posted June 4, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Marathi

This role involves testing conversational AI models as a safety expert by generating adversarial attacks, security bypasses, and misuse scenarios to identify vulnerabilities. You will produce high-quality annotation data and reproducible reports that help clients strengthen the robustness of their AI systems. The ideal profile must possess sound judgment about language and content, be rigorous in identifying subtle anomalies, and capable of clearly communicating observations.

Posted June 3, 2026

OpenRecently verified$16-22/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Vetto

This remote project asks contributors to role-play users holding assigned political views in multi-turn conversations with an AI system, then judge whether the AI stays neutral, balanced and responsible under pushback. It suits people who already take part in political debate in everyday life and can keep a persona separate from their own opinions while applying detailed written guidelines.

Posted May 22, 2026

OpenRecently verifiedPay not disclosed
  • Writing, Creative & Design
  • Platform: Vetto

This freelance job asks contributors to perform emotionally complex characters in structured multi-turn role-play conversations with an AI system. Clinical experts later use these performances to assess how AI handles sensitive topics such as emotional distress and suicidal ideation. It suits performers with acting or improvisation training, as well as people with crisis line or peer support experience.

Posted May 7, 2026

OpenRecently verifiedPay not disclosed
  • Health & Medicine
  • Platform: Vetto

This project-based role involves reviewing pre-generated health scenarios used to train AI health assistants. Reviewers rate how realistic each scenario is and how much safety risk it carries, then strengthen weak or missing constraints. It is aimed at generalist doctors and experienced nurses with strong patient-facing experience who can tell authentic care-seeking situations from generic ones.

Posted February 5, 2026

OpenRecently verifiedPay not disclosed

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Safety AI training jobs

How many AI Safety AI training jobs are open right now?

56 AI Safety AI training jobs are open on SideHustler today. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Safety AI training jobs pay?

Among the 53 open roles that publish an hourly rate in USD, pay runs from $15 to $120 per hour, with a median of $50. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Safety AI training jobs?

Mercor (37), Alignerr (16) and Vetto (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Safety AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Safety, Red Teaming, Adversarial Testing, AI Evaluation, Vulnerability Assessment. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.