OpenRecently verified

AI Safety Red Teamer

  • Data & Machine Learning
  • Platform: Mercor
  • Location: 40 countries
  • Level: Expert

$70-84/hrAs published by the platform.

About this job

This role involves stress-testing frontier AI models with adversarial prompts to uncover jailbreaks, unsafe behaviour and policy failures on sensitive topics. It is aimed at experienced safety, security, life sciences or policy professionals who can document weaknesses and help researchers strengthen model alignment.

What you'll do

  • Design adversarial prompts to stress-test frontier AI models.
  • Identify jailbreaks, unsafe behaviours, hallucinations and policy failures.
  • Evaluate model robustness across sensitive domains such as cyber, biosecurity and fraud.
  • Document vulnerabilities and contribute to safety benchmarks and red-teaming reports.

Requirements

AI SafetyAI Red TeamingAdversarial prompt designFrontier AI evaluationTrust & SafetyCybersecurityLife sciencesInvestigative journalism
  • 5+ years of experience
  • Bachelor's degree or higher in Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related discipline

Pay

$70-84/hr

Pay as published by the platform. It is not a guarantee of income or hours.

Location

Open to residents of: United States, Denmark, Estonia, Finland, Iceland, Ireland, Latvia, Lithuania, Norway, Sweden, Austria, Belgium, France, Germany, Liechtenstein, Luxembourg, Monaco, Netherlands, Switzerland, United Kingdom, Albania, Bosnia & Herzegovina, Croatia, Greece, Italy, Kosovo, Malta, North Macedonia, Portugal, San Marino, Serbia, Slovenia, Spain, Bulgaria, Czechia, Hungary, Moldova, Poland, Romania, Slovakia.

How to apply

You apply on work.mercor.com: create a profile, complete a skills assessment, then get matched to projects that fit your background.

All Mercor jobs and how the platform works

Source

Official posting: https://work.mercor.com/jobs/list_AAABn2uCTs39FcmSt0NAnYGQ/ai-safety-red-teamer

Last checked on October 8, 2026.

Similar jobs

  • Data & Machine Learning
  • Platform: Alignerr

This contract role asks security-minded people to probe AI systems built on the OpenClaw platform by running red-teaming exercises and writing adversarial prompts. The work suits cybersecurity practitioners who enjoy finding weaknesses and who can document their findings clearly for technical and non-technical readers.

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Danish

You join a red teaming team specialized in adversarial evaluation of conversational AI models. Your role consists of testing AI systems by exploring their vulnerabilities (jailbreaks, prompt injections, biases), generating high-quality data documenting these vulnerabilities, and producing reproducible reports to strengthen model safety. This position is for bilingual English-Danish experts with prior experience in red teaming or related fields (cybersecurity, adversarial ML, socio-technical analysis).

Posted July 30, 2026

OpenRecently verified$48-62/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Dutch

This role consists of testing the robustness and security of conversational AI models by subjecting them to sophisticated adversarial attacks. You will generate high-quality training data by identifying vulnerabilities, biases, and systemic risks that automated tests do not detect. The ideal profile has prior experience in red teaming, structured adversarial thinking, and the ability to clearly communicate security risks.

Posted July 30, 2026

OpenRecently verified$48-62/hr

Referral link: we may earn a fee. Apply without it

  • Software Engineering
  • Platform: Alignerr

This role involves analyzing security vulnerabilities and threat scenarios in AI systems and large language models to identify weaknesses and recommend mitigations. Security professionals will conduct adversarial testing, evaluate real-world attack scenarios, and provide structured feedback to improve AI safety and resilience.

  • Software Engineering
  • Platform: DataAnnotation

This role involves evaluating how AI models reason about offensive security concepts and identifying flaws in their exploit chain reasoning. Penetration testers with hands-on red team experience will test model outputs across reconnaissance, exploitation, privilege escalation, and lateral movement, then write accurate attack paths that reflect real-world tradecraft when models fall short.

Talent poolRecently verified$40-125/hr
  • Data & Machine Learning
  • Platform: Vetto

This remote project asks contributors to role-play users holding assigned political views in multi-turn conversations with an AI system, then judge whether the AI stays neutral, balanced and responsible under pushback. It suits people who already take part in political debate in everyday life and can keep a persona separate from their own opinions while applying detailed written guidelines.

Posted May 22, 2026

OpenRecently verifiedPay not disclosed

Keep exploring

Apply on Mercor (opens Mercor in a new tab)

This is a referral link: we may earn a fee if you are selected and meet the platform's conditions.

Apply without the referral link