AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

679
open jobs
26
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 637 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

601 to 620 of 679 jobs
  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Danish

You join a red teaming team specialized in adversarial evaluation of conversational AI models. Your role consists of testing AI systems by exploring their vulnerabilities (jailbreaks, prompt injections, biases), generating high-quality data documenting these vulnerabilities, and producing reproducible reports to strengthen model safety. This position is for bilingual English-Danish experts with prior experience in red teaming or related fields (cybersecurity, adversarial ML, socio-technical analysis).

Posted July 30, 2026

OpenRecently verified$48-62/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Dutch

This role consists of testing the robustness and security of conversational AI models by subjecting them to sophisticated adversarial attacks. You will generate high-quality training data by identifying vulnerabilities, biases, and systemic risks that automated tests do not detect. The ideal profile has prior experience in red teaming, structured adversarial thinking, and the ability to clearly communicate security risks.

Posted July 30, 2026

OpenRecently verified$48-62/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Finnish

This role involves testing the limits and vulnerabilities of conversational AI models using adversarial techniques to identify security risks. You generate high-quality data that enables clients to improve the robustness and reliability of their AI systems. The ideal profile masters red teaming, structured adversarial thinking, and the ability to reproducibly document discovered flaws.

Posted July 30, 2026

OpenRecently verified$48-62/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Indonesian

You will join an adversarial testing team to evaluate the robustness and safety of conversational AI models. Your job is to identify vulnerabilities by generating malicious inputs, annotating failures, and documenting systemic risks to enable clients to improve their AI systems. This role suits experts with experience in security testing and a natural ability to explore the limits of systems.

Posted July 30, 2026

OpenRecently verified$17-25/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Malay

This role involves training conversational AI models by testing them adversarially to identify vulnerabilities and improve their safety. You will generate high-quality training data by annotating failures, classifying risks, and documenting reproducible attack cases. The ideal profile has prior experience in red teaming or cybersecurity, and masters a structured and methodical approach to testing the limits of AI systems.

Posted July 30, 2026

OpenRecently verified$17-25/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Norwegian Bokmål

You join a team dedicated to adversarial testing of conversational AI models, identifying their vulnerabilities to malicious inputs, manipulations, and biases. This role consists of generating high-quality training data to strengthen the robustness and security of AI systems. You are ideally a red teaming expert with prior experience in adversarial security testing and fluent proficiency in English and Norwegian.

Posted July 30, 2026

OpenRecently verified$48-62/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Portuguese

Join a red teaming team specialized in AI safety: you will probe conversational models with adversarial attacks, discover vulnerabilities, and generate data that makes AI systems safer. This role, entirely text-based and team-oriented, will allow you to deploy ethical hacking techniques, risk classification, and documentation so that clients can act on identified flaws.

Posted July 30, 2026

OpenRecently verified$29-45/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Swedish

You join a specialized team focused on adversarial evaluation of conversational AI models. Your role is to identify vulnerabilities in these systems by testing them with malicious inputs, then generate the data and reports that will enable clients to strengthen the security of their products. This position requires fluent mastery of English and Swedish, as well as prior experience in adversarial testing or security.

Posted July 30, 2026

OpenRecently verified$48-62/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Thai

You will join a red teaming team to test the robustness and security of conversational AI models by simulating adversarial attacks. The role involves generating high-quality training data by identifying vulnerabilities, biases, and abuse cases through structured and documented techniques. This position is for experts capable of thinking like an attacker to strengthen the defense of AI systems.

Posted July 30, 2026

OpenRecently verified$24-35/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Language: Vietnamese

You will join a team of adversarial testers tasked with probing conversational AI models to identify security vulnerabilities, biases, and misinformation. The job involves generating high-quality training data by simulating attacks, classifying failures, and documenting systemic risks. This role is for experts who can think like an adversary, structure their testing according to established frameworks, and communicate results clearly.

Posted July 30, 2026

OpenRecently verified$17-25/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Level: Entry level

Mercor is recruiting experienced data scientists to join a talent network designed to evaluate the capabilities of AI systems to perform data science work. Future roles will primarily involve designing precise evaluation criteria and assessing data science deliverables according to established standards. This role is suited for professionals who have worked in top-tier organizations and master analytical best practices.

Posted July 29, 2026

Talent poolRecently verified$100-150/hr

Referral link: we may earn a fee. Apply without it

  • Business, Consulting & Operations
  • Platform: Mercor
  • Location: United States
  • Level: Entry level

Mercor is building a network of experienced management consultants who may work on future projects with AI research organizations. Your roles will consist of evaluating the quality of consulting work generated by AI systems, by designing precise evaluation criteria and providing detailed justifications for your assessments. This position is for professionals with solid experience in strategic consulting firms.

Posted July 28, 2026

Talent poolRecently verified$100-150/hr

Referral link: we may earn a fee. Apply without it

  • Business, Consulting & Operations
  • Platform: Mercor
  • Location: United States

Join a team tasked with driving innovative projects in the field of navigation capabilities for language models. You will manage performance tracking, quality assurance of annotations, and relationships with contributors, taking your projects from end to end with autonomy and rigor.

Posted July 27, 2026

OpenRecently verified$50-60/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Level: Intermediate

Mercor is recruiting research experts capable of improving and evaluating cutting-edge AI models. You will contribute to creating high-quality training materials (questions, answers, evaluations) for advanced language models. This role is aimed at candidates holding a Master's degree or with 2 to 3 years of experience, with strong research and communication skills.

Posted July 26, 2026

OpenRecently verified$50-60/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: 40 countries
  • Level: Expert

This role involves judging how frontier AI models respond to sensitive and ambiguous topics, checking answers for safety, accuracy and policy compliance. It suits experienced professionals from fields such as psychology, law, science or public policy who can apply safety rubrics and give structured feedback to model developers.

Posted July 16, 2026

OpenRecently verified$60-70/hr

Referral link: we may earn a fee. Apply without it

  • Data & Machine Learning
  • Platform: Mercor
  • Location: 40 countries
  • Level: Expert

This role involves stress-testing frontier AI models with adversarial prompts to uncover jailbreaks, unsafe behaviour and policy failures on sensitive topics. It is aimed at experienced safety, security, life sciences or policy professionals who can document weaknesses and help researchers strengthen model alignment.

Posted July 16, 2026

OpenRecently verified$70-84/hr

Referral link: we may earn a fee. Apply without it

  • Business, Consulting & Operations
  • Platform: Mercor
  • Location: United States
  • Level: Intermediate

This mission consists of writing and refining annotation guidelines for training state-of-the-art language models. You will transform ambiguous specifications into clear and non-contradictory directives for annotators, working across varied domains such as finance, insurance, and law. This role combines expertise in linguistics, instructional design, and technical writing to ensure the precision and consistency of training data.

Posted July 10, 2026

OpenRecently verified$45-65/hr

Referral link: we may earn a fee. Apply without it

  • Business, Consulting & Operations
  • Platform: Mercor
  • Location: United States
  • Level: Expert

This role involves coordinating data annotation operations for AI training models, with specialization in the financial domain. You will lead collaboration between finance-specialized annotators and the program team, ensuring deadlines and quality standards are met. The ideal candidate combines solid experience in finance or operational management with proven skills in project coordination and team leadership.

Posted July 10, 2026

OpenRecently verified$40-60/hr

Referral link: we may earn a fee. Apply without it

  • Finance & Accounting
  • Platform: Mercor
  • Location: United States
  • Level: Expert

This role consists of evaluating and improving the financial reasoning capabilities of AI models as a domain specialist. You will design rigorous financial tasks, evaluate model responses according to defined criteria, and collaborate with research teams to strengthen the quality of training data. This role is suited to finance professionals with solid experience in leading institutions and familiarity with evaluating AI systems.

Posted July 10, 2026

OpenRecently verified$65-90/hr

Referral link: we may earn a fee. Apply without it

  • Finance & Accounting
  • Platform: Mercor
  • Location: United States
  • Level: Expert

This role invites insurance experts to contribute to the development of generative AI models by rigorously evaluating algorithm outputs against structured evaluation frameworks. You will leverage your experience in underwriting, claims management, or risk assessment to refine training data and guide research teams on points of business expertise. The role combines designing complex insurance tasks, critically evaluating AI outputs, and developing domain-specific scoring criteria.

Posted July 10, 2026

OpenRecently verified$60-80/hr

Referral link: we may earn a fee. Apply without it

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

679 AI Evaluation AI training jobs are open on SideHustler today. 26 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 637 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (139), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.