AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

679
open jobs
26
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 637 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

1 to 20 of 679 jobs
  • Finance & Accounting
  • Platform: Alignerr

This role involves evaluating and improving AI-generated accounting and finance content to ensure it meets professional standards. Experienced accountants will review AI outputs on US GAAP, SEC reporting, financial modeling, and regulatory compliance, providing detailed feedback to enhance AI understanding of financial workflows and concepts.

  • Finance & Accounting
  • Platform: Alignerr
  • Level: Intermediate

This role involves evaluating and improving AI systems trained on accounting and financial reporting content. Experienced accounting instructors will review AI-generated accounting questions, explanations, and scenarios to assess accuracy and identify errors, providing structured feedback to enhance the quality of AI financial tools.

  • Software Engineering
  • Platform: Alignerr

This role involves analyzing cyber attacks, adversary tactics, and exploitation chains to produce structured assessments that help AI systems understand cybersecurity threats better. It is designed for experienced offensive security professionals who can evaluate threat behaviors and translate complex technical findings into clear documentation.

  • Data & Machine Learning
  • Platform: Alignerr

This role involves testing and evaluating AI chatbot responses across diverse topics by engaging in conversations, identifying issues, and providing structured feedback. The position is designed for individuals with strong critical thinking and communication skills who want to contribute to AI safety and improvement without requiring prior technical or AI experience.

  • Data & Machine Learning
  • Platform: Alignerr

Review and evaluate AI-generated outputs across text, images, and structured data to assess quality, accuracy, and consistency. This remote contract role is ideal for detail-oriented individuals who can identify errors and provide actionable feedback to guide AI model improvement, with no prior AI experience required.

  • Software Engineering
  • Platform: Alignerr

This role involves analyzing security vulnerabilities and threat scenarios in AI systems and large language models to identify weaknesses and recommend mitigations. Security professionals will conduct adversarial testing, evaluate real-world attack scenarios, and provide structured feedback to improve AI safety and resilience.

  • Data & Machine Learning
  • Platform: Alignerr

This role involves reviewing AI-generated content to identify and flag bias, safety issues, and ethical concerns that could affect global audiences. Reviewers apply fairness and safety guidelines to ensure AI systems are responsible and trustworthy, providing structured written feedback on diverse content types.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Chinese

This role involves evaluating and improving AI systems designed for Chinese language learning by assessing the linguistic accuracy, naturalness, and educational quality of AI-generated speech and text, as well as learner output across proficiency levels. The position is ideal for Chinese language experts with teaching or evaluation experience who want to directly influence how AI understands and teaches Chinese.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: French

Evaluate and improve AI systems trained on French-language content by assessing the linguistic accuracy, naturalness, and educational quality of AI-generated French speech and text. This role is for French language experts with teaching backgrounds who will provide structured feedback to enhance AI tutoring and language-learning models for learners at various proficiency levels.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Korean

Evaluate and improve Korean language AI systems by assessing AI-generated speech and text, as well as learner language across proficiency levels. This role combines linguistic expertise with educational knowledge to refine AI tutoring and language-learning models through detailed, structured feedback.

  • Business, Consulting & Operations
  • Platform: Alignerr

This role involves evaluating AI-generated content against safety, fairness, and ethical standards to ensure AI systems align with human values. The position suits those with strong critical thinking and clear communication skills who are interested in ethics and policy, with no prior AI or policy experience required.

  • Software Engineering
  • Platform: Alignerr

Penetration testing experts are needed to identify security vulnerabilities in AI-powered applications and infrastructure through offensive security testing. This remote contract role combines traditional penetration testing methodologies with emerging AI-specific attack scenarios such as prompt injection and model manipulation.

  • Data & Machine Learning
  • Platform: Alignerr

This role involves testing and evaluating AI models by designing adversarial prompts and scenarios to identify weaknesses, biases, and unsafe outputs. Red team testers document discovered failure modes and assess their severity to help improve AI safety before deployment.

  • Software Engineering
  • Platform: Alignerr

This role involves red-teaming and security testing of AI models to identify vulnerabilities and safety issues before deployment. The position is ideal for cybersecurity professionals who want to apply adversarial thinking and penetration testing skills to AI systems used at scale.

  • Software Engineering
  • Platform: Alignerr

This role focuses on stress-testing and securing AI models by identifying vulnerabilities through red-teaming and adversarial testing. The position suits cybersecurity professionals who want to apply penetration testing and ethical hacking expertise to emerging AI security challenges.

  • Software Engineering
  • Platform: Alignerr

This role involves conducting security testing and red-teaming exercises on advanced AI models to identify vulnerabilities, design adversarial prompts, and evaluate safety failures. The specialist will document findings and collaborate with engineering teams to strengthen AI system resilience and reliability.

  • Data & Machine Learning
  • Platform: Alignerr

This role involves completing structured evaluation and labeling tasks to train and improve AI systems. You will assess AI-generated content across multiple formats and provide feedback to help make AI systems more accurate and reliable, working remotely on a flexible schedule.

  • Data & Machine Learning
  • Platform: Alignerr

AI Training Specialists label, classify, rate, and evaluate AI-generated content across multiple formats to improve AI system accuracy and reliability. This remote, flexible contract role suits detail-oriented individuals committed to quality work on structured AI training tasks.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: French

This role involves recording high-quality French voice samples and evaluating AI-generated speech for naturalness and expressiveness. The contractor will provide feedback to refine AI voice outputs and review scripts for clarity, supporting the development of frontier AI models.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Italian

This role involves recording high-quality voice samples in Italian and evaluating AI-generated speech outputs for naturalness and expressiveness. Contractors will provide feedback to refine AI voice synthesis and review scripts for clarity, working flexibly on an asynchronous schedule to help develop frontier AI speech technology.

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

679 AI Evaluation AI training jobs are open on SideHustler today. 26 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 637 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (139), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.