AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

679
open jobs
26
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 637 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

61 to 80 of 679 jobs
  • Business, Consulting & Operations
  • Platform: Alignerr

A Clinical Business Intelligence Manager will leverage healthcare analytics expertise to support AI training projects by evaluating AI-generated clinical content, leading BI teams, and maintaining analytics infrastructure. This remote contract role is designed for experienced healthcare data professionals who can apply domain knowledge to ensure AI outputs are accurate and aligned with clinical practice.

  • Data & Machine Learning
  • Platform: Alignerr

A Clinical Data Strategist analyzes complex multi-source clinical datasets to identify trends and generate actionable insights that improve patient outcomes and healthcare operations. This fully remote contract role is suited for experienced healthcare analytics professionals who can transform raw clinical data into structured summaries for downstream use.

  • Business, Consulting & Operations
  • Platform: Alignerr

This role involves managing operational aspects of clinical trials while contributing to AI model training. Experienced Clinical Study Managers will oversee trial timelines, budgets, and multi-vendor coordination while helping evaluate AI-generated clinical content, working remotely on a flexible contract basis.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This role involves analyzing real-world cloud security scenarios to identify risks, misconfigurations, and failure patterns across AWS, GCP, and Azure. Contractors will evaluate AI-generated security reasoning and help create datasets for training AI models to better understand cloud infrastructure vulnerabilities and incidents.

  • Software Engineering
  • Platform: Alignerr

This role involves reviewing and evaluating C++ code generated by AI systems to assess its correctness, efficiency, and style. You will provide detailed feedback to help train AI models to better understand C++ programming across various domains and difficulty levels.

  • Software Engineering
  • Platform: Alignerr

This role involves reviewing and evaluating AI-generated Go code, writing clean Go solutions for backend problems, and creating coding challenges to test AI model capabilities. It is ideal for experienced Go developers who want to contribute to the training and improvement of AI coding models.

  • Software Engineering
  • Platform: Alignerr

Review and evaluate AI-generated Java code for correctness and efficiency, and write high-quality Java solutions to coding problems at varying difficulty levels. This role is ideal for Java developers who can also provide clear explanations of code logic and problem-solving strategies to support AI model training.

  • Software Engineering
  • Platform: Alignerr

Review and evaluate AI-generated JavaScript code for correctness and efficiency, and write high-quality JavaScript solutions to coding problems. Create clear explanations for code logic and identify edge cases. This role is ideal for JavaScript experts who want to contribute to improving AI models.

  • Software Engineering
  • Platform: Alignerr

Review and evaluate AI-generated Rust code for correctness, memory safety, and idiomatic practices. Create high-quality coding problems and provide structured feedback to improve AI model performance. This role is ideal for experienced Rust developers who want to shape how AI systems understand systems programming.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This role involves evaluating and writing SQL queries for AI model training projects. You will assess AI-generated SQL for accuracy and performance, write complex queries across relational databases, and document your reasoning to support AI development at leading research labs and enterprises.

  • Engineering
  • Platform: Alignerr

This role involves evaluating and improving AI models' understanding of computer engineering concepts by designing technical challenges, creating reference solutions, and auditing AI-generated code and hardware designs. The position is ideal for computer engineering experts who want to help train advanced language models on complex topics like architecture, embedded systems, and hardware design.

  • Data & Machine Learning
  • Platform: Alignerr

This role involves evaluating and improving AI models that process visual data, including images and video. Computer vision and machine learning experts will assess model outputs, identify failure modes, create training annotations, and provide technical feedback to guide the development of AI systems used at scale.

OpenRecently verified$100-150/hr
  • Business, Consulting & Operations
  • Platform: Alignerr

This role involves evaluating AI-generated text for quality, accuracy, and appropriateness across diverse topics. Content reviewers provide structured feedback to help improve AI tools used by millions, working independently on a flexible remote schedule.

  • Writing, Creative & Design
  • Platform: Alignerr

This role involves reading and evaluating AI-generated creative writing across multiple genres, assessing literary qualities such as style, coherence, tone, originality, and engagement. Evaluators provide constructive feedback to improve AI writing systems, requiring a strong grasp of language and the ability to articulate what makes writing effective.

  • Business, Consulting & Operations
  • Platform: Alignerr

This role involves designing realistic customer support task scenarios and evaluation rubrics for AI agent training. You will author ticket triage tasks, create scoring rubrics, set up task environments, and calibrate difficulty to ensure AI models learn to make sound support decisions similar to experienced professionals.

  • Business, Consulting & Operations
  • Platform: Alignerr

Design and author realistic customer support triage tasks that train and evaluate AI agents on real enterprise work, including ticket triage, account configuration checks, and escalation decisions. This role is for experienced support professionals who can create detailed task scenarios with clear scoring rubrics that reflect actual judgment calls made by support engineers.

  • Software Engineering
  • Platform: Alignerr

This role involves evaluating and stress-testing AI-generated cybersecurity content to ensure accuracy and safety. Cybersecurity Defence Analysts review threat analyses, incident response plans, and defensive recommendations while identifying errors and dangerous hallucinations, helping ensure AI systems provide trustworthy security guidance.

  • Software Engineering
  • Platform: Alignerr

This role involves evaluating and stress-testing AI-generated cybersecurity content to ensure accuracy and safety. Cybersecurity Defense Analysts will review threat analyses, incident response plans, and defensive strategies produced by AI systems, identifying errors and dangerous misconceptions while providing detailed feedback to improve AI security guidance.

  • Software Engineering
  • Platform: Alignerr

This role involves evaluating and stress-testing AI-generated cybersecurity content to ensure accuracy and safety. Cybersecurity experts with hands-on defense experience will review threat analyses, incident response plans, and security recommendations, identifying errors and gaps that could lead to real harm in production environments.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Czech

This role involves evaluating and improving AI systems' understanding and generation of Czech-language content. You will review AI-generated Czech translations and original content, identify linguistic issues, and provide structured feedback to enhance the quality and cultural appropriateness of AI outputs in Czech.

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

679 AI Evaluation AI training jobs are open on SideHustler today. 26 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 637 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (139), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.