AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

679
open jobs
26
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 637 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

81 to 100 of 679 jobs
  • Data & Machine Learning
  • Platform: Alignerr

This role involves categorizing, tagging, and annotating text, images, and audio data to train AI systems. Data Labeling Specialists work independently on a flexible remote contract, following structured guidelines to ensure consistent, high-quality dataset preparation for machine learning models.

  • Data & Machine Learning
  • Platform: Alignerr

This role involves managing and preparing datasets for AI model training by cleaning, organizing, and quality-checking data to ensure accuracy and consistency. The position suits detail-oriented individuals with strong organizational skills who want to contribute to AI development without requiring prior machine learning experience.

  • Software Engineering
  • Platform: Alignerr
  • Level: Expert

A senior Python engineer will design and optimize high-performance data pipelines, annotation tooling, and evaluation systems for AI model training at scale. This fully remote contract role requires building full-stack infrastructure that processes large datasets and supports research workflows across distributed teams.

  • Data & Machine Learning
  • Platform: Alignerr

This role involves designing advanced data science challenges, writing benchmark solutions, and auditing AI-generated code to improve AI reasoning and reliability. It is designed for experienced data scientists who want flexible, intellectually demanding work training and evaluating machine learning systems.

  • Data & Machine Learning
  • Platform: Alignerr

This role involves designing complex data science challenges and authoring ground-truth solutions to train and evaluate advanced AI models. Data scientists will audit AI-generated code, identify logical reasoning failures, and provide structured feedback to improve model performance across machine learning, statistical inference, and data engineering domains.

  • Data & Machine Learning
  • Platform: Alignerr
  • Level: Intermediate

This role involves analyzing data security and Data Loss Prevention scenarios to help train AI models that understand real-world data risk. You will classify sensitive information, evaluate security strategies, and create labeled datasets for AI model development based on your hands-on security expertise.

  • Science & Research
  • Platform: Alignerr
  • Level: Intermediate

This role involves evaluating and improving AI systems trained on development economics content by reviewing explanations, scenarios, and reasoning. Development economists assess the accuracy and quality of AI outputs on topics like poverty, inequality, and global development, ensuring the AI communicates economic concepts with appropriate depth and nuance.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

Design and evaluate infrastructure and automation challenges to train and benchmark AI systems used by leading research labs and Fortune 500 companies. This remote contract role is for experienced DevOps engineers to create realistic cloud infrastructure scenarios, write Infrastructure as Code, and assess how well AI models understand and reason about complex DevOps concepts.

  • Data & Machine Learning
  • Platform: Alignerr

Review, categorize, and refine product data for AI-powered shopping and recommendation systems. This role involves evaluating product listings, AI-generated descriptions, and ensuring data quality to improve product discovery. Ideal for detail-oriented individuals comfortable with online platforms and structured data.

  • Science & Research
  • Platform: Alignerr
  • Level: Intermediate

This role involves reviewing and validating econometric questions, models, and explanations used in AI training datasets. Instructors with expertise in econometrics and quantitative economics will assess statistical accuracy, identify errors in AI-generated economic reasoning, and provide structured feedback to improve model outputs. The work is remote and asynchronous, suited for experienced econometrics educators.

  • Business, Consulting & Operations
  • Platform: Alignerr
  • Level: Intermediate

An experienced postsecondary economics instructor evaluates and improves AI systems by designing questions, reviewing AI-generated content, and providing expert feedback on economic reasoning and explanations. This remote role leverages teaching expertise to enhance how artificial intelligence communicates economic concepts to learners and professionals.

  • Health & Medicine
  • Platform: Alignerr

This role involves annotating clinical data from Electronic Health Records to train AI systems. Clinically knowledgeable professionals will review and label medical information including diagnoses, procedures, and clinical notes to improve AI understanding of healthcare data.

  • Engineering
  • Platform: Alignerr

This role involves designing challenging electrical engineering problems, writing authoritative solutions, and auditing AI-generated technical outputs to improve how artificial intelligence reasons through complex engineering tasks. It is suited for experienced electrical engineers who want to contribute to AI model development remotely on a flexible contract basis.

  • Business, Consulting & Operations
  • Platform: Alignerr

This role involves evaluating and validating AI-generated content related to EMR/EHR systems, clinical workflows, and healthcare IT best practices. The specialist leverages hands-on experience with platforms like Epic and Cerner to assess accuracy, identify errors in AI outputs, and provide detailed feedback to improve AI models' understanding of real-world healthcare operations.

  • Software Engineering
  • Platform: Alignerr

This role involves writing, debugging, and evaluating AI-generated code to help train and improve AI coding assistants. Software engineers will review code quality, identify bugs and security issues, provide detailed feedback, and design test scenarios to stress-test AI capabilities across various programming languages and frameworks.

  • Business, Consulting & Operations
  • Platform: Alignerr
  • Level: Intermediate

This role involves evaluating and improving AI systems trained on startup strategy and business planning by reviewing AI-generated entrepreneurship scenarios for accuracy and practical relevance. Experienced entrepreneurship instructors and business educators will provide structured feedback on business models, go-to-market strategies, and startup advice to enhance AI outputs used by founders and investors.

  • Engineering
  • Platform: Alignerr

This role involves applying environmental engineering expertise to improve AI models' reasoning about physical and chemical processes. You will create technical problems, develop reference solutions, and audit AI-generated engineering outputs to identify and correct modeling failures across water treatment, air quality, remediation, and environmental compliance domains.

  • Science & Research
  • Platform: Alignerr

This role involves training advanced AI models to correctly understand environmental engineering concepts and identify flaws in their reasoning across water treatment, air quality, and hazardous waste remediation. You will design complex technical scenarios, write authoritative solutions, and critically evaluate AI outputs to improve how models think about environmental systems.

  • Science & Research
  • Platform: Alignerr
  • Level: Intermediate

This remote contract role involves evaluating AI-generated environmental management analyses and recommendations to improve their accuracy and real-world applicability. Environmental management professionals will assess AI reasoning on sustainability frameworks, impact mitigation, and resource planning while providing structured feedback to enhance AI outputs.

  • Science & Research
  • Platform: Alignerr

This remote position involves developing and reviewing advanced environmental science problems to support AI model training. Environmental science experts with advanced degrees will apply their knowledge in climate modeling, ecology, or sustainability to create complex problem statements and enhance AI model reasoning through collaboration with research teams.

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

679 AI Evaluation AI training jobs are open on SideHustler today. 26 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 637 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (139), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.