AI Evaluation AI training jobs

Open AI training and model evaluation roles that call for AI Evaluation expertise, gathered from 10 platforms. Part of Data & Machine Learning.

679
open jobs
26
posted this week
$75/hr
median published rate
$400/hr
highest published rate
10
platforms hiring

Pay figures cover the 637 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

21 to 40 of 679 jobs
  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Arabic

This role involves recording high-quality Saudi Arabic voice samples and evaluating AI-generated speech to help train and improve AI voice models. The position is ideal for native or near-native Saudi Arabic speakers with voice acting or narration experience who can provide constructive feedback on pronunciation, tone, and naturalness.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Spanish

This role involves recording Spanish voice samples, evaluating AI-generated speech outputs, and providing feedback to improve AI voice systems. It is designed for native or near-native Spanish speakers with voice acting or narration experience who can work remotely on a flexible contract basis.

  • Finance & Accounting
  • Platform: Alignerr

This role involves evaluating algorithmic trading strategies built into AI systems, stress-testing their logic, and identifying edge cases to ensure sound reasoning in automated trading decisions. It is designed for experienced algorithmic traders who can critically assess trading systems without requiring prior AI training experience.

  • Business, Consulting & Operations
  • Platform: Alignerr
  • Language: Spanish
  • Level: Intermediate

Lead a team of Spanish-speaking business analysts creating high-quality training content for AI models across finance, strategy, and economics domains. This remote role combines people management, domain expertise, and content development for an AI training platform, requiring business analysis experience and team leadership skills.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This role involves analyzing real-world application security scenarios and creating training datasets to teach AI models how to assess software vulnerabilities. You will evaluate code, APIs, and system behavior to distinguish genuinely exploitable risks from theoretical findings, providing structured analysis that helps frontier AI systems learn to reason about application security with the nuance of experienced engineers.

  • Science & Research
  • Platform: Alignerr

This role involves translating mathematical proofs into machine-verifiable Lean 4 formalizations for an AI research organization. Mathematicians with formal proof system experience will work on bridging informal mathematical arguments and automated reasoning, identifying gaps in current verification tools and helping advance mechanized mathematics at the intersection of AI and rigorous proof.

OpenRecently verified$170-200/hr
  • Science & Research
  • Platform: Alignerr

This role invites PhD-level applied physicists to evaluate and improve how AI models reason about physical systems. You will design advanced physics problems, author rigorous solutions, audit AI-generated reasoning for physical consistency, and provide expert feedback to refine model behavior across classical mechanics, electrodynamics, quantum mechanics, and related domains.

  • Science & Research
  • Platform: Alignerr

This role involves evaluating and improving large language models' understanding of physics by designing advanced problems, authoring ground-truth solutions, and auditing AI-generated reasoning across quantum mechanics, electrodynamics, thermodynamics, and other domains. It is suited for PhD-level applied physicists who want to contribute to AI training while working independently on a flexible schedule.

  • Languages, Translation & Voice
  • Platform: Alignerr
  • Language: Arabic

This role involves evaluating and improving AI systems trained on Arabic-language content. Arabic language experts will review AI-generated text for linguistic accuracy, cultural appropriateness, and natural phrasing, providing structured feedback to enhance how AI communicates with Arabic speakers worldwide.

  • Science & Research
  • Platform: Alignerr

This role invites atmospheric science experts with Masters or PhD credentials to develop rigorous, advanced problems in climate modeling, meteorology, and ecology that train AI systems to reason at the highest scientific level. You will design complex problem statements, evaluate AI-generated reasoning, and collaborate with AI researchers to advance scientific AI capabilities.

  • Writing, Creative & Design
  • Platform: Alignerr

An audio engineering role focused on evaluating and improving AI systems trained on audio content. You will produce, mix, and master audio while reviewing AI-generated outputs and providing feedback to refine audio processing tools.

  • Languages, Translation & Voice
  • Platform: Alignerr

Audio Transcription Specialists convert real audio recordings into accurate written text to power speech recognition and language models. This fully remote, flexible contract role is open to anyone with strong listening skills and attention to detail, with no professional transcription background required.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This role involves evaluating and providing feedback on AI-generated Python code for correctness, efficiency, and security while designing complex backend algorithmic solutions. It is designed for experienced backend Python developers who want to contribute to cutting-edge AI research by helping train the next generation of AI systems to write better code.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This role involves evaluating and providing feedback on AI-generated Python code while designing complex algorithmic challenges to test AI reasoning capabilities. Senior Python developers will work remotely on flexible contracts to help advance AI systems' ability to write and understand production-level code.

  • Software Engineering
  • Platform: Alignerr
  • Level: Expert

This role involves reviewing and evaluating AI-generated Rust code to ensure correctness, memory safety, and systems-level performance. Experienced Rust engineers will assess the quality of AI outputs, design challenging problems to test AI understanding, and provide structured feedback to improve how AI models reason about systems programming.

  • Data & Machine Learning
  • Platform: Alignerr

A fully remote contract role where PhD-level biochemists create expert-level problems, author gold-standard solutions, and audit AI-generated biochemical reasoning to train and improve scientific AI models. Ideal for biochemists passionate about scientific rigor who want to shape the accuracy of next-generation AI without requiring prior AI experience.

  • Science & Research
  • Platform: Alignerr

This role involves helping train and refine AI models on biochemistry topics by designing test problems, writing ground-truth solutions, and evaluating AI-generated scientific reasoning. The position is ideal for PhD-level biochemists who want to apply their expertise to AI development while working flexibly on a remote, hourly contract basis.

  • Science & Research
  • Platform: Alignerr

This is a contract position for biologists with PhDs to annotate and evaluate training data for AI models. The role involves labeling biology questions and concepts, mapping relationships between biological principles, reviewing AI-generated answers, and designing training datasets across biology subfields.

  • Data & Machine Learning
  • Platform: Alignerr

This role involves designing and evaluating advanced biology problems to help train AI models to reason rigorously about scientific topics like genetics, biotechnology, and bioinformatics. Masters and PhD holders in biology will apply their domain expertise to assess AI-generated reasoning for scientific accuracy and clarity in a fully remote, flexible contract position.

  • Science & Research
  • Platform: Alignerr

This role involves evaluating and improving AI models' understanding of biology by designing rigorous problems, developing solutions, and assessing AI-generated outputs for scientific accuracy and reasoning quality. It is designed for biology researchers and academics who can work independently on a flexible, remote contract basis.

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about AI Evaluation AI training jobs

How many AI Evaluation AI training jobs are open right now?

679 AI Evaluation AI training jobs are open on SideHustler today. 26 were posted in the last 7 days. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do AI Evaluation AI training jobs pay?

Among the 637 open roles that publish an hourly rate in USD, pay runs from $8 to $400 per hour, with a median of $75. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer AI Evaluation AI training jobs?

Alignerr (253), Mercor (196), micro1 (139), DataAnnotation (47), Terac (13), Vetto (11), Handshake AI (7), Mindrift (6), xAI (4) and Ethos (3). You apply on the platform itself, which handles screening, contracts and payment.

What skills do AI Evaluation AI training jobs ask for most?

The most requested areas of expertise on these listings are AI Evaluation, Data Annotation, Content Annotation, Code Quality & Review, Multilingual Expertise. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.