OpenRecently verified

AI Evaluation Specialist

  • Data & Machine Learning
  • Platform: Handshake AI
  • Location: United States

Up to $40/hrAs published by the platform.

About this job

Test and evaluate Large Language Models through remote, part-time projects in collaboration with AI research labs. Apply your analytical skills and reasoning abilities to help improve AI model performance.

What you'll do

  • Evaluate LLM responses and outputs according to detailed instructions and rubrics
  • Complete training and pass an assessment before starting project work
  • Identify issues and provide feedback to help improve model performance
  • Work independently on short-term AI training projects at your own pace

Requirements

LLM
  • Associate's degree, Bachelor's degree, or Master's degree

Pay

Up to $40/hr

Pay as published by the platform. It is not a guarantee of income or hours.

Location

Open to residents of: United States.

How to apply

You apply on Handshake AI: create or sign in to your Handshake account, complete the application for the program, then get placed on projects as they open.

All Handshake AI jobs and how the platform works

Source

Official posting: https://joinhandshake.com/ai/opportunities/generalist-bachelor-ai-eveluation-specialist

Last checked on October 8, 2026.

Similar jobs

  • Data & Machine Learning
  • Platform: Alignerr

This role involves testing and evaluating AI chatbot responses across diverse topics by engaging in conversations, identifying issues, and providing structured feedback. The position is designed for individuals with strong critical thinking and communication skills who want to contribute to AI safety and improvement without requiring prior technical or AI experience.

  • Data & Machine Learning
  • Platform: Alignerr

Review and evaluate AI-generated outputs across text, images, and structured data to assess quality, accuracy, and consistency. This remote contract role is ideal for detail-oriented individuals who can identify errors and provide actionable feedback to guide AI model improvement, with no prior AI experience required.

  • Data & Machine Learning
  • Platform: DataAnnotation
  • Level: Intermediate

This role involves evaluating how well AI models can perform computational biology tasks. Computational biologists and bioinformaticians will design realistic analysis scenarios from their own work, run them through AI systems, and grade the results against professional standards to help benchmark model capabilities.

Talent poolRecently verified$40-125/hr
  • Data & Machine Learning
  • Platform: DataAnnotation

As a Data Scientist, you will evaluate analyses generated by AI models on real datasets, checking whether the statistical methods and conclusions are sound. You will identify flaws in the reasoning, stress-test the logic, and write detailed assessments that serve as training data for future model improvements.

Talent poolRecently verified$40-150/hr

Keep exploring