OpenRecently verified

AI Agent Evaluator

  • Data & Machine Learning
  • Platform: micro1
  • Location: 58 countries
  • Level: Intermediate

$25-45/hrAs published by the platform.

About this job

This contract role asks you to test two AI agents on everyday personal errands such as paying bills, booking travel and buying groceries, using your own accounts and phone. You approve or stop the agent at key steps, record each attempt and score the results. It suits people who already use AI tools in daily life and want to help shape how these agents improve.

What you'll do

  • Run everyday errands with two AI agents side by side on your own accounts.
  • Approve or stop the agent at key steps and screen record every attempt.
  • Score each run on completion, quality, control, satisfaction and time.
  • Share clear written feedback on the agent experience.

Requirements

ChatGPTMuseAI agentsAttention to DetailFollowing InstructionsScreen Recording
  • 2+ years of experience
  • Bachelor's degree

Pay

$25-45/hr

Pay as published by the platform. It is not a guarantee of income or hours.

Location

Open to residents of: Bangladesh, Hong Kong SAR China, India, Indonesia, Japan, Kazakhstan, Kyrgyzstan, Malaysia, Pakistan, Philippines, Singapore, Sri Lanka, Taiwan, Thailand, Uzbekistan, Vietnam, Austria, Belarus, Belgium, Denmark, France, Germany, Greece, Italy, Netherlands, Portugal, Russia, Spain, Switzerland, United Kingdom, Argentina, Brazil, Chile, Colombia, Mexico, Peru, Algeria, Bahrain, Egypt, Iraq, Jordan, Kuwait, Lebanon, Libya, Morocco, Oman, Palestinian Territories, Qatar, Saudi Arabia, Tunisia, United Arab Emirates, United States, Canada, Nigeria, Kenya, South Africa, Ghana, Ethiopia.

How to apply

You apply on micro1: create an account, complete a screening interview with its AI interviewer, then join its vetted expert network to be matched to projects.

All micro1 jobs and how the platform works

Source

Official posting: https://jobs.micro1.ai/post/2fea94b5-2f19-4287-b776-f061b413a00f

Last checked on October 10, 2026.

Similar jobs

  • Data & Machine Learning
  • Platform: Alignerr

This role involves testing and evaluating AI chatbot responses across diverse topics by engaging in conversations, identifying issues, and providing structured feedback. The position is designed for individuals with strong critical thinking and communication skills who want to contribute to AI safety and improvement without requiring prior technical or AI experience.

  • Data & Machine Learning
  • Platform: Alignerr

This remote, hourly contract involves checking and assessing geographic information produced by AI systems, such as addresses, place names, and location search results. It suits people who know their own city or region well and can spot errors in digital maps, with no GIS or geography degree needed.

  • Data & Machine Learning
  • Platform: Mercor
  • Location: United States
  • Level: Intermediate

Mercor is recruiting research experts capable of improving and evaluating cutting-edge AI models. You will contribute to creating high-quality training materials (questions, answers, evaluations) for advanced language models. This role is aimed at candidates holding a Master's degree or with 2 to 3 years of experience, with strong research and communication skills.

Posted July 26, 2026

OpenRecently verified$50-60/hr

Referral link: we may earn a fee. Apply without it

  • Business, Consulting & Operations
  • Platform: DataAnnotation

This role asks experienced consultants to test cutting-edge AI models on business problems drawn from their own practice. Contributors write client-style briefs, review each response as they would an associate's work, and flag where recommendations are sharp and where they are generic. It suits people with consulting, strategy, PMO, or deal advisory backgrounds.

Talent poolRecently verified$40-125/hr
  • Finance & Accounting
  • Platform: micro1
  • Location: 58 countries

This job consists of evaluating and annotating U.S. tax return processes and financial workflows to train AI systems. You must apply your tax expertise to assess the accuracy and completeness of these workflows, comparing system-generated results with real documents and scenarios.

Posted September 30, 2026

OpenRecently verified$30-110/hr

Referral link to micro1's job list: search for this role there. Open this exact job

  • Business, Consulting & Operations
  • Platform: micro1
  • Location: 58 countries
  • Level: Intermediate

This role invites you to mobilize your expertise in interviewing and investigative questioning to contribute to training next-generation AI systems. You will conduct structured elicitation sessions, analyze responses in real time by detecting verbal and behavioral cues, and document your analyses according to precise frameworks. The ideal profile possesses solid experience in professional interviews, intelligence information gathering, or investigative questioning.

Posted September 18, 2026

OpenRecently verified$25-65/hr

Referral link to micro1's job list: search for this role there. Open this exact job

Keep exploring

Apply on micro1 (opens micro1 in a new tab)

This referral link opens micro1's list of open roles, not this exact job. Once there, search for AI Agent Evaluator. We may earn a fee if you are selected and meet the platform's conditions.

Open this exact job, without the referral link