OpenRecently verified

AI Chatbot Tester

  • Data & Machine Learning
  • Platform: Alignerr

$40-120/hrAs published by the platform.

About this job

This role involves testing and evaluating AI chatbot responses across diverse topics by engaging in conversations, identifying issues, and providing structured feedback. The position is designed for individuals with strong critical thinking and communication skills who want to contribute to AI safety and improvement without requiring prior technical or AI experience.

What you'll do

  • Engage in conversations with AI chatbots and evaluate their responses for helpfulness, accuracy, tone, and safety
  • Identify issues such as incorrect information, awkward phrasing, logical errors, or potential bias in AI responses
  • Craft creative and challenging prompts to stress-test AI capabilities and find limitations
  • Submit structured, detailed feedback to research teams to guide AI performance improvements

Requirements

Critical thinkingWritten communicationResponse evaluationPrompt engineering

    Pay

    $40-120/hr

    Pay as published by the platform. It is not a guarantee of income or hours.

    Location

    The platform has not published which countries are eligible for this job.

    How to apply

    You apply on Alignerr, Labelbox's expert network: create a profile, complete an AI-led interview and skills assessment, then get matched to projects in your field.

    All Alignerr jobs and how the platform works

    Source

    Official posting: https://www.alignerr.com/jobs/08b8a45c-bfb9-4d7b-a05f-13fc28095431

    Last checked on October 8, 2026.

    Similar jobs

    • Business, Consulting & Operations
    • Platform: Mercor
    • Location: United States
    • Language: Croatian

    This role involves evaluating and improving the safety of AI models by analyzing their behavior on sensitive topics in Croatian. You will write expert prompts, classify content according to structured guidelines, and identify adversarial formulations, as an English-Croatian bilingual speaker with no prior AI experience required.

    Posted September 4, 2026

    OpenRecently verified$38-42/hr

    Referral link: we may earn a fee. Apply without it

    • Data & Machine Learning
    • Platform: Alignerr

    This role involves testing and evaluating AI models by designing adversarial prompts and scenarios to identify weaknesses, biases, and unsafe outputs. Red team testers document discovered failure modes and assess their severity to help improve AI safety before deployment.

    • Data & Machine Learning
    • Platform: Alignerr

    This contract role asks security-minded people to probe AI systems built on the OpenClaw platform by running red-teaming exercises and writing adversarial prompts. The work suits cybersecurity practitioners who enjoy finding weaknesses and who can document their findings clearly for technical and non-technical readers.

    • Data & Machine Learning
    • Platform: Mercor
    • Location: United States
    • Language: Norwegian

    This role consists of evaluating and strengthening the safety of advanced AI models by analyzing their behavior on sensitive topics in Norwegian. You will use your bilingual proficiency (English-Norwegian) and cultural judgment to identify gaps and bypass attempts, with no prior AI experience required.

    Posted September 4, 2026

    OpenRecently verified$58-62/hr

    Referral link: we may earn a fee. Apply without it

    • Data & Machine Learning
    • Platform: Terac
    • Location: United States

    This is a paid one-hour trial for people who know a professional or personal workflow well. Participants turn that workflow into a demanding prompt, test it in ChatGPT to find where the model fails, and make the prompt harder when the model succeeds. The work suits domain experts and power users who can judge AI outputs quickly and who are open to ongoing evaluation work.

    Posted August 25, 2026

    OpenRecently verifiedPay not disclosed
    • Software Engineering
    • Platform: DataAnnotation

    This role involves evaluating how AI models reason about offensive security concepts and identifying flaws in their exploit chain reasoning. Penetration testers with hands-on red team experience will test model outputs across reconnaissance, exploitation, privilege escalation, and lateral movement, then write accurate attack paths that reflect real-world tradecraft when models fall short.

    Talent poolRecently verified$40-125/hr

    Keep exploring