OpenRecently verified

Professionals: Designing Challenging AI Prompts

  • Data & Machine Learning
  • Platform: Terac
  • Location: United States

About this job

This is a paid one-hour trial for people who know a professional or personal workflow well. Participants turn that workflow into a demanding prompt, test it in ChatGPT to find where the model fails, and make the prompt harder when the model succeeds. The work suits domain experts and power users who can judge AI outputs quickly and who are open to ongoing evaluation work.

What you'll do

  • Convert a familiar real-world workflow into a demanding prompt that requires reasoning and lookup.
  • Test the prompt in ChatGPT, record failure points, and refine it until the model breaks.
  • Write a clear grading rubric that another evaluator could apply to any AI attempt at the task.
  • Submit the prompt, failure notes, rubric, and the generated output file.

Requirements

Prompt engineeringAI output evaluationRubric design

    Pay

    Pay not disclosed

    The platform does not publish the pay for this job.

    Location

    Open to residents of: United States.

    How to apply

    You apply on Terac's posting: start its screening from the link in the listing. If you qualify, Terac invites you to the paid tasks.

    All Terac jobs and how the platform works

    Source

    Official posting: https://jobs.ashbyhq.com/terac/b7732bd7-0046-4a98-b6ed-2dac344b51f8

    Last checked on October 8, 2026.

    Similar jobs

    • Data & Machine Learning
    • Platform: Alignerr

    This role involves testing and evaluating AI chatbot responses across diverse topics by engaging in conversations, identifying issues, and providing structured feedback. The position is designed for individuals with strong critical thinking and communication skills who want to contribute to AI safety and improvement without requiring prior technical or AI experience.

    • Data & Machine Learning
    • Platform: Alignerr

    This role involves testing and evaluating AI models by designing adversarial prompts and scenarios to identify weaknesses, biases, and unsafe outputs. Red team testers document discovered failure modes and assess their severity to help improve AI safety before deployment.

    • Data & Machine Learning
    • Platform: Mercor
    • Location: United States
    • Language: Turkish

    You will participate in improving AI model safety by evaluating how they handle sensitive subjects in Turkish. Your linguistic and cultural judgments will help identify and strengthen weaknesses in these systems when facing delicate content. No prior AI experience is required.

    Posted September 4, 2026

    OpenRecently verified$23-27/hr

    Referral link: we may earn a fee. Apply without it

    • Data & Machine Learning
    • Platform: Mercor
    • Location: United States
    • Language: Danish

    You join a red teaming team specialized in adversarial evaluation of conversational AI models. Your role consists of testing AI systems by exploring their vulnerabilities (jailbreaks, prompt injections, biases), generating high-quality data documenting these vulnerabilities, and producing reproducible reports to strengthen model safety. This position is for bilingual English-Danish experts with prior experience in red teaming or related fields (cybersecurity, adversarial ML, socio-technical analysis).

    Posted July 30, 2026

    OpenRecently verified$48-62/hr

    Referral link: we may earn a fee. Apply without it

    • Software Engineering
    • Platform: DataAnnotation

    This role involves evaluating how AI models reason about offensive security concepts and identifying flaws in their exploit chain reasoning. Penetration testers with hands-on red team experience will test model outputs across reconnaissance, exploitation, privilege escalation, and lateral movement, then write accurate attack paths that reflect real-world tradecraft when models fall short.

    Talent poolRecently verified$40-125/hr
    • Software Engineering
    • Platform: DataAnnotation

    You will evaluate AI-generated code by running models on real engineering tasks, analyzing their output against production standards, and stress-testing them to find failures. This role is ideal for experienced software engineers who want to contribute to AI model improvement through rigorous code review and red-teaming.

    Talent poolRecently verified$40-150/hr

    Keep exploring