OpenRecently verified

Software Engineers: Paid Code Review for AI Agent Evaluation

  • Software Engineering
  • Platform: Terac
  • Location: United States

$65/hrAs published by the platform.

About this job

This remote study asks practicing software engineers to review programming tasks and the evaluation harnesses built to test AI agents on them. You will check whether the tasks, test cases and environment structure reflect realistic software engineering work. It suits engineers with solid experience building, testing and reviewing complex systems.

What you'll do

  • Review realistic programming tasks for accuracy and difficulty.
  • Verify the logic and test cases within provided evaluation harnesses.
  • Assess whether coding environments effectively measure software engineering skills.
  • Explain your reasoning aloud while analyzing complex code structures.

Requirements

Machine Learning

    Pay

    $65/hr

    Pay as published by the platform. It is not a guarantee of income or hours.

    Location

    Open to residents of: United States.

    How to apply

    You apply on Terac's posting: start its screening from the link in the listing. If you qualify, Terac invites you to the paid tasks.

    All Terac jobs and how the platform works

    Source

    Official posting: https://jobs.ashbyhq.com/terac/997d9c88-e299-468b-9c9a-88d4a5b77b88

    Last checked on October 8, 2026.

    Similar jobs

    • Software Engineering
    • Platform: Alignerr
    • Level: Intermediate

    This remote contract role asks senior TypeScript developers to review AI-generated code for type safety, scalability and modern practices. The work centers on writing structured feedback that helps AI models understand architectural choices and type logic. It is aimed at experienced developers who can work independently on flexible hourly commitments.

    • Software Engineering
    • Platform: Alignerr
    • Level: Intermediate

    This contract role asks TypeScript developers to review and improve AI-generated code so it is type-safe, idiomatic, scalable and ready for production. It is aimed at senior developers with deep TypeScript knowledge who can clearly explain why a piece of code is well designed. No prior AI background is required.

    • Software Engineering
    • Platform: Mercor
    • Location: United States
    • Level: Intermediate

    You will evaluate the quality and correctness of AI-assisted software development traces used to train and assess models at a leading AI laboratory. You must judge the accuracy of coding sessions, the consistency of workflows, and the quality of reasoning, then provide structured written feedback according to an evaluation grid.

    Posted August 28, 2026

    OpenRecently verified$70-90/hr

    Referral link: we may earn a fee. Apply without it

    • Software Engineering
    • Platform: DataAnnotation

    You will evaluate AI-generated code by running models on real engineering tasks, analyzing their output against production standards, and stress-testing them to find failures. This role is ideal for experienced software engineers who want to contribute to AI model improvement through rigorous code review and red-teaming.

    Talent poolRecently verified$40-150/hr
    • Software Engineering
    • Platform: Terac
    • Location: Argentina

    This paid study asks experienced software engineers to review proposed programming tasks and the harnesses that test AI agents, checking their logic, structure and realism. The work takes place in a screen-shared session that mixes code review, technical discussion and direct feedback on task design. It is aimed at engineers who have already built, reviewed or tested evaluation harnesses or coding tasks.

    Posted October 1, 2026

    Keep exploring