OpenRecently verified

Senior Software Engineer - AI Evaluation & Benchmarks

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

$80-100/hrAs published by the platform.

About this job

This remote contract role involves building coding benchmarks and data pipelines that test how well frontier AI models reason about, debug, and write software. It is aimed at senior software engineers with several years of professional experience who can work across large, multi-language codebases.

What you'll do

  • Design and implement coding benchmarks for evaluating frontier AI models.
  • Build and maintain scalable data pipelines for evaluation workflows.
  • Analyze AI-generated code for correctness, reliability, and edge-case failures.
  • Provide detailed technical feedback on model performance and failure patterns.

Requirements

PythonJavaScriptC++GitCI/CDMachine LearningLLM
  • 4+ years of experience

Pay

$80-100/hr

Pay as published by the platform. It is not a guarantee of income or hours.

Location

The platform has not published which countries are eligible for this job.

How to apply

You apply on Alignerr, Labelbox's expert network: create a profile, complete an AI-led interview and skills assessment, then get matched to projects in your field.

All Alignerr jobs and how the platform works

Source

Official posting: https://www.alignerr.com/jobs/51c15c59-a780-4d32-a969-b4e5ea7ee6a3

Last checked on October 8, 2026.

Similar jobs

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This remote contract role involves designing and building the software that measures how well AI models perform, including evaluation pipelines, automated testing harnesses, and internal dashboards and APIs. It suits senior engineers who have shipped production systems and want to work closely with AI research teams on reliable, repeatable evaluation tooling.

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This role involves evaluating and providing feedback on AI-generated Python code for correctness, efficiency, and security while designing complex backend algorithmic solutions. It is designed for experienced backend Python developers who want to contribute to cutting-edge AI research by helping train the next generation of AI systems to write better code.

  • Software Engineering
  • Platform: DataAnnotation

You will evaluate AI-generated code by running models on real engineering tasks, analyzing their output against production standards, and stress-testing them to find failures. This role is ideal for experienced software engineers who want to contribute to AI model improvement through rigorous code review and red-teaming.

Talent poolRecently verified$40-150/hr
  • Software Engineering
  • Platform: micro1
  • Location: 58 countries

You will contribute to training next-generation AI models by creating realistic coding tasks and implementing deterministic verifiers. This flexible contracting role is aimed at experienced backend engineers capable of generating relevant technical challenges and automatically evaluating solutions.

Posted October 6, 2026

OpenRecently verified$30-100/hr

Referral link to micro1's job list: search for this role there. Open this exact job

  • Software Engineering
  • Platform: Mercor
  • Location: United States
  • Level: Intermediate

This role involves evaluating the quality, correctness, and reproducibility of software engineering tasks used to train and test the AI models of a leading laboratory. You audit tasks at the repository level, baseline patches, test harnesses, and scoring system integrity, providing structured feedback based on an evaluation grid.

Posted August 28, 2026

OpenRecently verified$70-90/hr

Referral link: we may earn a fee. Apply without it

Keep exploring