OpenRecently verified

Senior Software Engineer - AI Evaluation

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

$60-120/hrAs published by the platform.

About this job

This remote contract role involves designing and building the software that measures how well AI models perform, including evaluation pipelines, automated testing harnesses, and internal dashboards and APIs. It suits senior engineers who have shipped production systems and want to work closely with AI research teams on reliable, repeatable evaluation tooling.

What you'll do

  • Design and build scalable evaluation pipelines and frameworks for AI model performance.
  • Develop automated testing harnesses, scoring systems, and benchmarking tools for language models.
  • Create and maintain APIs, dashboards, and internal tools for running and comparing evaluations.
  • Identify edge cases and failure modes in AI outputs through systematic engineering approaches.

Requirements

PythonAWSGCPAzureGitCI/CDPandasNumPyLLM
  • 4+ years of experience

Pay

$60-120/hr

Pay as published by the platform. It is not a guarantee of income or hours.

Location

The platform has not published which countries are eligible for this job.

How to apply

You apply on Alignerr, Labelbox's expert network: create a profile, complete an AI-led interview and skills assessment, then get matched to projects in your field.

All Alignerr jobs and how the platform works

Source

Official posting: https://www.alignerr.com/jobs/cfa052f0-e721-4ec5-8d38-d697e534081b

Last checked on October 8, 2026.

Similar jobs

  • Software Engineering
  • Platform: Alignerr
  • Level: Expert

This contract role asks a senior software engineer to build and scale the backend infrastructure behind AI products, working closely with machine learning engineers and researchers. It suits experienced engineers who can work independently on production-grade services, data pipelines, and APIs.

  • Software Engineering
  • Platform: Mercor
  • Location: United States
  • Level: Expert

This role invites experienced software engineers to evaluate and design migration strategies for complex codebases, particularly in mission-critical environments. You will analyze existing systems to identify safe modernization approaches and verify the quality of solutions generated by AI, while preserving functionality and security.

Posted August 22, 2026

OpenRecently verified$200/hr

Referral link: we may earn a fee. Apply without it

  • Software Engineering
  • Platform: Alignerr
  • Level: Intermediate

This role involves evaluating and providing feedback on AI-generated Python code for correctness, efficiency, and security while designing complex backend algorithmic solutions. It is designed for experienced backend Python developers who want to contribute to cutting-edge AI research by helping train the next generation of AI systems to write better code.

  • Software Engineering
  • Platform: xAI

This role supports AI model training by writing, correcting, and curating code examples in specialized programming languages. It suits experienced software engineers who can judge whether code meets professional standards for performance, scalability, and reliability.

Posted February 26, 2026

OpenRecently verifiedPay not disclosed
  • Software Engineering
  • Platform: DataAnnotation

Evaluate how AI models handle DevOps tasks such as CI/CD pipelines, infrastructure-as-code, and cloud operations. Identify failures in automation, security configurations, and deployment processes, then write correct solutions that a platform engineer would trust.

Talent poolRecently verified$40-150/hr

Keep exploring