Software Engineer , AI Code Evaluation & Benchmarking (US candidates only)

Legal & Compliance • Independent Contractor • Remote • via Turing

Evaluate and improve AI-generated code for accuracy and quality.

About Turing

Turing is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems. Turing helps leading AI labs improve the reasoning, problem-solving, and decision-making capabilities of large language models (LLMs) through high-quality human feedback, evaluation, and training data.

Role Overview

We are looking for experienced Software Engineers to help evaluate, benchmark, and improve the coding capabilities of frontier AI models. In this role, you will assess AI-generated code, validate solutions against real-world software engineering tasks, identify correctness and quality issues, and contribute to the development of high-quality evaluation datasets and benchmarks.

This position is ideal for engineers who enjoy code review, debugging, problem-solving, and applying strong software engineering judgment to complex technical scenarios. Your work will directly contribute to measuring and improving the performance of advanced AI coding systems.

What Does Day-to-Day Look Like?

Requirements

Perks of Freelancing With Turing

Offer Details

Evaluation Process

  1. Online automated coding challenge for Python and Docker test (RHLF)

Details

Apply directly through Turing using the link below. Turing matches vetted experts to frontier AI labs, with fast onboarding and remote flexibility.

Stay Updated on Roles Like This

Subscribe to receive fresh openings aligned with Legal & Compliance and AI training roles across Turing, Mercor and JobHub by NeonLabs