Back to remote jobs

QA/Test Engineer

Mercor

AI Expert - Software Engineering Full-time Contingent Role
United States $60 – $90/hr July 25, 2026

Job description

Role Overview

Cincinnatus LLC is seeking experienced QA/Test Engineers to help develop next-generation evaluation benchmarks for frontier AI models. In this fully remote W-2 role, you will ensure complex AI evaluation tasks are accurate, unambiguous, and robust by designing comprehensive test cases, validating reference solutions, and identifying issues before benchmarks are released.

Working alongside researchers at a leading AI lab, you will serve as the quality assurance backbone for AI benchmark development by verifying that every task measures the intended skills and produces reliable, trustworthy evaluation results.

Key Responsibilities

Design Quality Assurance Tests

  • Create comprehensive test cases that validate benchmark tasks, including edge cases and failure scenarios.
  • Develop quality assurance checks that verify correctness, robustness, and grading accuracy.
  • Ensure evaluation benchmarks function as intended under diverse conditions.

Review and Validate Tasks

  • Review benchmark tasks and reference solutions for clarity, completeness, and technical correctness.
  • Identify ambiguous instructions, grading gaps, and implementation issues before publication.
  • Verify that benchmarks accurately assess the intended engineering skills.

Debug and Improve Benchmarks

  • Debug Python implementations, testing environments, and validation scripts.
  • Investigate unexpected behavior and resolve technical issues affecting benchmark quality.
  • Improve testing workflows and quality assurance processes.

Collaborate with Research Teams

  • Work closely with researchers and task authors to refine benchmark quality.
  • Develop repeatable quality review processes and testing standards.
  • Identify shortcuts or grading weaknesses that could affect AI evaluation results.

Required Qualifications

  • Master's degree or PhD in:

- STEM discipline

- Or equivalent practical experience in a research-intensive or engineering-focused domain

  • At least 1 year of experience in:

- Quality Assurance

- Test Engineering

- Software Engineering

- Research Engineering with significant quality ownership

  • Strong experience with:

- Test case design

- Quality review processes

- End-to-end debugging

  • Proficiency with:

- Python

- Git

- Software development environments

  • Excellent written documentation and communication skills.
  • Exceptional attention to detail.
  • Available to work approximately 35 hours per week.

Preferred Qualifications

  • Experience with:

- AI model evaluation

- AI training

- Quality review of AI-generated work

  • Experience designing benchmark validation frameworks.
  • Strong problem-solving skills in complex engineering environments.

Compensation

  • $60–$90 per hour
  • Full-time W-2 employment

Work Arrangement

  • Fully Remote (United States)
  • W-2 Contingent Role
  • Approximately 35 hours per week
  • Opportunity to work with a leading AI laboratory through Cincinnatus LLC
Apply now

You will be redirected to the company's website to complete your application.

Apply now

Stay in the loop.

One email per week, 5 hand-picked roles.