Back to remote jobs

Machine Learning Engineer — Model Evaluation & Experimentation

Mercor

AI Expert - Machine Learning & AI Research Full-time Contingent Role
United States $60 – $90/hr July 25, 2026

Job description

Role Overview

Cincinnatus LLC is seeking experienced Machine Learning Engineers to help develop next-generation evaluation benchmarks for frontier AI models. In this fully remote W-2 role, you will design complex machine learning experiments, implement research ideas, run end-to-end training workflows, and evaluate how advanced AI systems perform on realistic ML research tasks.

Working alongside researchers at a leading AI lab, you will transform open-ended machine learning concepts into rigorous evaluation benchmarks that test the reasoning, experimentation, and implementation capabilities of frontier AI models.

Key Responsibilities

Design Machine Learning Tasks

  • Create realistic, multi-step machine learning challenges based on real-world research workflows.
  • Translate research ideas into well-defined evaluation tasks involving implementation, experimentation, and analysis.
  • Design benchmarks that push the capabilities of frontier AI models.

Run ML Experiments

  • Implement machine learning solutions in Python.
  • Configure, execute, and analyze end-to-end training experiments.
  • Produce reference implementations and document expected outcomes.

Evaluate AI Models

  • Assess how frontier AI models perform on machine learning tasks.
  • Identify implementation errors, reasoning failures, and experimental shortcomings.
  • Provide detailed feedback to improve AI evaluation benchmarks.

Collaborate with Research Teams

  • Work closely with researchers and fellow experts to refine evaluation tasks.
  • Review benchmark quality, consistency, and technical accuracy.
  • Contribute to the development of robust AI evaluation methodologies.

Required Qualifications

  • Master's degree or PhD in:

- Machine Learning

- Computer Science

- Another STEM discipline

- Or equivalent practical experience in a research-intensive domain

  • At least 1 year of experience in:

- Machine Learning Research

- Research Engineering

- Research-focused software development

  • Hands-on experience with:

- Training machine learning models

- Running end-to-end ML experiments

- Experiment design, execution, and analysis

  • Strong familiarity with:

- Large Language Models (LLMs)

- Model evaluation techniques

  • Proficiency with:

- Python

- Git

- Notebook-based development environments

  • Excellent written communication skills.
  • High attention to detail and ability to work independently.
  • Available to work approximately 35 hours per week.

Preferred Qualifications

  • Understanding of reinforcement learning concepts, including:

- Reward functions

- Policy training

  • Experience with:

- AI model evaluation

- AI training

- Benchmark development

- Task authoring

  • Strong problem-solving skills in ambiguous, research-oriented environments.

Compensation

  • $60–$90 per hour
  • Full-time W-2 employment

Work Arrangement

  • Fully Remote (United States)
  • W-2 Contingent Role
  • Approximately 35 hours per week
  • Opportunity to work with a leading AI laboratory through Cincinnatus LLC
Apply now

You will be redirected to the company's website to complete your application.

Apply now

Stay in the loop.

One email per week, 5 hand-picked roles.