Design computational research benchmarks, build reproducible scientific analyses, and evaluate frontier AI models on complex research and experimental reasoning tasks in a full-time remote W-2 role.
Machine Learning Engineer — Model Evaluation & Experimentation
Job description
Role Overview
Cincinnatus LLC is seeking experienced Machine Learning Engineers to help develop next-generation evaluation benchmarks for frontier AI models. In this fully remote W-2 role, you will design complex machine learning experiments, implement research ideas, run end-to-end training workflows, and evaluate how advanced AI systems perform on realistic ML research tasks.
Working alongside researchers at a leading AI lab, you will transform open-ended machine learning concepts into rigorous evaluation benchmarks that test the reasoning, experimentation, and implementation capabilities of frontier AI models.
Key Responsibilities
Design Machine Learning Tasks
- Create realistic, multi-step machine learning challenges based on real-world research workflows.
- Translate research ideas into well-defined evaluation tasks involving implementation, experimentation, and analysis.
- Design benchmarks that push the capabilities of frontier AI models.
Run ML Experiments
- Implement machine learning solutions in Python.
- Configure, execute, and analyze end-to-end training experiments.
- Produce reference implementations and document expected outcomes.
Evaluate AI Models
- Assess how frontier AI models perform on machine learning tasks.
- Identify implementation errors, reasoning failures, and experimental shortcomings.
- Provide detailed feedback to improve AI evaluation benchmarks.
Collaborate with Research Teams
- Work closely with researchers and fellow experts to refine evaluation tasks.
- Review benchmark quality, consistency, and technical accuracy.
- Contribute to the development of robust AI evaluation methodologies.
Required Qualifications
- Master's degree or PhD in:
- Machine Learning
- Computer Science
- Another STEM discipline
- Or equivalent practical experience in a research-intensive domain
- At least 1 year of experience in:
- Machine Learning Research
- Research Engineering
- Research-focused software development
- Hands-on experience with:
- Training machine learning models
- Running end-to-end ML experiments
- Experiment design, execution, and analysis
- Strong familiarity with:
- Large Language Models (LLMs)
- Model evaluation techniques
- Proficiency with:
- Python
- Git
- Notebook-based development environments
- Excellent written communication skills.
- High attention to detail and ability to work independently.
- Available to work approximately 35 hours per week.
Preferred Qualifications
- Understanding of reinforcement learning concepts, including:
- Reward functions
- Policy training
- Experience with:
- AI model evaluation
- AI training
- Benchmark development
- Task authoring
- Strong problem-solving skills in ambiguous, research-oriented environments.
Compensation
- $60–$90 per hour
- Full-time W-2 employment
Work Arrangement
- Fully Remote (United States)
- W-2 Contingent Role
- Approximately 35 hours per week
- Opportunity to work with a leading AI laboratory through Cincinnatus LLC
You will be redirected to the company's website to complete your application.