Design and execute machine learning experiments, build evaluation benchmarks, and analyze frontier AI model performance in a full-time remote W-2 role supporting a leading AI research lab.
STEM Researcher — Computational Fields
Job description
Role Overview
Cincinnatus LLC is seeking experienced STEM Researchers specializing in computational fields to help develop next-generation evaluation benchmarks for frontier AI models. In this fully remote W-2 role, you will transform real-world scientific research workflows into complex, multi-step evaluation tasks that measure the reasoning, experimentation, and analytical capabilities of advanced AI systems.
Working alongside researchers at a leading AI lab, you will design rigorous benchmark tasks based on experimental design, hypothesis testing, computational analysis, and scientific methodology, helping define the standards for high-quality AI reasoning across computational research disciplines.
Key Responsibilities
Design Research Benchmarks
- Create realistic, multi-step research tasks based on computational scientific workflows.
- Design benchmarks involving study design, hypothesis testing, coding, data analysis, and scientific interpretation.
- Develop evaluation challenges that reflect authentic research practices.
Author Reference Solutions
- Complete your own benchmark tasks using Python and notebook environments.
- Produce reproducible analyses and well-documented reference solutions.
- Demonstrate rigorous scientific reasoning and methodology.
Evaluate AI Performance
- Assess AI-generated solutions for methodological soundness, analytical accuracy, and scientific rigor.
- Identify weaknesses in experimental design, implementation, and interpretation.
- Provide detailed feedback to improve AI evaluation benchmarks.
Collaborate with Research Teams
- Review benchmark tasks created by fellow researchers.
- Work closely with AI researchers to refine evaluation standards.
- Help define what distinguishes robust scientific reasoning from superficial analysis.
Required Qualifications
- Master's degree or PhD in:
- STEM discipline
- Computational Social Science
- Computational Humanities
- Or equivalent practical experience in a research-intensive computational field
- At least 1 year of experience in:
- Academic research
- Industry research
- National laboratory research
- Hands-on experience with:
- Python
- Computational analysis
- Simulation
- Modeling
- Data pipelines
- Strong understanding of:
- Experimental design
- Hypothesis testing
- Scientific evaluation
- Proficiency with:
- Git
- Integrated Development Environments (IDEs)
- Jupyter Notebook
- Google Colab
- Excellent written communication skills.
- High attention to detail and ability to work independently.
- Available to work approximately 35 hours per week.
Preferred Qualifications
- Experience with:
- AI model evaluation
- AI training
- Benchmark development
- Task authoring
- Strong problem-solving skills in research-oriented, open-ended environments.
Compensation
- $60–$90 per hour
- Full-time W-2 employment
Work Arrangement
- Fully Remote (United States)
- W-2 Contingent Role
- Approximately 35 hours per week
- Opportunity to work with a leading AI laboratory through Cincinnatus LLC
You will be redirected to the company's website to complete your application.