Role Overview In this hourly, remote contractor role, you will review AI-generated English-language responses and/or create expert-level language content. You will evaluate reasoning quality, perform step-by-step edits, and provide precise, actionable feedback to improve AI…
Remote job listings
Find your next remote opportunity from thousands of listings across the globe.
Mercor is seeking bilingual AI Safety Experts fluent in both English and Punjabi to help evaluate and strengthen the safety of frontier AI systems. This role focuses on red teaming AI models by identifying vulnerabilities, testing misuse scenarios, and generating high-quality…
Evaluate the safety, quality, and alignment of frontier AI models by reviewing policy-sensitive content, applying AI safety standards, identifying model failures, and helping improve next-generation AI systems.
Help secure frontier AI models by designing adversarial prompts, identifying jailbreaks and safety vulnerabilities, evaluating AI behavior across high-risk domains, and improving AI alignment through structured red-team assessments.
Design and evaluate performance engineering tasks for frontier AI systems using C++, Python, and Rust while improving runtime performance, systems optimization, compiler engineering, and AI infrastructure training data.
Design and evaluate MLOps engineering tasks for frontier AI systems using JAX, PyTorch, and custom GPU kernels with Pallas or Triton while improving ML infrastructure, distributed training, and AI model evaluation.
Evaluate frontier AI systems by testing complex real-world workflows using Model Context Protocol (MCP) and AI plugins/connectors, creating evaluation rubrics, and assessing AI performance across travel, productivity, health, and personal planning tasks.
Help train frontier AI models by reviewing, evaluating, and annotating AI-generated content using academic research expertise and structured evaluation guidelines.
Seeking experienced marketing professionals to design evaluation rubrics, assess AI-generated marketing work, and improve frontier AI systems through rigorous, evidence-based evaluation.
Seeking experienced B2B sales professionals to design evaluation rubrics, assess AI-generated enterprise sales work, and improve frontier AI systems through rigorous, evidence-based evaluation.
Seeking experienced accounting professionals to design evaluation rubrics, assess AI-generated accounting work, and improve frontier AI systems through rigorous, evidence-based evaluation.
Seeking experienced Data Scientists to design evaluation rubrics, assess AI-generated data science work, and improve frontier AI systems through rigorous, evidence-based evaluation.
Seeking experienced marketing professionals with AI evaluation experience to design marketing tasks, evaluate LLM outputs, develop scoring rubrics, and improve the marketing reasoning capabilities of frontier AI models.
Seeking experienced insurance professionals with AI evaluation experience to design insurance tasks, evaluate LLM outputs, develop scoring rubrics, and improve the insurance reasoning capabilities of frontier AI models.
Seeking experienced retail professionals with AI evaluation experience to design retail tasks, evaluate LLM outputs, develop scoring rubrics, and improve the retail reasoning capabilities of frontier AI models.
Seeking experienced finance professionals with AI evaluation experience to design finance tasks, evaluate LLM outputs, develop scoring rubrics, and improve the financial reasoning capabilities of frontier AI models.
Seeking experienced linguists, instructional designers, and technical writers to create AI rater guidelines, evaluation rubrics, and annotation instructions that improve the consistency and quality of frontier AI training data.
Collaborate with autonomous AI coding agents, evaluate AI-generated code, and help improve agentic software engineering workflows for next-generation AI systems.
Evaluate AI-generated responses to complex accounting problems, helping improve AI systems' accounting capabilities. Collaborate with leading AI researchers and apply your professional accounting expertise to real-world AI evaluation.
Seeking users with hands-on experience using enterprise AI agent systems to provide structured feedback on real-world workflows and tool behavior. This fully remote contractor role offers flexible scheduling and pays $30 per hour.
Graduate students evaluate and validate AI outputs requiring deep academic expertise across disciplines. Flexible, project-based, and remote work with a commitment of up to 10 hours per week.
This is a remote, project-based opportunity for experienced professionals across any domain who can apply real-world expertise to AI data annotation and evaluation tasks. Contributors will review, label, and evaluate: Text samples Documents AI model outputs Domain-specific…
Job Summary micro1 is engaging Math Experts (PhD) to contribute to an advanced AI training project for a customer. In this role, you will apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and perform through…
Job Summary micro1 is engaging Chemistry Experts (PhD) to contribute their advanced subject-matter expertise to a customer's AI training initiative. In this role, you will apply your expertise to help train next-generation AI systems. Your work will shape how models learn,…
Job Summary micro1 is engaging Biology Experts (PhD) to contribute to a dynamic customer project focused on advancing artificial intelligence. In this role, you will apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and…
Job Summary micro1 is engaging Physics (PhD) experts to contribute their advanced knowledge to a high-impact AI development project. In this role, you will apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and perform…
Job Summary micro1 is engaging PhD-level Engineers in Electrical, Mechanical, or Chemical disciplines to contribute to a high-impact customer project. In this role, you will apply your expertise to help train next-generation AI systems. Your work will shape how models learn,…
Mercor is seeking Automated CS Planning QA Reviewers for a premier project with one of the world's leading AI laboratories. In this role, you will review and quality-check advanced symbolic AI and automated planning problems created by domain experts. Your work will help ensure…
Mercor is seeking Automated CS Planning PhDs and Domain Experts for a premier project with one of the world's leading AI laboratories. In this role, you will apply advanced expertise in automated planning, symbolic AI, artificial intelligence, machine learning, and related…
Mercor is partnering with a leading frontier AI research laboratory on an advanced mathematics project focused on improving AI reasoning capabilities. The project seeks individuals with demonstrated expertise in Olympiad-level mathematics, mathematical problem writing, and…
About the Role Cincinnatus LLC — in partnership with Mercor — is recruiting accomplished research mathematicians to contribute to the development of next-generation AI systems capable of advanced mathematical and scientific reasoning. Contributors will work with leading AI…
About the Role Mercor is seeking a senior Computer Vision Machine Learning Engineer (MLE) for a remote, contingent W2 engagement focused on evaluating the feasibility of a computer vision system that identifies and grades physical objects from images. The initial engagement is…
About the Role Mercor is hiring Cybersecurity Software Engineers to help improve the safety and reliability of advanced AI systems. In this role, you will apply your cybersecurity and software engineering expertise to evaluate, annotate, and strengthen how AI models handle…
About the Role Mercor is seeking experienced Mechanical Engineering professionals to support a high-impact quality assurance and technical review initiative focused on improving engineering content quality and accuracy. In this role, you will evaluate, validate, and improve…
Mercor is seeking bilingual AI Safety Experts fluent in both English and Marathi to help evaluate and strengthen the safety of frontier AI systems. This role focuses on red teaming AI models by identifying vulnerabilities, testing misuse scenarios, and generating high-quality…
Mercor is building a benchmark dataset to evaluate AI models on professional document understanding and instruction following within the Technology domain. Contributors will create high-quality evaluation tasks based on real-world technical materials including: Technical…
About the Project We're building a large-scale evaluation benchmark for advanced AI reasoning across scientific and engineering domains. Task designers create challenging computational problems that test whether AI systems can use real scientific software tools to solve…
About the Project We're building a large-scale evaluation benchmark for advanced AI reasoning across scientific and engineering domains. Task designers create challenging computational problems that test whether AI systems can use real scientific software tools to solve…
About the Project We're building a large-scale evaluation benchmark for advanced AI reasoning across scientific and engineering domains. Task designers create challenging computational problems that test whether AI systems can use real scientific software tools to solve…
About the Project We're building a large-scale evaluation benchmark for advanced AI reasoning across scientific and engineering domains. Task designers create challenging computational problems that test whether AI systems can use real scientific software tools to solve…
About the Project We're building a large-scale evaluation benchmark for advanced AI reasoning across scientific and engineering domains. Task designers create challenging computational problems that test whether AI systems can use real scientific software tools to solve…
About the Project We're building a large-scale evaluation benchmark for advanced AI reasoning across scientific and engineering domains. Task designers create challenging computational problems that test whether AI systems can use real scientific software tools to solve…
About the Project We're building a large-scale evaluation benchmark for advanced AI reasoning across scientific and engineering domains. Task designers create challenging computational problems that test whether AI systems can use real scientific software tools to solve…
About the Project We're building a large-scale evaluation benchmark for advanced AI reasoning across scientific and engineering domains. Task designers create challenging computational problems that test whether AI systems can use real scientific software tools to solve…
<h2>About the Role</h2> <p>As a Senior Python Engineer, you will review AI-generated Python solutions and technical explanations while creating high-quality reference content that helps improve advanced AI systems.</p> <p>You will evaluate:</p> <ul> <li>Code correctness</li>…
Mercor is seeking bilingual AI Safety Experts fluent in both English and Odia to help evaluate and strengthen the safety of frontier AI systems. This role focuses on red teaming AI models by identifying vulnerabilities, testing misuse scenarios, and generating high-quality…
Mercor is seeking bilingual AI Safety Experts fluent in both English and Gujarati to help evaluate and strengthen the safety of frontier AI systems. This role focuses on red teaming AI models by identifying vulnerabilities, testing misuse scenarios, and generating high-quality…
Mercor is seeking bilingual AI Safety Experts fluent in both English and Assamese to help evaluate and strengthen the safety of frontier AI systems. This role focuses on red teaming AI models by identifying vulnerabilities, testing misuse scenarios, and generating high-quality…
Tips for finding remote jobs
- Set up job alerts on multiple platforms to never miss an opportunity.
- Highlight your remote work experience and self-management skills.
- Prepare for video interviews and remote work assessments.
- Customize your resume and cover letter for each remote position.
- Build a strong online presence on LinkedIn and professional networks.