Help improve frontier AI models by designing advanced mathematics problems, evaluating AI reasoning, developing computational solutions in Python, and contributing to formal theorem-proving tasks using Lean in a flexible remote contractor role.
Remote job listings
Find your next remote opportunity from thousands of listings across the globe.
Help improve frontier AI models by creating advanced chemistry problems, evaluating AI reasoning, and developing multimodal scientific content featuring chemical equations, reaction mechanisms, and detailed explanations in a flexible remote contractor role.
Help improve frontier AI models by designing advanced Biology questions, evaluating AI reasoning, and creating high-quality scientific solutions in this flexible remote contractor role for Master's and PhD-level Biology experts.
Review AI-generated Indian domestic tax research, validate statutory authorities, improve grading rubrics, and ensure technical accuracy across Indian tax topics in this flexible remote contractor role.
Review AI-generated Singapore domestic tax research, validate IRAS guidance and statutory authorities, improve grading rubrics, and ensure technical accuracy across Singapore tax topics in this flexible remote contractor role.
Review AI-generated UK domestic tax research, validate legal authorities, improve grading rubrics, and ensure technical accuracy across UK tax topics in this flexible remote contractor role.
Mercor is seeking bilingual AI Safety Experts fluent in both English and Punjabi to help evaluate and strengthen the safety of frontier AI systems. This role focuses on red teaming AI models by identifying vulnerabilities, testing misuse scenarios, and generating high-quality…
Oerview Mercor is seeking Generalist Experts to evaluate AI-generated responses and provide structured written feedback that helps improve frontier AI systems. This role is ideal for analytical thinkers with strong reasoning and communication skills who can assess complex…
LIVE VIRTUAL ENROLLMENT EVENT – Thai AI Evaluators Needed! Earn $250 USD Plus, if you don't have a Google Pixel phone, we'll provide one for you! We're looking for Thai speakers living in the United States to participate in an exciting Google Pixel AI evaluation project.
LIVE VIRTUAL ENROLLMENT EVENT – Polish AI Evaluators Needed! Earn $250 USD Plus, if you don't have a Google Pixel phone, we'll provide one for you! We're looking for Polish speakers living in the United States to participate in an exciting Google Pixel AI evaluation project.
LIVE VIRTUAL ENROLLMENT EVENT – Korean AI Evaluators Needed! Earn $300 USD Plus, if you don't have a Google Pixel phone, we'll provide one for you! We're looking for Korean speakers living in the United States to participate in an exciting Google Pixel AI evaluation project.
LIVE VIRTUAL ENROLLMENT EVENT – Indian English AI Evaluators Needed! Earn $225 USD Plus, if you don't have a Google Pixel phone, we'll provide one for you!
Review and improve AI-generated Catalan content, create gold-standard training data, evaluate language quality, and help build more accurate, culturally appropriate AI systems in this flexible remote contractor role.
Evaluate AI-generated Catalan text and audio, provide grammar and cultural corrections, and help improve AI language understanding through high-quality language evaluations completed on a flexible weekly schedule.
Get Paid as a Dutch Expert Rater | Remote | $27.97/hr Are you an experienced Dutch-speaking Rater looking for your next remote opportunity? Join Welo Data and help improve the quality of AI systems while working from anywhere.
Get Paid as a German Expert Rater | Remote | $23.71/hr Are you an experienced German Rater looking for your next remote opportunity? Join Welo Data and help improve the quality of AI systems while working from home.
Evaluate the safety, quality, and alignment of frontier AI models by reviewing policy-sensitive content, applying AI safety standards, identifying model failures, and helping improve next-generation AI systems.
Help secure frontier AI models by designing adversarial prompts, identifying jailbreaks and safety vulnerabilities, evaluating AI behavior across high-risk domains, and improving AI alignment through structured red-team assessments.
Design and evaluate performance engineering tasks for frontier AI systems using C++, Python, and Rust while improving runtime performance, systems optimization, compiler engineering, and AI infrastructure training data.
Design and evaluate MLOps engineering tasks for frontier AI systems using JAX, PyTorch, and custom GPU kernels with Pallas or Triton while improving ML infrastructure, distributed training, and AI model evaluation.
Evaluate frontier AI systems by testing complex real-world workflows using Model Context Protocol (MCP) and AI plugins/connectors, creating evaluation rubrics, and assessing AI performance across travel, productivity, health, and personal planning tasks.
Evaluate AI-generated legal analyses involving housing, family law, consumer protection, foreclosure, and debt matters while improving legal reasoning and access-to-justice guidance through expert civil legal review.
Review AI-generated civil legal guidance involving housing, family law, consumer debt, and bankruptcy while evaluating procedural accuracy, documentation quality, and real-world legal workflows.
Evaluate and refine AI-generated consulting presentations, executive slide decks, and partner-ready deliverables using MBB-level engagement management and presentation expertise. This is a flexible remote consulting opportunity for experienced consulting leaders who want to apply their expertise to improving frontier AI systems.
Role Title: AI Evaluation Analyst Role Type: Contractor Location: Remote micro1 is engaging AI Evaluation Analysts to contribute to a customer’s project focused on advancing frontier language model capabilities.
In this remote contractor role, you will work as an Attorney / Lawyer Subject Matter Expert (SME) reviewing AI-generated legal responses and creating expert legal content. The work focuses on evaluating legal reasoning, issue spotting, rule application, procedural analysis, and…
In this remote contractor role, you will work as a Contract Review Subject Matter Expert (SME) reviewing AI-generated contract analyses and drafting outputs while creating expert contract review content. The work focuses on evaluating contract interpretation, risk allocation,…
In this remote contractor role, you will work as a Compliance & Regulatory Subject Matter Expert (SME) reviewing AI-generated compliance, regulatory, and risk-related responses and creating expert compliance content. The work focuses on evaluating regulatory interpretation,…
In this remote contractor role, you will work as a Paralegal Subject Matter Expert (SME) reviewing AI-generated legal support responses and creating expert paralegal content. The work focuses on evaluating reasoning quality, legal support workflows, document drafting, legal…
In this remote contractor role, you will work as an Intellectual Property (IP) Subject Matter Expert (SME) reviewing AI-generated IP-related responses and creating expert intellectual property content. The work focuses on evaluating legal reasoning, IP issue spotting, rights…
Evaluate AI-generated responses to complex medical and biomedical problems, helping improve the clinical reasoning and safety of frontier AI systems.
What We're Researching We are running a paid project focused on creating and validating complex coding tasks for AI systems. The goal is to build realistic, self-contained software engineering problems that push the boundaries of automated coding agents.
Review and evaluate AI training tasks using structured rubrics, provide actionable feedback, and collaborate with contributors in a fast-paced project for a leading AI company.
Evaluate AI-generated investment banking deliverables and define what excellent work looks like by creating evaluation criteria and scoring AI- and human-generated work with rigorous, evidence-based written feedback.
Evaluate AI-generated design work and develop evaluation criteria for real-world design deliverables. This is a flexible, remote contract opportunity for experienced graphic, brand, and visual designers.
Seeking experienced marketing professionals to design evaluation rubrics, assess AI-generated marketing work, and improve frontier AI systems through rigorous, evidence-based evaluation.
Seeking experienced B2B sales professionals to design evaluation rubrics, assess AI-generated enterprise sales work, and improve frontier AI systems through rigorous, evidence-based evaluation.
Seeking experienced accounting professionals to design evaluation rubrics, assess AI-generated accounting work, and improve frontier AI systems through rigorous, evidence-based evaluation.
Seeking experienced Data Scientists to design evaluation rubrics, assess AI-generated data science work, and improve frontier AI systems through rigorous, evidence-based evaluation.
Seeking experienced marketing professionals with AI evaluation experience to design marketing tasks, evaluate LLM outputs, develop scoring rubrics, and improve the marketing reasoning capabilities of frontier AI models.
Seeking experienced insurance professionals with AI evaluation experience to design insurance tasks, evaluate LLM outputs, develop scoring rubrics, and improve the insurance reasoning capabilities of frontier AI models.
Seeking experienced retail professionals with AI evaluation experience to design retail tasks, evaluate LLM outputs, develop scoring rubrics, and improve the retail reasoning capabilities of frontier AI models.
Seeking experienced finance professionals with AI evaluation experience to design finance tasks, evaluate LLM outputs, develop scoring rubrics, and improve the financial reasoning capabilities of frontier AI models.
Collaborate with autonomous AI coding agents, evaluate AI-generated code, and help improve agentic software engineering workflows for next-generation AI systems.
Design realistic business documentation scenarios, evaluate AI-generated Excel, PowerPoint, and Word deliverables, and provide expert feedback to improve next-generation AI systems.
Part-time remote internship for university students to work with AI labs, training and evaluating AI systems, and providing structured feedback.
Evaluate AI-generated legal responses, compare model outputs, and provide expert feedback to improve frontier AI systems' legal reasoning, accuracy, and practical usefulness.
Evaluate AI-generated responses to complex accounting problems, helping improve AI systems' accounting capabilities. Collaborate with leading AI researchers and apply your professional accounting expertise to real-world AI evaluation.
Tips for finding remote jobs
- Set up job alerts on multiple platforms to never miss an opportunity.
- Highlight your remote work experience and self-management skills.
- Prepare for video interviews and remote work assessments.
- Customize your resume and cover letter for each remote position.
- Build a strong online presence on LinkedIn and professional networks.