Oerview Mercor is seeking Generalist Experts to evaluate AI-generated responses and provide structured written feedback that helps improve frontier AI systems. This role is ideal for analytical thinkers with strong reasoning and communication skills who can assess complex…
Remote job listings
Find your next remote opportunity from thousands of listings across the globe.
Please submit your resume in English and indicate your level of English. Mindrift is looking for a highly skilled Web Designer to join the Tendem project ( https://tendem.ai/ ) and drive specialized data scraping workflows within our hybrid AI + human system.
Evaluate the safety, quality, and alignment of frontier AI models by reviewing policy-sensitive content, applying AI safety standards, identifying model failures, and helping improve next-generation AI systems.
AI Ethics Reviewer About the Role What if your sense of fairness and critical thinking could directly influence how AI behaves for millions of people?
Please submit your resume in English and indicate your level of English. Mindrift is looking for a highly skilled Brand Designer to join the Tendem project ( https://tendem.ai/ ) and drive specialized data scraping workflows within our hybrid AI + human system.
Datavant is the data collaboration platform trusted for healthcare. Guided by our mission to make the world’s health data secure, accessible and actionable, we provide critical data solutions for organizations across the healthcare ecosystem - including providers, health plans,…
Help train frontier AI models by reviewing, evaluating, and annotating legal AI outputs using professional legal expertise and structured evaluation guidelines.
Help train frontier AI models by reviewing, evaluating, and annotating AI-generated content using academic research expertise and structured evaluation guidelines.
Review and evaluate AI training tasks using structured rubrics, provide actionable feedback, and collaborate with contributors in a fast-paced project for a leading AI company.
Seeking experienced marketing professionals to design evaluation rubrics, assess AI-generated marketing work, and improve frontier AI systems through rigorous, evidence-based evaluation.
Seeking experienced B2B sales professionals to design evaluation rubrics, assess AI-generated enterprise sales work, and improve frontier AI systems through rigorous, evidence-based evaluation.
Seeking experienced accounting professionals to design evaluation rubrics, assess AI-generated accounting work, and improve frontier AI systems through rigorous, evidence-based evaluation.
Seeking experienced Data Scientists to design evaluation rubrics, assess AI-generated data science work, and improve frontier AI systems through rigorous, evidence-based evaluation.
Seeking experienced marketing professionals with AI evaluation experience to design marketing tasks, evaluate LLM outputs, develop scoring rubrics, and improve the marketing reasoning capabilities of frontier AI models.
Seeking experienced insurance professionals with AI evaluation experience to design insurance tasks, evaluate LLM outputs, develop scoring rubrics, and improve the insurance reasoning capabilities of frontier AI models.
Seeking experienced retail professionals with AI evaluation experience to design retail tasks, evaluate LLM outputs, develop scoring rubrics, and improve the retail reasoning capabilities of frontier AI models.
Role Title: Generalist for Data Annotator Role Type: Contractor Location: Remote micro1 is engaging Generalists for Data Annotation to support a customer's project focused on building high-quality datasets.
Part-time remote internship for university students to work with AI labs, training and evaluating AI systems, and providing structured feedback.
Help train frontier AI systems by creating legal reasoning tasks and evaluating AI-generated legal analyses. Create realistic legal scenarios and develop high-quality reference answers and grading rubrics.
Seeking fluent Russian and English speakers in the San Francisco Bay Area to transcribe, annotate, and evaluate Russian audio content for training and benchmarking advanced AI language models. This short-term contract offers flexible hours and requires US work authorization.
Seeking fluent Hindi and English speakers in the San Francisco Bay Area to transcribe, annotate, and evaluate Hindi audio content for training and benchmarking advanced AI language models. This short-term contract offers flexible hours and requires US work authorization.
Seeking experienced musicians, producers, composers, and musicologists to evaluate AI-generated music, annotate musical content, and help improve frontier generative music AI systems. This part-time, remote role involves structured evaluation, captioning, lyric alignment, and quality assessment to enhance AI understanding of music across genres and languages.
Seeking Juris Doctor (JD) professionals to create legal reasoning challenges and evaluate AI-generated legal analysis for cutting-edge AI research projects. Contribute to cutting-edge AI research while applying your legal knowledge to challenging work.
We are offering a project-based opportunity for independent contributors who are curious about artificial intelligence and excited to learn how advanced AI models are built, evaluated, and improved.
About micro1 micro1 is a leading AI data lab that builds training datasets and evaluation systems for frontier AI models. The company works with domain experts across multiple fields to generate high-quality human feedback, annotations, and evaluation signals that improve model…
AI Generalist (No Experience Required) We are offering a project-based opportunity for independent contributors who are curious about artificial intelligence and excited to learn how advanced AI models are built, evaluated, and improved.
Mercor is seeking Phylogenetics PhDs and Domain Experts for a premier project with one of the world's leading AI laboratories. In this role, you will apply advanced subject matter expertise to help develop high-quality datasets used to train and evaluate state-of-the-art large…
Role Overview Mercor is seeking a Basque Audio Generalist Evaluator Expert to contribute to a high-impact audio AI research project with a leading research lab. In this role, you will work on transcription, annotation, evaluation, and quality assurance tasks that help train and…
Mercor is seeking an English (Australia) Audio Generalist Evaluator Expert to contribute to a high-impact AI audio research project with a leading research laboratory. In this role, you will work on transcription, annotation, evaluation, and quality assurance tasks that help…
About the Role Mercor is seeking an English (New Zealand) Audio Generalist Evaluator Expert to contribute to a high-impact AI audio research project with a leading research laboratory. In this role, you will work on transcription, annotation, evaluation, and quality assurance…
Mercor is seeking a Spanish (Mexico) Audio Generalist Evaluator Expert to contribute to a high-impact AI audio research project with a leading research laboratory. In this role, you will work on transcription, annotation, evaluation, and quality assurance tasks that help train…
Mercor is seeking a Spanish (Spain) Audio Generalist Evaluator Expert to contribute to a high-impact AI audio research project with a leading research laboratory. In this role, you will work on transcription, annotation, evaluation, and quality assurance tasks that help train…
About the Role A leading AI research organization is seeking experienced AI power users to evaluate how effectively advanced language models perform real-world, highly personalized life-assistance tasks. Contributors will assess AI outputs across domains such as productivity,…
At Underdog, we make sports more fun. Our thesis is simple: build the best products and we’ll build the biggest company in the space, because there’s so much more to be built for sports fans.
Mercor is seeking Dutch Audio Generalist Evaluator Experts to contribute to a high-impact AI research project focused on audio understanding, transcription, annotation, and model evaluation. Contributors will help train and benchmark advanced language models by converting audio…
<h2>About Fleet AI</h2> <p>Fleet AI is an applied artificial intelligence company focused on human–AI collaboration at scale.</p> <p>Our goal is to accelerate the transition to the allocation economy—a future where humans direct work rather than perform it. This capability has…
About the Role Mercor is seeking expert evaluators with deep experience in humanities, arts, and cultural studies to assess AI-generated work products for accuracy, rigor, and domain quality. This is a remote, hourly contract opportunity focused on reviewing documents,…
As an Accountant Expert, you will apply your accounting and financial analysis expertise to help train and evaluate next-generation AI systems. Your work will support the development of AI models through accounting-based reasoning, data analysis, rubric creation, and…
As a Consultant Expert, you will apply your professional expertise to help train and evaluate next-generation AI systems. Your work will involve reviewing, creating, and assessing domain-specific content while helping improve AI performance through expert feedback and…
Mercor is seeking domain experts in Mechanical Engineering and Materials Science to evaluate simulated research environments and AI agent trajectories. Contributors will review AI-generated research workflows and outputs, assessing performance across multiple dimensions…
Mercor is seeking a Japanese Audio Generalist Evaluator Expert to contribute to a high-impact audio AI research project with a leading research lab. This role focuses on transcription, annotation, evaluation, benchmarking, and quality assurance tasks that help train and assess…
Mercor is seeking an English (US) Audio Generalist Evaluator Expert to contribute to a high-impact audio AI research project with a leading research lab. This role focuses on transcription, annotation, evaluation, benchmarking, and quality assurance tasks that help train and…
DataAnnotation is seeking Secondary Education Teachers to help train and evaluate AI models. In this role, contributors will assess AI-generated responses, evaluate reasoning quality, and complete writing and editing tasks that help improve AI performance and accuracy.…
About Project Axiom — Track 01: Axis Sanity Check Mercor is recruiting clinicians and health-content researchers to evaluate and stress-test a four-axis taxonomy used for classifying health claims appearing across social media and podcast content. This project focuses on…
Are you an English language expert eager to shape the future of AI? Large‑scale language models are evolving from clever chatbots into powerful engines of linguistic discovery. With high‑quality training data, tomorrow’s AI can democratize world‑class education, keep pace with…
micro1 is seeking Biology Experts to contribute domain expertise toward training and improving next-generation AI systems. This role focuses on helping AI models learn biological reasoning, scientific interpretation, analytical problem-solving, and research comprehension…
micro1 is seeking Physics Experts to contribute domain expertise toward training and improving next-generation AI systems. This role focuses on helping AI models learn scientific reasoning, physics concepts, analytical problem-solving, and technical interpretation through…
Mercor is seeking experienced arts and culture program leaders, humanities educators, and curators to contribute expertise toward frontier AI training and evaluation projects focused on humanities, arts, and cultural reasoning. This role focuses on supporting AI systems…
Tips for finding remote jobs
- Set up job alerts on multiple platforms to never miss an opportunity.
- Highlight your remote work experience and self-management skills.
- Prepare for video interviews and remote work assessments.
- Customize your resume and cover letter for each remote position.
- Build a strong online presence on LinkedIn and professional networks.