Machine Learning Evaluation Specialist
Posted 1 Sep 2026
Machine Learning Evaluation Specialist (AI Training) About the Role What if your years of research and domain expertise could directly shape how the next generation of AI models th…
Data and machine learning AI training jobs are the meta-layer of this market: labs hiring ML-literate people to design the evaluations, curate the datasets and audit the reward signals that everyone else’s work feeds into. If the other categories produce training data, this one decides what good data looks like.
The live listings below come from Mercor, Outlier, Alignerr, micro1, DataAnnotation and Terac, refreshed daily, with rates as published.
90 open data & machine learning listings live right now.
Filtering by rate hides per-task listings, which have no hourly equivalent.
Posted 1 Sep 2026
Machine Learning Evaluation Specialist (AI Training) About the Role What if your years of research and domain expertise could directly shape how the next generation of AI models th…

Posted 25 Sep 2026
Role Title: Data Science Specialist - Executive Presentation Review Role Type: Contractor Location: Remote (US, UK or CAN) micro1 is engaging Data Science Specialists to contribute…

Posted 27 Aug 2026
Role Title: Data Scientist Role Type: Contractor Location: Remote About the Role micro1 is partnering with a leading AI lab to hire data scientists and quantitative professionals f…
Posted 27 Feb 2026
About Mercor’s talent network Join our Machine Learning Engineer Expert Network to connect with leading AI labs and companies seeking your expertise. This is an open application fo…

Posted 25 Aug 2026
Job Title: AI/ML Engineer Job Type: Full-Time About Us: micro1 is the end-to-end human data infrastructure behind AGI. Our AI recruiter model is used by frontier AI labs and Fortun…
Posted 13 Jun 2026
About the Opportunity A leading AI research organization is seeking advanced LLM power users to evaluate how well AI systems handle personalized, real-world life tasks. This role i…
Posted 21 Sep 2026
About the role We are hiring expert Evaluators in Data Science to review and assess AI-generated work products (slides, spreadsheets, and documents) for real-world quality. You wil…

Posted 11 Sep 2026
Job Description Pay: $100–$150/hour Location: Global, fully remote Job Type: Contractor (~15 hours per week) Schedule: Flexible—you choose the hours and days you work, including we…
Posted 7 Sep 2026
Computer Vision & ML Expert (AI Training) About the Role What if your expertise in computer vision could directly shape how the next generation of AI systems perceives, interprets,…
Posted 29 Jul 2026
Mercor is building a network of experienced data scientists for potential future projects with leading AI research organizations. These projects may focus on evaluating how effecti…
Use your machine learning expertise to improve how AI reasons Earn up to $150/hr training AI models with Outlier.ai. Put your ML/AI expertise to work evaluating frontier models rem…

Machine Learning Engineer Overview Models are surprisingly bad at reasoning about themselves: training dynamics, evaluation design, data pipelines, and deployment trade-offs. Plaus…

Data Scientist Overview The newest models can spin up a full analysis in minutes — load the data, pick a method, produce charts and a confident conclusion. Someone has to check whe…
Statistics Experts Wanted to Train the Next Generation of AI Use your statistics expertise to train AI models. Work remotely, set your own hours, and earn up to $150/hr. Statistics…

Posted 9 Sep 2026
Role Title: Paid Photo Contributor Opportunity (Siblings) Location: Remote Job Summary micro1 is recruiting sibling pairs or groups to participate in a paid photo data collection p…

Posted 17 Apr 2026
Role Title: Machine Learning Engineer Role Type: Contractor Location: Remote micro1 is engaging Machine Learning Engineers to contribute expertise to a dynamic customer project. In…

Posted 14 Sep 2026
Job Title: Member of Technical Staff, Frontier AI Job Type: Full time Location: Remote The Role We’re hiring a Member of Technical Staff (MTS) to act as a technical owner operating…

Posted 14 Sep 2026
Job Title: Member of Technical Staff, Enterprise AI Job Type: Full-time Location: Remote The Role As a Member of Technical Staff, you will function as a forward-deployed research p…

Posted 16 Sep 2026
Role Title: Statistician Role Type: Contractor Location: Remote micro1 is selecting Statistician to contribute expert knowledge to a customer project focused on advancing data-driv…
Posted 15 Sep 2026
Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. 1. Ov…
Posted 15 Sep 2026
The research teams behind the best-known AI models come to Mercor for data judgment they can't generate on their own. Your standards and expertise become the rubric that frontier A…
Posted 15 Sep 2026
The research teams behind the best-known AI models come to Mercor for machine learning judgment they can't generate on their own. Your standards and expertise become the rubric tha…
Posted 7 Sep 2026
SWE - Machine Learning (Contract) Labelbox • Remote (United States preferred) Shape the data that powers frontier AI --- Quick facts Engagement | Hourly, at‑will contractor Schedul…
Posted 4 Sep 2026
Search Quality Evaluator (AI Training) About the Role What if your everyday curiosity and sharp judgment could directly influence how millions of people find information online? We…
Mercor and micro1 both hire data scientists to grade how AI models handle statistics, experimentation and machine learning, and both pay at the top of this category. The work is evaluation, not modelling. You read a model’s analysis, find the flawed methodology or the misread p-value, and write up why it is wrong.
The two listings are built differently. The Mercor data scientist role is a Talent Network: you apply once, sit its AI interview, and wait to be matched when a lab project opens. The micro1 data scientist role is a named contract with a fixed rate range and a set number of seats, so a passed screen leads to assigned hours rather than a waiting list. A talent network with no live project pays nothing until one starts, so read the listing before you pick.
Both screens want the same profile: a year or more at a top technology, finance or quant firm, working Python and SQL, and experimentation or causal inference you can explain in plain written English. micro1 also asks for a degree from a well-ranked university and a base in an English-speaking country. Kaggle medals count for less than the ability to spot a broken A/B test in someone else’s notebook.
7 open data scientist listings right now, published at $40 to $350 an hour.

Posted 25 Sep 2026

Posted 27 Aug 2026

Posted 29 Jul 2026
Posted 21 Sep 2026
Posted 1 Sep 2026
Typical projects: writing evaluation rubrics and grading guidelines, building benchmark tasks, reviewing preference data for label quality, red-teaming models for failure modes, and analyzing where a fine-tune went wrong. Some roles are hands-on with Python and data tooling; others are pure judgment work on other contributors’ output.
This is the one category where ML knowledge is the product. Everywhere else the labs want domain experts without AI backgrounds; here they want people who know what a held-out set is and why a reward model drifts. Data scientists, ML engineers and quantitative analysts fit the bill.
Six marketplaces publish this work openly. Every listing above links to the one that posted it, and each has its own page here covering pay, screening and who gets hired.
Evaluation design, dataset curation, quality auditing of labeled data, red-teaming, error analysis on model outputs, and reviewer roles that oversee other contributors. Titles vary by marketplace: Outlier calls some of this "quality management", Mercor posts it as evaluation and research roles, Alignerr as expert review. The common thread is judging data and model behavior rather than producing domain content.
Published rates mostly sit between $30 and $80 an hour, above generalist annotation and below scarce-specialist domains like medicine. Reviewer and evaluation-design roles pay more than labeling because they gate everyone else’s output. Per-task pricing is rare here; nearly all listings are hourly.
Working experience with ML systems: data science, ML engineering, analytics or research. A PhD helps for research-adjacent roles but is not the filter; the screens test whether you can spot bad labels, write an unambiguous rubric, and explain why a metric is misleading. Portfolio evidence, a Kaggle history, published analysis or production ML work, moves applications faster than credentials.
Annotation is producing labels; this category is deciding what should be labeled, how, and whether the result is any good. Annotation pays less and scales to more people. Data and ML roles are fewer, better paid, and screened harder, usually with an assessment that hands you messy real data and asks what is wrong with it.
Mercor and micro1 post matched contract roles for data scientists and ML engineers. Outlier and Alignerr run continuous evaluation and reviewer tracks. Terac runs paid expert studies with ML practitioners. The mix shifts weekly, which is why the feed above is worth checking rather than any single board.
It is a standing applicant pool, not an open project. Mercor screens data scientists with its AI interviewer and holds them until a lab needs graders for statistical modelling, A/B test write-ups or ML pipelines. The listing itself says there is no immediate opening, and you are paid only once matched. Apply anyway if you fit the profile: the interview is done once and the match can arrive weeks later. If you want paid hours sooner, the micro1 data scientist contract has a fixed rate and a set number of seats, and the Data Scientist section on this page shows whichever of the two is open.
Every marketplace above screens with an interview or assessment in your own domain before assigning paid work. Run that interview with Skillora first and get feedback on how you explained your reasoning, not just on what you said.
Practise the screen free