Machine Learning Evaluation Specialist
Posted 1 Sep 2026
Machine Learning Evaluation Specialist (AI Training) About the Role What if your years of research and domain expertise could directly shape how the next generation of AI models th…
Data and machine learning AI training jobs are the meta-layer of this market: labs hiring ML-literate people to design the evaluations, curate the datasets and audit the reward signals that everyone else’s work feeds into. If the other categories produce training data, this one decides what good data looks like.
The live listings below come from Mercor, Outlier, Alignerr, micro1, DataAnnotation and Terac, refreshed daily, with rates as published.
89 open data & machine learning listings live right now.
Filtering by rate hides per-task listings, which have no hourly equivalent.
Posted 1 Sep 2026
Machine Learning Evaluation Specialist (AI Training) About the Role What if your years of research and domain expertise could directly shape how the next generation of AI models th…

Posted 14 Jul 2026
Job Title: AI/ML Engineer, Internal Platforms Job Type: Full-time Location: Remote The Role Join our team to build and improve a next-generation AI recruiting agent. You'll develop…

Posted 27 Aug 2026
Role Title: Data Scientist Role Type: Contractor Location: Remote About the Role micro1 is partnering with a leading AI lab to hire data scientists and quantitative professionals f…
Posted 27 Feb 2026
About Mercor’s talent network Join our Machine Learning Engineer Expert Network to connect with leading AI labs and companies seeking your expertise. This is an open application fo…
Posted 4 Sep 2026
Mercor is recruiting UK-Based Data Engineering experts for a short, intensive project with a globally leading AI lab. What the work involves You will produce the kind of profession…

Posted 25 Aug 2026
Job Title: AI/ML Engineer Job Type: Full-Time About Us: micro1 is the end-to-end human data infrastructure behind AGI. Our AI recruiter model is used by frontier AI labs and Fortun…

Posted 4 Aug 2026
Role Title: AI Data Science Domain Expert Role Type: Contractor (Part Time) Location: Remote micro1 is engaging AI Data Science Domain Experts to participate in a customer's projec…

Posted 3 Aug 2026
Role Title: AI Domain Expert Role Type: Contractor (Part Time) Location: Remote micro1 is engaging AI Domain Experts to contribute to a high-impact project that advances artificial…
Posted 13 Jun 2026
About the Opportunity A leading AI research organization is seeking advanced LLM power users to evaluate how well AI systems handle personalized, real-world life tasks. This role i…

Posted 30 Jul 2026
Role Title: AI trainer Role Type: Contractor Location: Remote micro1 is engaging AI trainers to contribute to a customer's project focused on advancing artificial intelligence capa…
Posted 30 Apr 2026
Overview We are hiring Analyst / Insights professionals with experience in analytics, research, or insights roles. In this role, you will review, assess, and provide structured fee…

Posted 11 Sep 2026
Job Description Pay: $100–$150/hour Location: Global, fully remote Job Type: Contractor (~15 hours per week) Schedule: Flexible—you choose the hours and days you work, including we…
Posted 7 Sep 2026
Computer Vision & ML Expert (AI Training) About the Role What if your expertise in computer vision could directly shape how the next generation of AI systems perceives, interprets,…
Posted 29 Jul 2026
Mercor is building a network of experienced data scientists for potential future projects with leading AI research organizations. These projects may focus on evaluating how effecti…
Use your machine learning expertise to improve how AI reasons Earn up to $150/hr training AI models with Outlier.ai. Put your ML/AI expertise to work evaluating frontier models rem…

Machine Learning Engineer Overview Models are surprisingly bad at reasoning about themselves: training dynamics, evaluation design, data pipelines, and deployment trade-offs. Plaus…

Data Scientist Overview The newest models can spin up a full analysis in minutes — load the data, pick a method, produce charts and a confident conclusion. Someone has to check whe…
Statistics Experts Wanted to Train the Next Generation of AI Use your statistics expertise to train AI models. Work remotely, set your own hours, and earn up to $150/hr. Statistics…
Your ML expertise shapes how frontier AI models think — up to $150/hr Earn up to $150/hr training AI as an ML/AI PhD on Outlier. Evaluate model outputs, craft domain questions. Rem…
Computer Science experts — Earn up to $150/hr improving frontier AI models. Fully remote. Earn up to $150/hr training AI as a CS PhD on Outlier. Evaluate reasoning, craft domain qu…

Posted 9 Sep 2026
Role Title: AI Facial Data Collection Associate (Siblings) Role Type: Contractor Location: Remote micro1 is engaging AI Facial Data Collection Associates to contribute to an impact…

Posted 15 Aug 2026
Role Title: AI Facial Data Collection Contributor Location: Remote Job Summary Get paid from home in this fully remote AI training project by providing a few selfies and short vide…

Posted 6 Aug 2026
Role Title: AI Facial Data Collection Contributor Location: Remote Job Summary Get paid from home in this fully remote AI training project by providing a few selfies and short vide…

Posted 17 Apr 2026
Role Title: Machine Learning Engineer Role Type: Contractor Location: Remote micro1 is engaging Machine Learning Engineers to contribute expertise to a dynamic customer project. In…
Mercor and micro1 both hire data scientists to grade how AI models handle statistics, experimentation and machine learning, and both pay at the top of this category. The work is evaluation, not modelling. You read a model’s analysis, find the flawed methodology or the misread p-value, and write up why it is wrong.
The two listings are built differently. The Mercor data scientist role is a Talent Network: you apply once, sit its AI interview, and wait to be matched when a lab project opens. The micro1 data scientist role is a named contract with a fixed rate range and a set number of seats, so a passed screen leads to assigned hours rather than a waiting list. A talent network with no live project pays nothing until one starts, so read the listing before you pick.
Both screens want the same profile: a year or more at a top technology, finance or quant firm, working Python and SQL, and experimentation or causal inference you can explain in plain written English. micro1 also asks for a degree from a well-ranked university and a base in an English-speaking country. Kaggle medals count for less than the ability to spot a broken A/B test in someone else’s notebook.
6 open data scientist listings right now, published at $40 to $280 an hour.

Posted 27 Aug 2026

Posted 4 Aug 2026

Posted 29 Jul 2026
Posted 1 Sep 2026
Posted 4 Sep 2026
Typical projects: writing evaluation rubrics and grading guidelines, building benchmark tasks, reviewing preference data for label quality, red-teaming models for failure modes, and analyzing where a fine-tune went wrong. Some roles are hands-on with Python and data tooling; others are pure judgment work on other contributors’ output.
This is the one category where ML knowledge is the product. Everywhere else the labs want domain experts without AI backgrounds; here they want people who know what a held-out set is and why a reward model drifts. Data scientists, ML engineers and quantitative analysts fit the bill.
Six marketplaces publish this work openly. Every listing above links to the one that posted it, and each has its own page here covering pay, screening and who gets hired.
Evaluation design, dataset curation, quality auditing of labeled data, red-teaming, error analysis on model outputs, and reviewer roles that oversee other contributors. Titles vary by marketplace: Outlier calls some of this "quality management", Mercor posts it as evaluation and research roles, Alignerr as expert review. The common thread is judging data and model behavior rather than producing domain content.
Published rates mostly sit between $30 and $80 an hour, above generalist annotation and below scarce-specialist domains like medicine. Reviewer and evaluation-design roles pay more than labeling because they gate everyone else’s output. Per-task pricing is rare here; nearly all listings are hourly.
Working experience with ML systems: data science, ML engineering, analytics or research. A PhD helps for research-adjacent roles but is not the filter; the screens test whether you can spot bad labels, write an unambiguous rubric, and explain why a metric is misleading. Portfolio evidence, a Kaggle history, published analysis or production ML work, moves applications faster than credentials.
Annotation is producing labels; this category is deciding what should be labeled, how, and whether the result is any good. Annotation pays less and scales to more people. Data and ML roles are fewer, better paid, and screened harder, usually with an assessment that hands you messy real data and asks what is wrong with it.
Mercor and micro1 post matched contract roles for data scientists and ML engineers. Outlier and Alignerr run continuous evaluation and reviewer tracks. Terac runs paid expert studies with ML practitioners. The mix shifts weekly, which is why the feed above is worth checking rather than any single board.
It is a standing applicant pool, not an open project. Mercor screens data scientists with its AI interviewer and holds them until a lab needs graders for statistical modelling, A/B test write-ups or ML pipelines. The listing itself says there is no immediate opening, and you are paid only once matched. Apply anyway if you fit the profile: the interview is done once and the match can arrive weeks later. If you want paid hours sooner, the micro1 data scientist contract has a fixed rate and a set number of seats, and the Data Scientist section on this page shows whichever of the two is open.
Every marketplace above screens with an interview or assessment in your own domain before assigning paid work. Run that interview with Skillora first and get feedback on how you explained your reasoning, not just on what you said.
Practise the screen free