Data & Machine Learning AI Training Jobs

Data and machine learning AI training jobs are the meta-layer of this market: labs hiring ML-literate people to design the evaluations, curate the datasets and audit the reward signals that everyone else’s work feeds into. If the other categories produce training data, this one decides what good data looks like.

The live listings below come from Mercor, Outlier, Alignerr, micro1, DataAnnotation and Terac, refreshed daily, with rates as published.

89 open data & machine learning listings live right now.

FiltersActive

Data Scientist AI Training Jobs

Mercor and micro1 both hire data scientists to grade how AI models handle statistics, experimentation and machine learning, and both pay at the top of this category. The work is evaluation, not modelling. You read a model’s analysis, find the flawed methodology or the misread p-value, and write up why it is wrong.

The two listings are built differently. The Mercor data scientist role is a Talent Network: you apply once, sit its AI interview, and wait to be matched when a lab project opens. The micro1 data scientist role is a named contract with a fixed rate range and a set number of seats, so a passed screen leads to assigned hours rather than a waiting list. A talent network with no live project pays nothing until one starts, so read the listing before you pick.

Both screens want the same profile: a year or more at a top technology, finance or quant firm, working Python and SQL, and experimentation or causal inference you can explain in plain written English. micro1 also asks for a degree from a well-ranked university and a base in an English-speaking country. Kaggle medals count for less than the ability to spot a broken A/B test in someone else’s notebook.

6 open data scientist listings right now, published at $40 to $280 an hour.

Search all data scientist listings

What the work looks like

Typical projects: writing evaluation rubrics and grading guidelines, building benchmark tasks, reviewing preference data for label quality, red-teaming models for failure modes, and analyzing where a fine-tune went wrong. Some roles are hands-on with Python and data tooling; others are pure judgment work on other contributors’ output.

This is the one category where ML knowledge is the product. Everywhere else the labs want domain experts without AI backgrounds; here they want people who know what a held-out set is and why a reward model drifts. Data scientists, ML engineers and quantitative analysts fit the bill.

Where these listings come from

Six marketplaces publish this work openly. Every listing above links to the one that posted it, and each has its own page here covering pay, screening and who gets hired.

Data & machine learning AI training jobs, answered

What roles fall under data and ML in AI training work?

Evaluation design, dataset curation, quality auditing of labeled data, red-teaming, error analysis on model outputs, and reviewer roles that oversee other contributors. Titles vary by marketplace: Outlier calls some of this "quality management", Mercor posts it as evaluation and research roles, Alignerr as expert review. The common thread is judging data and model behavior rather than producing domain content.

How much does this work pay?

Published rates mostly sit between $30 and $80 an hour, above generalist annotation and below scarce-specialist domains like medicine. Reviewer and evaluation-design roles pay more than labeling because they gate everyone else’s output. Per-task pricing is rare here; nearly all listings are hourly.

What background gets you hired?

Working experience with ML systems: data science, ML engineering, analytics or research. A PhD helps for research-adjacent roles but is not the filter; the screens test whether you can spot bad labels, write an unambiguous rubric, and explain why a metric is misleading. Portfolio evidence, a Kaggle history, published analysis or production ML work, moves applications faster than credentials.

How is this different from data annotation?

Annotation is producing labels; this category is deciding what should be labeled, how, and whether the result is any good. Annotation pays less and scales to more people. Data and ML roles are fewer, better paid, and screened harder, usually with an assessment that hands you messy real data and asks what is wrong with it.

Which marketplaces post the most data and ML work?

Mercor and micro1 post matched contract roles for data scientists and ML engineers. Outlier and Alignerr run continuous evaluation and reviewer tracks. Terac runs paid expert studies with ML practitioners. The mix shifts weekly, which is why the feed above is worth checking rather than any single board.

Is the Mercor data scientist talent network a real job?

It is a standing applicant pool, not an open project. Mercor screens data scientists with its AI interviewer and holds them until a lab needs graders for statistical modelling, A/B test write-ups or ML pipelines. The listing itself says there is no immediate opening, and you are paid only once matched. Apply anyway if you fit the profile: the interview is done once and the match can arrive weeks later. If you want paid hours sooner, the micro1 data scientist contract has a fixed rate and a set number of seats, and the Data Scientist section on this page shows whichever of the two is open.

The screen is where these roles are won

Every marketplace above screens with an interview or assessment in your own domain before assigning paid work. Run that interview with Skillora first and get feedback on how you explained your reasoning, not just on what you said.

Practise the screen free

Browse other fields