Data & Machine Learning AI Training Jobs

Data and machine learning AI training jobs are the meta-layer of this market: labs hiring ML-literate people to design the evaluations, curate the datasets and audit the reward signals that everyone else’s work feeds into. If the other categories produce training data, this one decides what good data looks like.

The live listings below come from Mercor, Outlier, Alignerr, micro1, DataAnnotation and Terac, refreshed daily, with rates as published.

88 open data & machine learning listings live right now.

FiltersActive

What the work looks like

Typical projects: writing evaluation rubrics and grading guidelines, building benchmark tasks, reviewing preference data for label quality, red-teaming models for failure modes, and analyzing where a fine-tune went wrong. Some roles are hands-on with Python and data tooling; others are pure judgment work on other contributors’ output.

This is the one category where ML knowledge is the product. Everywhere else the labs want domain experts without AI backgrounds; here they want people who know what a held-out set is and why a reward model drifts. Data scientists, ML engineers and quantitative analysts fit the bill.

Where these listings come from

Six marketplaces publish this work openly. Every listing above links to the one that posted it, and each has its own page here covering pay, screening and who gets hired.

Data & machine learning AI training jobs, answered

What roles fall under data and ML in AI training work?

Evaluation design, dataset curation, quality auditing of labeled data, red-teaming, error analysis on model outputs, and reviewer roles that oversee other contributors. Titles vary by marketplace: Outlier calls some of this "quality management", Mercor posts it as evaluation and research roles, Alignerr as expert review. The common thread is judging data and model behavior rather than producing domain content.

How much does this work pay?

Published rates mostly sit between $30 and $80 an hour, above generalist annotation and below scarce-specialist domains like medicine. Reviewer and evaluation-design roles pay more than labeling because they gate everyone else’s output. Per-task pricing is rare here; nearly all listings are hourly.

What background gets you hired?

Working experience with ML systems: data science, ML engineering, analytics or research. A PhD helps for research-adjacent roles but is not the filter; the screens test whether you can spot bad labels, write an unambiguous rubric, and explain why a metric is misleading. Portfolio evidence, a Kaggle history, published analysis or production ML work, moves applications faster than credentials.

How is this different from data annotation?

Annotation is producing labels; this category is deciding what should be labeled, how, and whether the result is any good. Annotation pays less and scales to more people. Data and ML roles are fewer, better paid, and screened harder, usually with an assessment that hands you messy real data and asks what is wrong with it.

Which marketplaces post the most data and ML work?

Mercor and micro1 post matched contract roles for data scientists and ML engineers. Outlier and Alignerr run continuous evaluation and reviewer tracks. Terac runs paid expert studies with ML practitioners. The mix shifts weekly, which is why the feed above is worth checking rather than any single board.

The screen is where these roles are won

Every marketplace above screens with an interview or assessment in your own domain before assigning paid work. Run that interview with Skillora first and get feedback on how you explained your reasoning, not just on what you said.

Practise the screen free

Browse other fields