Senior Software Engineer – AI Evaluation
Build the systems that measure AI intelligence. Senior engineers — shape how we evaluate cutting-edge AI, fully remote and flexible....
Data and machine learning AI training jobs are the meta-layer of this market: labs hiring ML-literate people to design the evaluations, curate the datasets and audit the reward signals that everyone else’s work feeds into. If the other categories produce training data, this one decides what good data looks like.
The live listings below come from Mercor, Outlier, Alignerr, micro1, DataAnnotation and Terac, refreshed daily, with rates as published.
88 open data & machine learning listings live right now.
Filtering by rate hides per-task listings, which have no hourly equivalent.
Build the systems that measure AI intelligence. Senior engineers — shape how we evaluate cutting-edge AI, fully remote and flexible....
Seeking backend development experts to create, review, and refine high-quality UI coding problems and solutions, ensuring rigor and accuracy in evalua...
Put your statistics expertise to work shaping frontier AI. Flexible, remote, and well-compensated for experts in probability, inference, and experimen...
Turn your forecasting instincts into impact. Help shape smarter AI by predicting real-world events with structured, probabilistic reasoning....
Turn complex clinical data into decisions that improve patient outcomes. Bring your healthcare analytics expertise to a fully remote, high-impact cont...
Get paid $50–$75/hr to build AI infrastructure that powers the world's leading models. Senior Python engineers wanted — fully remote and high-impact....
Get paid $50–$75/hr to build AI data pipelines at the frontier. Senior Python engineers wanted — fully remote, flexible contract....
Get paid $50–$75/hr to build AI infrastructure that matters — Python engineers wanted for real-world pipelines at leading AI labs. Fully remote....
Put your PhD or Masters to work shaping the future of AI — solve complex, real-world problems for the world's leading research labs. Remote and flexib...
Get paid $60–$80/hr to teach AI how to think. Senior ML experts wanted — shape the reasoning behind next-gen models, fully remote....
Get paid to teach AI how to think — use your ML expertise to shape the reasoning behind next-gen models. Remote, flexible, and highly competitive pay....
Get paid $60–$80/hr to teach AI how to think. Use your ML expertise to shape how the world's most advanced models reason and make decisions....
Get paid to do chores — and help build the robots that will eventually do them for you. Flexible, remote-friendly AI training role....
Put your data science expertise to work training the world's most advanced AI — fully remote, flexible, and paying up to $80/hr....
Get paid up to $80/hr to challenge and improve cutting-edge AI — put your data science expertise to work shaping the future of machine learning....
Put your econometrics expertise to work shaping the future of AI. Flexible, remote contract with competitive pay....
Get paid to shape the future of AI — no experience needed. Flexible remote contract, up to $120/hr....
Get paid to shape how AI talks about products — evaluate reviews, catch bad content, and help build better AI. Fully remote and flexible....
Get paid to shape how AI sees the world — describe images from anywhere. No experience needed, flexible hours....
Get paid to shape the future of AI shopping — review and classify product data from anywhere. No e-commerce experience required....
Turn your sharp eye into AI impact. Describe images that teach machines to see — flexible, remote, no experience needed....
Get paid to challenge AI — craft clever prompts that push the limits of machine reasoning. Fully remote, no tech background needed....
Get paid to shape the future of AI — label, tag, and annotate data from anywhere. No experience needed, flexible hours....
Get paid to put your local knowledge to work — help shape how AI understands the real world, one location at a time. Fully remote....
Typical projects: writing evaluation rubrics and grading guidelines, building benchmark tasks, reviewing preference data for label quality, red-teaming models for failure modes, and analyzing where a fine-tune went wrong. Some roles are hands-on with Python and data tooling; others are pure judgment work on other contributors’ output.
This is the one category where ML knowledge is the product. Everywhere else the labs want domain experts without AI backgrounds; here they want people who know what a held-out set is and why a reward model drifts. Data scientists, ML engineers and quantitative analysts fit the bill.
Six marketplaces publish this work openly. Every listing above links to the one that posted it, and each has its own page here covering pay, screening and who gets hired.
Evaluation design, dataset curation, quality auditing of labeled data, red-teaming, error analysis on model outputs, and reviewer roles that oversee other contributors. Titles vary by marketplace: Outlier calls some of this "quality management", Mercor posts it as evaluation and research roles, Alignerr as expert review. The common thread is judging data and model behavior rather than producing domain content.
Published rates mostly sit between $30 and $80 an hour, above generalist annotation and below scarce-specialist domains like medicine. Reviewer and evaluation-design roles pay more than labeling because they gate everyone else’s output. Per-task pricing is rare here; nearly all listings are hourly.
Working experience with ML systems: data science, ML engineering, analytics or research. A PhD helps for research-adjacent roles but is not the filter; the screens test whether you can spot bad labels, write an unambiguous rubric, and explain why a metric is misleading. Portfolio evidence, a Kaggle history, published analysis or production ML work, moves applications faster than credentials.
Annotation is producing labels; this category is deciding what should be labeled, how, and whether the result is any good. Annotation pays less and scales to more people. Data and ML roles are fewer, better paid, and screened harder, usually with an assessment that hands you messy real data and asks what is wrong with it.
Mercor and micro1 post matched contract roles for data scientists and ML engineers. Outlier and Alignerr run continuous evaluation and reviewer tracks. Terac runs paid expert studies with ML practitioners. The mix shifts weekly, which is why the feed above is worth checking rather than any single board.
Every marketplace above screens with an interview or assessment in your own domain before assigning paid work. Run that interview with Skillora first and get feedback on how you explained your reasoning, not just on what you said.
Practise the screen free