Senior Software Engineer - Pairwise Evaluation
Build the AI infrastructure millions rely on — senior-level engineering, fully remote, and genuinely impactful work....
Labelbox’s expert network: per-project listings with published rates, weighted toward coding and language work.
Operated by
Labelbox
Work type
Per-project applications
Screening
AI interview + qualification tasks
Payouts
Weekly cycle
486 open Alignerr listings in the live feed, refreshed daily. Apply links go to Alignerr directly; where a referral programme exists, Skillora may earn a fee at no cost to you, and it never affects what is shown. Official site: alignerr.com
Filtering by rate hides per-task listings, which have no hourly equivalent.
Build the AI infrastructure millions rely on — senior-level engineering, fully remote, and genuinely impactful work....
Get paid to break AI. Use your security expertise to probe, test, and harden cutting-edge AI systems — fully remote and flexible....
Turn your social media instincts and people skills into real impact: grow and energize a thriving online community. Fully remote and flexible....
Turn your ear for audio into AI training gold — transcribe and refine English audio from anywhere. Flexible, remote, start quickly....
Put your investment expertise to work shaping the future of AI. Evaluate portfolio strategies and risk frameworks — remote, flexible, and well-paid....
Put your finance expertise to work shaping the future of AI. Fully remote, flexible contract — up to $95/hr....
Design AI agent training tasks that test end-to-end support triage, from ticket to fix-or-escalate decision. ...
Design AI agent training tasks that test accounting reconciliation with numerically verifiable answers....
Design AI agent training tasks that test data-grounded marketing work: announcements, competitive comparisons, and campaign attribution.You said: ...
Build the systems that power AI — write clean, scalable back end code from anywhere. Remote, flexible, and cutting-edge work....
Train the robots of tomorrow. Bring your ML and MuJoCo expertise to cutting-edge AI simulation projects — fully remote and flexible....
Design, build, test, deploy, and maintain AI software. Develop backend, frontend, APIs, databases, and cloud systems. Debug, optimize, review code, an...
Design, build, test, deploy, and maintain AI software. Develop backend, frontend, APIs, databases, and cloud systems. Debug, optimize, review code, an...
Get paid to shape how AI writes Lua code. Put your Roblox and scripting expertise to work on cutting-edge AI projects — fully remote and flexible....
Build the systems that measure AI intelligence. Senior engineers — shape how we evaluate cutting-edge AI, fully remote and flexible....
Design AI agent training tasks that test real engineering work: log debugging, bug tracing, and live code fixes. ...
Use your finance expertise and Python skills to teach AI how real analysts think. Fully remote, flexible, and meaningful....
Design AI agent training tasks that test multi-system revenue reconciliation across CRM, orders, and email....
Seeking generalist Math Reasoning Experts to create and refine high‑quality questions in causal inference and experimental design, ensuring rigor and ...
Put your PhD in mathematics to work shaping frontier AI. Remote, flexible, and highly paid — your expertise has never mattered more....
Seeking generalist Physics Reasoning Experts to create and refine high-quality questions across fundamental and applied physics topics, ensuring rigor...
Put your physics PhD to work shaping the future of AI. Help frontier models think like a scientist — fully remote, flexible, and well paid....
Seeking generalist Biology Reasoning Experts to create and refine high-quality questions across a broad range of biological topics, ensuring rigor and...
Seeking PhD-level Biologists to develop, review, and refine advanced questions across specialized areas of biology, ensuring the highest standards of ...
Alignerr is Labelbox’s network for AI training work. Labelbox sells data infrastructure to AI labs, and Alignerr is how it recruits the humans those labs need: preference labeling, response evaluation, and expert reference answers across coding, language, STEM and professional domains.
It runs as a job board. Each project is its own listing with its own published rate, you apply to specific projects, and accepted work happens inside the Labelbox platform. That per-project structure is why Alignerr usually has one of the largest catalogues in this feed.
Pay runs on a weekly cycle. Headline expert rates reach $150 an hour; most contributors work standard tiers well below that, and both ends are visible on the listings.
$10–200
per hour, published range
$60/hr
median listing
100
live listings publishing a rate
Computed from Alignerr’s own published rates on the live feed, not from survey estimates.
Every Alignerr listing publishes a rate, usually as a range. Generalist projects post in the teens to twenties per hour, expert projects from the $30s up, with advertised specialist ceilings around $150. Pay periods run Monday to Sunday, paid out the following week.
The catalogue leans hard into software engineering and language work, with expert tracks in STEM, medicine and law. The bar sits between DataAnnotation and Mercor: a professional or academic background in your subject helps and gates the expert tiers, but generalist projects exist and accept strong writers.
Because applications are per project, persistence works differently here: not getting one project means little, and contributors typically apply across several listings until one converts.
Alignerr screens with a recorded AI interview plus qualification tasks per project. The interview covers your background and domain reasoning; the qualification tasks are short samples of the actual work, graded against the project rubric before you get paid access.
The two failure modes are familiar: talking in generalities during the interview instead of reasoning through specifics, and rushing qualification tasks that are graded on exactness. Both are practisable, and the interview format is the same one the rest of this market uses.
Skillora runs the same format: a domain interview with an AI interviewer, followed by feedback on how you explained your reasoning. Take it as many times as you want before the one that counts.
Start a free practice interviewYes. Alignerr is operated by Labelbox, an established AI data company, and pays contributors on a weekly cycle for project work. As across this market, the real risks are operational: projects fill or pause quickly, qualification work is unpaid, and headline rates apply to a small expert tier rather than to most contributors.
Per project, with the rate on every listing. Generalist work posts in the high teens to twenties per hour; domain-expert projects post from the $30s upward; advertised specialist ceilings reach $150 an hour for scarce credentials. Community-tracked averages land in the $30s. The live cards above are the current truth.
A recorded conversation with an AI interviewer about your background and how you reason in your field, followed by project-specific qualification tasks. It is closer to Mercor’s screen than to a quiz: it scores explanation quality. Answer with specifics from real work, structure your reasoning aloud, and treat the qualification tasks as graded samples, because they are.
You work inside Labelbox’s platform on the project you applied to: labeling, ranking, writing or reviewing, against a rubric with reviewer feedback. Hours are flexible within project deadlines. When a project ends, you apply to the next one; tenure on one project makes the next application easier.
Moderately. It is easier to get into than Mercor, which filters on senior credentials, and harder than DataAnnotation, which has a single generalist gate. Expert tiers check degrees and professional history; generalist projects mostly check writing and rubric discipline. Applying to several projects at once is normal and improves odds.
Yes, like everywhere in this market. Per-project hiring means supply tracks lab demand: a strong month can be followed by a quiet one. The mitigation is the same as on Outlier: hold qualifications on multiple projects and platforms, and treat any single marketplace as one feed among several.