SWE-Bench Task Auditor

Posted 28 Aug 2026

$70–90/hr

Apply

Work

Remote · Contract · 40 hrs/wk · Remote — United States

Experience

3+ years

Eligibility

USA

Field

Software engineering

Skills

Software EngineeringPythonJavaGoTypeScriptC++DockerOpen-Source Software MaintenanceCode ReviewBenchmark EvaluationTest HarnessesReference Patch AuditingRepository-Level TestingReward Hacking Detection

About this role

Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric-based written feedback.

Basic Qualifications • 3+ years professional software engineering • Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles) • Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking • Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++)

Preferred Qualifications • Familiarity with SWE-Bench (Verified) or similar repository benchmarks • Maintainer history on major Python OSS (Django, Flask, scikit-learn, sympy, pytest, etc.) • Prior code-review or task-grading experience

Disclosure: Skillora is independent of Mercor. The apply button carries our referral code, and if you are placed, Mercor pays Skillora a fee. It does not change what you are paid, our ranking does not know which listings pay us, and selection is decided by Mercor’s own vetting.

All AI training jobs