MercorPosted 25 Aug 2026

GPU Kernel Expert

Apply on Mercor

Pay

$70–90/hr

Work

Remote · Contract · 40 hrs/wk · Remote — United States

Experience

3+ years

Eligibility

USA

Field

Software engineering

Skills

GPU Kernel DevelopmentCUDATritonNKIPallasJAXNumerical CorrectnessPerformance ProfilingBenchmarkingNsightNsight ComputeRoofline AnalysisGPU CompilationRuntime DebuggingOperator FusionCompiler EngineeringMLIRIntermediate Representation LoweringMemory Hierarchy OptimizationShared-Memory TilingRegister Pressure OptimizationMemory CoalescingcuBLAScuDNNXLA

About this role

Evaluate the quality, correctness, and completeness of GPU/accelerator kernel development tasks used to train and evaluate a frontier AI lab's models. You'll assess numerical correctness, performance-benchmarking fairness, task scoping, and compilation/runtime validity across diverse kernel task types — and provide clear, rubric-based written feedback.

Basic Qualifications • 3+ years of hands-on experience developing, optimizing, or verifying GPU/accelerator kernels in at least two of: CUDA, Triton, NKI, or Pallas (JAX) • Strong understanding of numerical-correctness criteria for kernels (absolute/relative/ULP tolerances, reference-implementation selection) • Demonstrated experience with performance profiling and benchmarking (nsight, ncu, roofline analysis, or framework-native profilers) • Familiarity with common compilation and runtime failure modes (driver mismatches, OOM, launch-configuration errors, shape/stride mismatches, autotuning failures) • Experience with at least three kernel task types: generation from specification, translation/lowering across frameworks, migration between hardware targets, debugging, performance optimization, or operator fusion

Preferred Qualifications • Experience across both NVIDIA GPU (CUDA/Triton) and custom-accelerator (NKI/Pallas/TPU) ecosystems • Background in compiler engineering, MLIR, or intermediate-representation lowering • Understanding of memory-hierarchy optimization (shared-memory tiling, register pressure, bank conflicts, coalescing patterns) • Contributions to kernel libraries (cuBLAS, cuDNN, Triton community kernels, JAX/XLA custom calls)

Similar AI training jobs

All AI training jobs