AI Research Engineer

CodeRound is hiring for this role.

Bengaluru · Onsite · Full-time · ₹25-35L · 3+ yrs

We’re looking for an AI Research Engineer focused on evaluations to turn ambiguous notions of model capability into rigorous, reproducible measurements. You will own evaluations end-to-end: defining what to measure, building datasets and verifiers, running evaluations at scale, analyzing model behavior, and translating failures into new benchmarks and training signals. You’ll work closely with model researchers throughout the training lifecycle.

What you'll do

Must have

Strong Python programming and software engineering fundamentals. Experience building benchmarks, evaluation systems, automated graders, research infrastructure, or similar testbeds. Strong understanding of LLMs and modern evaluation methodologies. Strong analytical and experimental thinking; you care deeply about whether a metric actually measures the intended capability. Experience debugging complex systems and investigating unexpected experimental results. Strong written and verbal communication.

Good to have

Experience with LLM coding/reasoning evaluations or agent benchmarks. C/C++/Rust and systems programming experience. Experience with QEMU, Renode, Verilator, Spike, or embedded/RTOS environments. Familiarity with formal verification, static analysis, fuzzing, differential testing, or AST-based program transformation. Experience with evaluation frameworks such as SWE-bench, EvalPlus, HumanEval, lm-evaluation-harness, or BigCode. Background in statistics, experimental design, observability, or large-scale ML infrastructure.

Apply for this role on CodeRound AI