AI Research Engineer
CodeRound is hiring for this role.
- AI Research
- Research Engineer
- LLMs
- Benchmarking
- Model Evaluation
- AI Evaluations
- Reasoning Models
- Coding Evaluation
- Agentic AI
- Python
- Software Engineering
- Automated Grading
- Research Infrastructure
- Reinforcement Learning
- SFT
- RL Post-Training
- C++
- Rust
- Embedded Systems
- Formal Verification
- Static Analysis
- Fuzzing
- Differential Testing
- QEMU
- Renode
- Verilator
- Spike
- RTOS
- SWE-bench
- EvalPlus
- HumanEval
- lm-evaluation-harness
- BigCode
- ML Infrastructure
- Experimental Design
- Statistics
We’re looking for an AI Research Engineer focused on evaluations to turn ambiguous notions of model capability into rigorous, reproducible measurements. You will own evaluations end-to-end: defining what to measure, building datasets and verifiers, running evaluations at scale, analyzing model behavior, and translating failures into new benchmarks and training signals. You’ll work closely with model researchers throughout the training lifecycle.
What you'll do
- Design rigorous benchmarks for reasoning, coding, and agentic capabilities across C/C++, Rust, embedded systems, hardware specifications, and engineering tasks.
- Develop deterministic evaluators using compilers, emulators, simulators, static analysis, formal verification, and hardware-in-the-loop systems.
- Build distributed infrastructure to run evaluations reliably against models and live training checkpoints.
- Diagnose regressions and anomalous results, separating model failures from issues in prompts, data, evaluators, or infrastructure.
- Design robust metrics, difficulty curricula, contamination-resistant tasks, and experiments around prompting, sampling, and scaffolding.
- Turn evaluation failures into targeted datasets, reward signals, and new training objectives in collaboration with RL researchers.
- Create dashboards, experiment tracking, and evaluation libraries that make model performance easy to understand and trust.
Must have
Strong Python programming and software engineering fundamentals. Experience building benchmarks, evaluation systems, automated graders, research infrastructure, or similar testbeds. Strong understanding of LLMs and modern evaluation methodologies. Strong analytical and experimental thinking; you care deeply about whether a metric actually measures the intended capability. Experience debugging complex systems and investigating unexpected experimental results. Strong written and verbal communication.
Good to have
Experience with LLM coding/reasoning evaluations or agent benchmarks. C/C++/Rust and systems programming experience. Experience with QEMU, Renode, Verilator, Spike, or embedded/RTOS environments. Familiarity with formal verification, static analysis, fuzzing, differential testing, or AST-based program transformation. Experience with evaluation frameworks such as SWE-bench, EvalPlus, HumanEval, lm-evaluation-harness, or BigCode. Background in statistics, experimental design, observability, or large-scale ML infrastructure.