ML Research Engineer
CodeRound is hiring for this role.
- Research Engineer
- Data Synthesis
- Synthetic Data
- AI Research
- LLMs
- SFT
- RL Post-Training
- GRPO
- PPO
- DPO
- Code Generation
- Structured Documents
- Technical Specifications
- Data Validation
- Compiler-Based Validation
- Execution-Based Validation
- Test Harnesses
- Formal Verification
- Reward Models
- Verifiers
- C
- C++
- Rust
- Python
- Embedded Systems
- Safety-Critical Software
- Automotive
- Aerospace
- Telecom
- Defense
- EDA
- LLVM
- Clang
- CBMC
- Frama-C
- KLEE
- ESBMC
- ML Training Tooling
- Dataset Quality
- Data Pipelines
We’re looking for a Research Engineer to design and build pipelines that generate, execute, and validate synthetic training data for code and technical specifications. The role focuses on solving data scarcity for safety-critical embedded systems by building high-signal datasets validated through compilers, simulators, test harnesses, and formal verification tools. You’ll work closely with the post-training team to connect data gaps and model failures, while contributing to large-scale synthetic data generation and evaluation tooling.
What you'll do
- Design and build pipelines that generate, execute, and validate synthetic training data—code and specifications alike—using compilers, test harnesses, and formal verification tools, not just LLM-judge scoring.
- Build and maintain reward models / verifiers for RL post-training (e.g. GRPO) on code and spec-generation tasks.
- Own data quality end to end, including decontamination, filtering, coverage analysis, and dataset documentation.
- Work closely with the post-training team to close the loop between data gaps and model failures.
- Contribute to internal tooling for large-scale synthetic data generation and evaluation.
Must have
Experience building synthetic or instruction-tuning datasets for code or structured technical documents (specs, requirements), with a named, shipped dataset, tool, or paper you can point to. Direct experience with execution-based, compiler-based, or test-based data validation. Familiarity with RL post-training methods (e.g. GRPO, PPO, DPO) and/or SFT data pipelines. Strong software engineering fundamentals. Comfortable in one or more systems languages (C, C++, Rust, or similar) and at least one scripting language.
Good to have
Exposure to embedded, safety-critical, or hardware-adjacent software (automotive, aerospace, telecom, defense, or EDA tooling). Experience with formal verification tools such as CBMC, Frama-C, KLEE, or ESBMC. Experience with compiler internals such as LLVM or Clang. Contributions to open-source ML training/eval tooling such as axolotl, unsloth, TRL, distilabel, EvalPlus, or bigcode-project.