ML Research Engineer

CodeRound is hiring for this role.

Bengaluru · Onsite · Full-time · ₹25-35L · 3+ yrs

We’re looking for a Research Engineer to design and build pipelines that generate, execute, and validate synthetic training data for code and technical specifications. The role focuses on solving data scarcity for safety-critical embedded systems by building high-signal datasets validated through compilers, simulators, test harnesses, and formal verification tools. You’ll work closely with the post-training team to connect data gaps and model failures, while contributing to large-scale synthetic data generation and evaluation tooling.

What you'll do

Must have

Experience building synthetic or instruction-tuning datasets for code or structured technical documents (specs, requirements), with a named, shipped dataset, tool, or paper you can point to. Direct experience with execution-based, compiler-based, or test-based data validation. Familiarity with RL post-training methods (e.g. GRPO, PPO, DPO) and/or SFT data pipelines. Strong software engineering fundamentals. Comfortable in one or more systems languages (C, C++, Rust, or similar) and at least one scripting language.

Good to have

Exposure to embedded, safety-critical, or hardware-adjacent software (automotive, aerospace, telecom, defense, or EDA tooling). Experience with formal verification tools such as CBMC, Frama-C, KLEE, or ESBMC. Experience with compiler internals such as LLVM or Clang. Contributions to open-source ML training/eval tooling such as axolotl, unsloth, TRL, distilabel, EvalPlus, or bigcode-project.

Apply for this role on CodeRound AI