AI Research Engineer
CodeRound is hiring for this role for a VC-backed startup.
- AI Research Engineer
- Applied Research
- ML Research
- Machine Learning
- LLM
- RAG
- Retrieval
- Embeddings
- Vector Search
- LLM Evaluation
- Evaluation Harnesses
- Python
- SQL
- Statistics
- Experiment Design
- Cohort Analysis
- Survival Analysis
- Personalization
- Recommender Systems
- Longitudinal Data
- Panel Data
- Clinical Research
- Epidemiology
- Multilingual Text
- Model Evaluation
Work across applied AI research and production engineering, exploring how AI can better understand people, learn from longitudinal health data, and measure real-world outcomes. The role combines literature review, experimentation, system design, production Python, LLM/RAG systems, statistical analysis, and evaluation.
What you'll do
- Read and evaluate relevant research literature
- Test research findings against proprietary data
- Design and build AI systems rather than limiting work to notebooks
- Write production Python and ship research into real-world systems
- Build systems around embeddings and retrieval
- Develop RAG and LLM-based systems using structured outputs and context construction
- Build evaluation harnesses and benchmarks
- Define what a good result looks like before running experiments
- Determine which questions should be evaluated using deterministic code versus models
- Design experiments and valid comparison groups
- Analyze cohorts and survival data
- Build measurement frameworks before shipping changes
- Track experiments and maintain reproducible research pipelines
- Evaluate whether research translates into measurable real-world outcomes
- Investigate how AI can learn, retain, and recall information about individuals over long periods
- Explore health behaviour, personalization, and decision-making patterns
- Form evidence-based conclusions and revise hypotheses when the data points elsewhere
Must have
2+ years of applied research or ML engineering experience, or a research degree alongside real-world shipping experience Strong Python and SQL fundamentals Working knowledge of embeddings and retrieval Hands-on experience with vector search, chunking, indexing, and retrieval evaluation Experience building with LLMs using structured outputs and context construction Experience with retrieval-augmented systems Experience using models as evaluators and understanding their limitations Strong understanding of AI/ML evaluation and experiment design Experience building evaluation harnesses and benchmarks Strong statistical fundamentals including sample size, confounding, and base rates Understanding of valid comparison/control groups Experience with cohort and survival analysis Ability to build reproducible pipelines and track experiments Strong version control and engineering practices Ability to read research literature and test findings against real-world data Ability to design, build, ship, and evaluate production systems independently
Good to have
Clinical research or epidemiology experience Experience with longitudinal or panel data Experience with personalization systems Experience with recommender systems at scale Experience fine-tuning smaller models Experience distilling smaller models Experience working with noisy text Experience working with multilingual text