Applied AI Engineer
CodeRound is hiring for this role for a VC-backed startup.
- AI Engineer
- Applied AI
- Artificial Intelligence
- LLM
- Large Language Models
- RAG
- Retrieval Augmented Generation
- AI Evals
- LLM Evaluation
- Evaluation Harness
- AI Agents
- Agentic AI
- ReAct
- Tool Use
- AI Research
- Applied Research
- Backend Engineering
- AI Systems
- Retrieval Systems
- Reranking
- Chunking
- AI Architecture
- Machine Learning
- Generative AI
- Open Source AI
We are looking for an AI Engineer / Applied AI Engineer with 3–6 years of experience to work on AI-powered customer support systems for D2C e-commerce brands. The role is focused on building and improving AI systems that handle real customer conversations every day. This role leans more toward applied research than conventional backend engineering. You’ll work on evaluation, retrieval, agent architectures, reasoning, memory and AI system failure analysis while designing experiments and making architectural decisions based on evidence.
What you'll do
- Design experiments to determine whether changes to AI systems actually improve response quality.
- Build evaluation harnesses, datasets and failure taxonomies that reflect real-world distributions rather than cherry-picked queries.
- Diagnose AI system regressions and identify underlying classes of failures rather than isolated instances.
- Investigate retrieval failures and determine how to fix the broader class of retrieval problems.
- Reason about where system behaviour should be deterministic versus LLM-driven.
- Explore how AI systems should remember information over long-running customer relationships.
- Argue for architectural decisions using evidence, with experiments and real production feedback informing what gets shipped.
Must have
3–6 years of relevant experience in AI, ML, backend engineering or related technical roles. Strong interest in modern AI and follows the field closely, with the ability to form opinions about what works and what is hype. Strong understanding of modern AI and backend systems, including LLM pipelines, retrieval systems and agent architectures. Ability to reason about AI system failure modes and enough backend fundamentals to ship what you design. Comfortable working with ambiguity and turning vague concerns into measurable questions. Strong ability to think in abstractions and identify the class of failure rather than just the individual instance. Willingness to openly disagree, argue from first principles and change your mind when evidence supports a different conclusion.
Good to have
Hands-on experience with AI evals, including building evaluation harnesses, designing datasets and measuring LLM output quality. Experience identifying eval inflation and dataset contamination while evaluating AI systems. Experience with RAG systems in production, including retrieval debugging, reranking and chunking strategy. Familiarity with agentic patterns, including ReAct-style loops, tool use and orchestration. Understanding of where agentic systems and architectures break and how to diagnose these failures. Prior experience working at an AI-first startup or research lab. Serious open-source AI contributions or substantial personal AI projects.