Site Reliability Engineer

CodeRound is hiring for this role for a VC-backed startup.

Bengaluru · Onsite · Full-time · ₹15-25L · 2+ yrs

Join a fast-growing GenAI startup as an SRE and take ownership of the reliability, scalability, and operational excellence of a platform powering the end-to-end ML lifecycle. You’ll work across Kubernetes, cloud infrastructure, production systems, observability, automation, and model-serving workloads, while helping build strong SRE and incident-management practices.

What you'll do

Must have

Experience in SRE, DevOps, or Platform Engineering Strong hands-on knowledge of Linux, networking, and Kubernetes Experience with AWS, GCP, or Azure Hands-on experience with Terraform, Helm, and CI/CD Strong troubleshooting, incident management, and problem-solving skills Scripting/programming experience in Python, Bash, or Go

Good to have

Experience with MLOps / AI infrastructure Exposure to GPU clusters and model serving Knowledge of release engineering Understanding of reliability and operational best practices

Apply for this role on CodeRound AI