Site Reliability Engineer
CodeRound is hiring for this role for a VC-backed startup.
- SRE
- Site Reliability Engineering
- DevOps
- Platform Engineering
- Kubernetes
- AWS
- GCP
- Azure
- Linux
- Networking
- Terraform
- Helm
- CI/CD
- Python
- Bash
- Go
- MLOps
- AI Infrastructure
- GPU
- Model Serving
- Observability
- Monitoring
- Incident Management
- Production Infrastructure
- Cloud Infrastructure
- Release Engineering
- Reliability Engineering
- GenAI
Join a fast-growing GenAI startup as an SRE and take ownership of the reliability, scalability, and operational excellence of a platform powering the end-to-end ML lifecycle. You’ll work across Kubernetes, cloud infrastructure, production systems, observability, automation, and model-serving workloads, while helping build strong SRE and incident-management practices.
What you'll do
- Own platform uptime, reliability, scalability, and performance
- Manage Kubernetes clusters, cloud infrastructure, and production environments
- Establish and improve incident response, on-call, RCA, and postmortem processes
- Drive deployment, rollback, and change-management practices
- Handle capacity planning and disaster recovery
- Build and enhance monitoring, alerting, and operational dashboards
- Automate deployments, scaling, and repetitive operational workflows
- Troubleshoot complex production infrastructure and application issues
- Support GPU workloads and model-serving infrastructure
Must have
Experience in SRE, DevOps, or Platform Engineering Strong hands-on knowledge of Linux, networking, and Kubernetes Experience with AWS, GCP, or Azure Hands-on experience with Terraform, Helm, and CI/CD Strong troubleshooting, incident management, and problem-solving skills Scripting/programming experience in Python, Bash, or Go
Good to have
Experience with MLOps / AI infrastructure Exposure to GPU clusters and model serving Knowledge of release engineering Understanding of reliability and operational best practices