Senior Site Reliability Engineer
Infosys
Reliability Engineering Design, build, and maintain highly available and fault-tolerant production systems. Define and monitor SLIs, SLOs, and SLAs for critical services. Drive reliability improvements through automation and proactive engineering. Conduct capacity planning and performance optimization activities. Production Support & Operations Manage production environments and ensure service uptime. Lead incident response, troubleshooting, and root cause analysis (RCA). Develop runbooks, operational playbooks, and disaster recovery procedures. Participate in on-call rotations and major incident management processes.
Cloud & Infrastructure Deploy and manage cloud-native infrastructure across AWS, Azure, or GCP. Automate infrastructure provisioning using Infrastructure as Code (IaC). Implement scalable and secure infrastructure solutions. Support Kubernetes-based platforms and containerized workloads.
Security & Compliance Ensure production environments meet security and compliance requirements. Manage secrets, access controls, and vulnerability remediation. Partner with security teams to implement security best practices. Leadership & Mentoring Mentor junior SREs, DevOps Engineers, and Production Support Engineers. Lead incident reviews and reliability improvement initiatives. Conduct knowledge-sharing sessions and technical workshops. Drive operational excellence and engineering best practices.
Don't want to miss the next one?
Subscribe to daily email alerts for roles matching your interests.

