Senior Site Reliability Engineer
Infosys
We are looking for an experienced Site Reliability Engineer (SRE) / Production Engineer to build and operate highly available, scalable, and resilient production systems. The ideal candidate will have strong expertise in cloud infrastructure, automation, observability, incident management, performance engineering, and DevOps practices.
The role requires close collaboration with software engineering, infrastructure, security, and platform teams to ensure operational excellence and improve system reliability at scale.
Reliability Engineering Design, build, and maintain highly available and fault-tolerant production systems. Define and monitor SLIs, SLOs, and SLAs for critical services. Drive reliability improvements through automation and proactive engineering. Conduct capacity planning and performance optimization activities. Production Support & Operations Manage production environments and ensure service uptime. Lead incident response, troubleshooting, and root cause analysis (RCA). Develop runbooks, operational playbooks, and disaster recovery procedures.
Automation & DevOps Automate deployments, infrastructure management, and operational workflows. Improve CI/CD pipelines and release processes. Implement self-healing, auto-scaling, and operational automation solutions. Promote DevOps and SRE best practices across engineering teams. Security & Compliance Ensure production environments meet security and compliance requirements. Manage secrets, access controls, and vulnerability remediation. Partner with security teams to implement security best practices.
Don't want to miss the next one?
Subscribe to daily email alerts for roles matching your interests.

