AWS Site Reliability Engineer
Infosys
We are seeking a skilled and motivated Site Reliability Engineer with hands-on expertise in application operations, DevOps tools, and SRE principles. The ideal candidate will have experience in supporting production systems, DEVOPS hands-on, a solid understanding of observability, and a foundational grasp of SRE principles. The role also requires basic to intermediate programming skills and familiarity with modern development practices.
Provide production support for Production applications, ensuring the stability and availability of systems.
- Diagnose, troubleshoot, and resolve production issues in real time.
- Design, implement, and maintain CI/CD pipelines using Jenkins, Git/Bitbucket,
and other DevOps tools.
- Manage and deploy applications using containerization and orchestration tools
like Docker and Kubernetes.
- Set up and maintain observability tools (Grafana, Prometheus, Instana) for
monitoring, logging, and alerting.
- Write and maintain infrastructure as code using Terraform.
- Collaborate with development teams to implement SRE practices and
principles, ensuring reliability, scalability, and performance.
- Assist in incident management and post-mortem analysis to improve system
reliability.
- Contribute to the automation of repetitive tasks and system processes.
Must have skills:
Production Support Expertise: Experience in application operations, troubleshooting, and system monitoring.
• DevOps Tools and Platforms:
o CI/CD: Jenkins, Git/Bitbucket. o Virtualization & Orchestration: Docker, Kubernetes. o Observability: Grafana, Prometheus, Instana. o Infrastructure as Code: Terraform.
Don't want to miss the next one?
Subscribe to daily email alerts for roles matching your interests.


