Site Reliability Engineering Lead_Truist
Infosys
The Site Reliability Engineering Lead is a senior, hands-on technical leader within the Corporate Technology and Operations organization. This teammate is accountable for elevating the reliability, resiliency, and operational excellence of critical enterprise platforms across hybrid cloud and onprem environments.
Reliability Engineering & Automation
Architect and deliver automation solutions that eliminate toil, reduce MTTR, and increase service resilience. Experience in Ansible, Puppet or Chef is a plus.
Implement intelligent alerting, anomaly detection, and event correlation leveraging AI and AIOps tools.
Guide and enforce SLO/SLI adoption across product teams, ensuring metrics inform decision-making and prioritization.
Utilize Infrastructure-as-Code (IaC) tools for automating deployment of assets within cloud tenants.
Observability & Operational Excellence
Ensure operational readiness of applications and platforms through resiliency testing, chaos engineering, and failure-mode validation.
Cross-Functional Leadership & Influence
Partner with Delivery, Architecture, Security, and Risk teams to embed reliability and resilience into design and execution.
Standardization & Documentation
Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns.
Follow best practices and internal processes for Non-Functional requirements to improve resiliency and reliability.
Mentorship & Technical Development
Coach and mentor Associate, Professional, and Senior SREs to build technical depth and operational discipline.
Provide thought leadership in SRE methodologies, cloud-native operational patterns, and automated reliability engineering.
Incident Leadership & Production Operations
Lead P1/P0 incident bridges and direct technical investigation efforts.
Perform hands-on triage using logs, traces, metrics, and application telemetry.
Don't want to miss the next one?
Subscribe to daily email alerts for roles matching your interests.
