We are seeking a highly technical SRE Leader to bridge software engineering and operations. You will lead a high-performing team to design, build, and maintain fault-tolerant, scalable, and highly available distributed systems. You will champion SRE best practices, reduce manual operations (toil) via automation, and partner with product development squads to embed reliability into the software delivery lifecycle
Key Responsibilities:
Reliability Strategy & SLO Management
Define and maintain Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets alongside product teams.
Establish reliability benchmarks and enforce architectural standards across the enterprise.
Act as an evangelist for the SRE mindset, driving a proactive rather than reactive operational culture.
Incident Response & Observability
Lead cross-team efforts during critical production outages and high-severity incidents.
Conduct blameless postmortems to identify root causes and implement systemic preventive measures.
Optimize observability platforms, telemetry, and distributed tracing to surface actionable alerts without noise
Automation & Architecture
Write robust software to automate infrastructure provisioning, configuration management, and deployment pipelines.
Spearhead capacity planning, performance tuning, and chaos engineering experiments.
Don't want to miss the next one?
Subscribe to daily email alerts for roles matching your interests.