Site Reliability Engr II
Honeywell
We are seeking a highly technical SRE Engineer to design, build, and maintain fault-tolerant, scalable, and highly available distributed systems. You will champion SRE best practices, reduce manual operations (toil) via automation, and partner with product development squads to embed reliability into the software delivery lifecycle
Your role will include overseeing, supervising and reviewing tasks performed by team members to ensure effective execution of work; managing end-to-end processes and projects for both internal and external clients with responsibility for timely and accurate delivery; issuing clear instructions and directions to team members on tasks to be performed; and mentoring and guiding junior colleagues to Support their skill development, professional growth, and overall success.
Key Responsibilities:
Reliability & Availability
- Ensure high availability and uptime of production services.
- Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs).
- Design and implement disaster recovery (DR) strategies, including RTO/RPO targets.
- Conduct capacity planning and scalability assessments.
Monitoring & Observability
- Implement and maintain monitoring, logging, and alerting systems.
- Create dashboards to track system health and performance.
- Improve observability using tools such as Prometheus, Grafana, Azure Monitor, Dynatrace, or Elastic.
- Proactively detect, investigate, and resolve system issues.
- Incident Management
Don't want to miss the next one?
Subscribe to daily email alerts for roles matching your interests.

