via Career pages·Today
Site Reliability Engineer
Infosys
Full-timeOn-site
Location:Mangalore, IndiaType:Full-timePosted:Today
- Manage capacity and performance to help scale the infrastructure both on public and private clouds around the world
- Define and implement standards and best practices related to: System Architecture, Deployment, metrics, operational tasks
- Support services through activities such as monitoring availability, system health, and incident response
- Improve system performance, application delivery and efficiency through automation, process refinement, postmortem reviews, and in-depth configuration analysis
- Engage in Communications across all areas of the organization
- Troubleshooting and monitoring production systems to ensure the highest uptimes are maintained
- Support and improve upon existing high-availability architecture solutions as well as manage the operational activity.
- Integrate Generative AI (GenAI) and AIOps tools to automate incident detection, root cause analysis, and resolution workflows (e.g., self-healing scripts, intelligent runbooks), reducing manual toil and accelerating response times.
- Apply Prompt Engineering techniques to enhance interactions with AI-based observability and automation platforms improving accuracy and efficiency of AI responses.
- Leverage platform-specific AI capabilities (e.g., AWS Bedrock, Azure OpenAI, GCP Vertex AI) to architect intelligent SRE solutions tailored to cloud environments.
Don't want to miss the next one?
Subscribe to daily email alerts for roles matching your interests.