Senior Software Engineer
NVIDIA
We are seeking a Senior Software Engineer with strong infrastructure expertise to design, build, and operate the next generation of our enterprise Observability, Automation, and AI-driven Reliability Platform.
This role will build highly scalable distributed systems and platform services spanning Storage, Compute, Network, VMware, OpenShift, and bare-metal infrastructure. The engineer will help transform infrastructure operations from reactive monitoring and manual remediation to proactive, predictive, and AI-driven autonomous operations.
What You Will Be Doing:
- Design, build, and operate distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability at large scale.
- Develop reusable platform services, APIs, automation frameworks, and control planes that enable self-service, reduce operational toil, and automate infrastructure operations across multiple engineering teams.
- Build scalable telemetry and event-processing systems spanning metrics, logs, traces, events, topology, and alerts, with the performance and efficiency to process billions of infrastructure signals.
- Build intelligent and AI-native reliability capabilities, including agentic workflows for anomaly detection, forecasting, root-cause analysis, automated debugging, and closed-loop remediation.
- Drive technical architecture and engineering direction across Storage, Compute, Network, and Platform domains, solving complex and ambiguous problems that span multiple teams.
- Engineer for production at scale, with strong focus on software quality, scalability, security, performance, observability, maintainability, and operational readiness.
- Provide technical leadership and mentorship, influence engineering standards and architecture decisions, and deliver measurable improvements in reliability, MTTR, operational toil, engineering productivity, and infrastructure efficiency.
What We Need To See:
- Bachelor's or Master's degree in Computer Science, Engineering, or equivalent practical experience, with 10+ years of software engineering, SRE, infrastructure, or distributed-systems experience and demonstrated technical leadership.
- Strong software engineering expertise in Go, Python, or equivalent languages, with experience designing and building production-grade distributed systems, platform services, APIs, and automation.
Don't want to miss the next one?
Subscribe to daily email alerts for roles matching your interests.

