Site Reliability Engineer
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. What you'll be doing: Support and contribute to SRE initiatives that improve reliability, scalability, and developer efficiency across enterprise systems. Assist in building and maintaining distributed systems that power NVIDIA’s AI-powered enterprise products and services, learning modern architectural patterns along the way. Help automate database operations — including provisioning, scaling, backup, and failover — for relational and vector database services. Contribute to observability and monitoring efforts by building dashboards, alerts, and automation scripts to improve system performance and reliability. Participate in incident response processes, learning to triage issues, reduce mean time to resolution (MTTR), and contribute to post-incident reviews. Collaborate with Cloud, Platform, Security, and AI/ML teams to support platform reliability and help implement SRE best practices. Learn to operate and troubleshoot complex systems — including Kubernetes-based and cloud-native infrastructure — following established standards in system design and incident management. Explore and adopt AI-assisted engineering practices, including coding agents and LLM-powered tooling, to accelerate day-to-day development workflows. What we need to see: BS degree in Computer Science or a related technical field (e.g., physics, mathematics), or equivalent practical experience. Foundational proficiency in at least one programming language such as Python, TypeScript, JavaScript, or Go. Basic understanding of cloud platforms (AWS, Azure, or GCP) and containerisation technologies like Docker and Kubernetes. Exposure to or coursework in infrastructure-as-code tools (e.g., Terraform, AWS CDK, CloudFormation) or willingness to learn. Familiarity with Linux/Unix systems, networking fundamentals, and version control (Git). Interest in observability concepts (logging, metrics, tracing) and tools such as OpenTelemetry, Prometheus, or Grafana. Basic knowledge of relational databases (e.g., PostgreSQL, MySQL) — understanding of SQL, indexing, and simple query optimisation. Strong problem-solving skills, curiosity, and a willingness to learn in a fast-paced, collaborative environment. Good communication and teamwork skills, with the ability to ask the right questions and learn from senior engineers. Ways to stand out from the crowd: Personal projects, internships, or coursework involving cloud infrastructure, automation, or DevOps/SRE practices. Contributions to open-source projects or active participation in hackathons, coding competitions, or technical communities. Exposure to AI/ML concepts — e.g., building or deploying a simple ML model, experimenting with LLM APIs, or using AI-powered developer tools (Copilot, Cursor, etc.). Hands-on experience with CI/CD pipelines, scripting for automation, or container orchestration (even in personal or academic projects). A strong sense of ownership, curiosity, and initiative — you turn challenges into learning opportunities and aren’t afraid to dive into unfamiliar systems. NVIDIA leads the charge in innovative breakthroughs in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, functions as the visual cortex of today’s computers and forms the core of our products and services. Our work opens new realms to explore, encourages outstanding creativity and discovery, and powers inventions once thought of as science fiction — from artificial intelligence to autonomous systems. NVIDIA is searching for outstanding talent like you to help us advance the next wave of artificial intelligence! Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/
Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.
Visa jobbet och ansökUrsprunglig annons: hitmarker.net