Principal Software Engineer - DevOps / Site Reliability Engineer

Riot Games Singapore, Singapore Publicerat 12 augusti 2026
full_timeonsitesenior
the opportunity to responsibly bring new AI-native automation patterns into real engineering workflows, thoughtfully applying emerging capabilities to reduce friction, improve reliability, and enhance how engineers interact with production systems without compromising safety or control. Responsibilities: Own and continuously improve the reliability, availability, scalability, performance, and operational health of the Efficiency team’s (web) platform and the tools deployed within it Design, build, and maintain the infrastructure, deployment systems, and operational foundations required to support a growing portfolio of production AI services and internal tools Improve CI/CD pipelines, release engineering practices, environment management, and deployment automation so software can be shipped safely, quickly, and consistently Establish production-readiness standards and ensure new utilities have appropriate monitoring, alerting, ownership, documentation, rollback strategies, and support plans before launch Define and operationalize service health indicators, SLIs, SLOs, error budgets, and reliability metrics that guide engineering priorities and tradeoffs between reliability, velocity, cost, and complexity Build comprehensive observability across applications, infrastructure, service dependencies, and user workflows using metrics, logs, traces, dashboards, synthetic monitoring, and actionable alerts Establish sustainable incident-management and on-call practices, including escalation paths, runbooks, severity definitions, communication protocols, and clear service ownership Lead or contribute to the diagnosis and resolution of production incidents, coordinating across teams and driving blameless post-incident reviews and durable corrective actions Build automation that reduces operational toil, improves mean time to detect and recover, and eliminates recurring sources of failure or manual intervention Implement safe deployment patterns such as automated validation, progressive delivery, canary releases, feature flags, health checks, rollback mechanisms, and controlled environment promotion Perform capacity planning, load testing, performance analysis, and resource forecasting to ensure the Toolkit can support increasing adoption and usage across Riot Design and validate resilience, backup, recovery, failover, and disaster-recovery strategies for critical services, data, configurations, and infrastructure Identify single points of failure and systemic risks across applications, cloud infrastructure, networking, databases, queues, caches, third-party dependencies, and operational workflows Improve developer experience by building self-service workflows, reusable infrastructure components, local development environments, test environments, deployment tooling, and clear operational documentation Establish and maintain infrastructure-as-code, configuration-management, secrets-management, and environment-governance practices that make infrastructure changes safe, repeatable, and auditable Partner with engineers throughout the software development lifecycle to embed reliability, operability, security, and maintainability into system design rather than addressing them only after launch Troubleshoot complex production issues across web applications, APIs, distributed services, containerized workloads, cloud infrastructure, network boundaries, authentication systems, and external service dependencies Partner with ML Platform Engineers to ensure model-serving and inference systems integrate cleanly with the team’s broader observability, deployment, incident-management, and reliability standards Collaborate with Riot infrastructure, information security, IT, developer-platform, and compliance teams to ensure the team follows appropriate operational and security requirements Evaluate and implement AI-assisted operational workflows such as automated anomaly investigation, log analysis, remediation recommendations, regression detection, and runbook automation Define guardrails, approval requirements, auditability, and escalation paths for agentic or automated operational systems that can interact with production environments Champion operational excellence through technical leadership, mentoring, documentation, standards, architecture reviews, and tooling that raise the reliability bar across the team Required Qualifications: Bachelor’s degree in Computer Science or a related field, or equivalent professional experience 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, Platform Engineering, Production Engineering, Developer Experience, or a similar role supporting production systems Strong programming and automation skills in one or more languages such as Python, Go, JavaScript, or TypeScript Experience designing, operating, and improving cloud-based production systems in AWS, GCP, Azure, or comparable environments Experience building and maintaining CI/CD pipelines, release systems, deployment automation, and environment-management workflows Strong understanding of observability practices, including metrics, logging, distributed tracing, dashboards, synthetic monitoring, and alert design Experience participating in or leading incident response, on-call support, root-cause analysis, and post-incident improvement work Experience improving the reliability, availability, scalability, and performance of distributed systems, service-oriented architectures, APIs, or web platforms Strong understanding of containerized environments and orchestration technologies such as ECS, Docker, Kubernetes, or comparable systems Experience with infrastructure-as-code and configuration-management tools such as Terraform, Pulumi, CloudFormation, or similar technologies Working knowledge of Linux systems, networking, DNS, load balancing, service discovery, authentication, secrets management, and cloud security fundamentals Ability to identify systemic operational risks and drive durable

Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.

Visa jobbet och ansök

Ursprunglig annons: www.riotgames.com