Principal Software Engineer - ML Platform Engineer
the opportunity to responsibly bring new AI-native automation patterns into real engineering workflows. Responsibilities: Design, implement, and evolve internal platform capabilities that make AI Efficiency services easier to build, ship, observe, secure, and operate Build and maintain self-service workflows, reusable platform abstractions, and golden paths that improve developer productivity while preserving reliability, security, and governance Improve platform reliability through better monitoring, alerting, observability, deployment safety, release practices, and incident readiness Define and operationalize service health indicators, SLIs, SLOs, and related reliability metrics that help teams make informed tradeoffs between reliability, velocity, and cost Build automation that reduces operational toil and improves mean time to detect, respond, and recover from incidents Partner with engineers throughout the software development lifecycle to embed operability, production readiness, and maintainability into system design, implementation, rollout, and ongoing support Improve CI/CD systems, developer workflows, and release pipelines so shipping becomes safer, faster, and more repeatable Identify platform and reliability risks across distributed systems, infrastructure, service dependencies, and operational workflows, and drive durable improvements Troubleshoot AI model-serving issues across frameworks, runtimes, and hardware environments, including diagnosing configuration, compatibility, and performance issues across different GPU platforms and supporting model format conversion workflows when needed Design and run resilience, recovery, and failure-mode testing to validate system behavior under stress and uncover hidden weaknesses before they impact users Evaluate, integrate, and operate AI-assisted engineering tools that improve code quality, reliability, security, performance, and developer productivity across the software delivery lifecycle Build and evolve automation pipelines that combine conventional CI/CD systems with agentic workflows such as automated code review, bug detection, regression analysis, test generation, remediation suggestions, and workflow verification Partner with engineers to introduce safe, auditable, and measurable uses of AI agents in areas such as pull request review, operational diagnostics, UI and UX validation, accessibility checks, and production readiness checks Define guardrails, approval workflows, observability, reporting, and escalation paths for AI-assisted automation to ensure these systems remain safe, trustworthy, and operationally effective Establish evaluation frameworks and success metrics for AI-native development tooling, including quality lift, false positive rates, latency, cost, operational risk, and impact on engineering throughput Lead or contribute to incident response and post-incident improvement work for critical internal platforms and services, with a focus on systemic fixes and long-term resilience Champion platform and operational excellence through documentation, runbooks, standards, and tooling that raise the engineering bar across the broader organization Required Qualifications: Bachelor’s degree in Computer Science or a related field, or equivalent professional experience 5+ years of experience in Platform Engineering, Infrastructure Engineering, Site Reliability Engineering, DevOps, Developer Experience, or a similar role supporting production systems and engineering workflows Strong programming and automation skills in one or more languages such as Python, Go, or JavaScript / TypeScript Experience designing, building, or operating internal platforms, developer tooling, CI/CD systems, or shared infrastructure used by multiple engineering teams Experience operating and improving cloud-based production systems in AWS, GCP, Azure, or comparable environments Strong understanding of observability practices, including metrics, logs, traces, dashboards, and alert design Experience improving reliability and operability for distributed systems, service-oriented architectures, APIs, or platform infrastructure Experience with incident management, root cause analysis, and driving durable operational improvements after production issues Strong understanding of containerized environments and orchestration platforms such as Kubernetes, ECS, or similar technologies Ability to collaborate across teams, influence technical direction, and communicate clearly with both engineers and non-engineers Desired Qualifications: Experience supporting AI/ML platforms, inference services, model-serving systems, data pipelines, or GPU-backed workloads Experience defining and using SLOs, error budgets, and reliability metrics to guide prioritization and engineering decisions Experience with platform product thinking, including designing self-service experiences, paved roads, or golden paths for internal users Experience improving the reliability and usability of developer platforms, internal tools, or enterprise-facing services Familiarity with infrastructure as code and configuration management systems such as Terraform, Pulumi, or similar tools Experience with security, access control, secrets management, and operational hardening in production environments Experience balancing availability, latency, efficiency, cost, and ease of use in systems operating at scale Experience mentoring other engineers and raising platform or reliability standards through technical leadership Experience evaluating or integrating AI-assisted software engineering tools for code review, static analysis, test generation, incident investigation, or operational automation Familiarity with emerging agentic engineering workflows, including how AI agents can safely interact with source control, CI/CD pipelines, browser automation, and developer platforms Experience with end-to-end browser automation and testing frameworks such as Playwright, including their use for UI validation, a
Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.
Visa jobbet och ansökUrsprunglig annons: www.riotgames.com