Director of Production Engineering

Legion Remote, United States Publicerat 7 september 2026
full_timeremotemid
<p> <strong>Headquarters:</strong> Remote, United States </p> <h1>Director of Production Engineering </h1> <h2><strong>Remote, United States</strong></h2> <p><strong>About this Position</strong></p> <p>Are you passionate about building the reliability, automation, and security foundations that let engineering teams move fast with confidence? At Legion, we are seeking a Director of Engineering, DevOps & SRE to lead the teams responsible for the availability, scalability, and security of our production environment. Our production infrastructure runs on AWS, leveraging services such as EKS, RDS, and a broad set of AWS-native technologies. You will partner closely with engineering and IT to build resilient systems, drive operational excellence, and ensure our platform meets the highest standards of security and compliance.</p> <p>This is a hands-on leadership role where you'll spend ~20-30% of your time contributing directly to architecture, tooling, and incident response, and the rest driving vision, roadmap, and cross-team execution.</p> <p><strong>Responsibilities</strong></p> <ul> <li>Hire and build a globally-distributed DevOps/SRE engineering team. Recruit, mentor, and manage engineers, and foster a culture of ownership, collaboration, and continuous improvement.</li> <li>Own the reliability and infrastructure roadmap for our AWS-based production environment, including EKS, RDS, and related AWS services, ensuring scalability, high availability, and cost efficiency.</li> <li>Lead the organization's security operations (SecOps) practice, including vulnerability management, threat detection, incident response, and remediation, to proactively identify and resolve security issues before they impact customers.</li> <li>Define and drive engineering OKRs for infrastructure reliability, automation, and security, and track progress against measurable outcomes.</li> <li>Champion observability and alerting best practices (e.g., Datadog), including automating alert triage and response to reduce mean-time-to-resolution.</li> <li>Solid understanding of agentic AI infrastructure and how AI agentic workflows apply to SDLC and DevOps processes (e.g., automated investigation, remediation, and PR-generation pipelines).</li> <li>Drive Infrastructure-as-Code, CI/CD, and automation practices to increase engineering velocity and reduce operational toil.</li> <li>Work closely with engineering and IT teams to align on infrastructure standards, access controls, tooling, and compliance requirements across the organization.</li> <li>Ensure the platform meets the highest standards of security, compliance, and data protection; implement and maintain robust security controls and audit-readiness.</li> <li>Lead and participate in the Incident Management on-call rotation, working with SRE and development teams to meet and exceed availability goals.</li> <li>Stay current on cloud, DevOps, and security best practices, and provide technical guidance and thought leadership to the broader engineering organization.</li> </ul> <p><strong>Required Qualifications</strong></p> <ul> <li>8-12 years of experience in DevOps, Site Reliability Engineering, or production infrastructure roles, including people management experience.</li> <li>Deep hands-on experience running production workloads on AWS, including EKS (Kubernetes), RDS, and other core AWS services (e.g., VPC, IAM, Lambda, S3).</li> <li>Demonstrated experience running security operations (SecOps) — vulnerability management, incident response, and remediation of production security issues.</li> <li>5+ years of experience leveraging observability platforms (e.g., Datadog, Prometheus, Grafana) to drive reliability, performance, and alerting improvements.</li> <li>Strong experience with Infrastructure-as-Code (e.g., Terraform, CloudFormation) and CI/CD automation.</li> <li>Proficiency in at least one of Go, Python, or Bash, with day-to-day use of Git and test automation pipelines.</li> <li>Hands-on experience operating Linux/Unix production platforms (Amazon Linux, Ubuntu, RHEL/CentOS).</li> <li>Proven track record partnering cross-functionally with engineering and IT teams to align on infrastructure, tooling, and security standards.</li> <li>Demonstrated experience leading incident management and on-call practices for high-availability production systems.</li> <li>Bachelor's degree in Computer Science, Engineering, or related field required; Master's degree preferred.</li> </ul> <p><strong>Preferred Qualifications</strong></p> <ul> <li>Experience with major cloud providers beyond AWS, such as Google Cloud Platform or Oracle Cloud Infrastructure (OCI).</li> <li>Relevant security certifications (e.g., CISSP, AWS Security Specialty, CKS).</li> <li>Experience with compliance frameworks such as SOC 2, ISO 27001, or HIPAA.</li> <li>3+ years of experience with Kubernetes or other containerization/orchestration platforms at scale.</li> <li>Experience with Kubernetes-native delivery tooling, including Argo Workflows and Helm.</li> <li>5+ years of experience leading teams in an agile/scrum environment.</li> <li>Experience building or scaling automated investigation and remediation pipelines for production error classes.</li> </ul> <p><strong>COMPENSATION & BENEFITS</strong></p> <p><em>Salary Range: Base Salary Range  </em>$220,000 - $265,000 <em>+ Bonus + </em&

Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.

Visa jobbet och ansök

Ursprunglig annons: weworkremotely.com