Director of Production Engineering
<p>
<strong>Headquarters:</strong> Remote, United States
</p>
<h1>Director of Production Engineering&nbsp;</h1>
<h2><strong>Remote, United States</strong></h2>
<p><strong>About this Position</strong></p>
<p>Are you passionate about building the reliability, automation, and security foundations that let engineering teams move fast with confidence? At Legion, we are seeking a Director of Engineering, DevOps &amp; SRE to lead the teams responsible for the availability, scalability, and security of our production environment. Our production infrastructure runs on AWS, leveraging services such as EKS, RDS, and a broad set of AWS-native technologies. You will partner closely with engineering and IT to build resilient systems, drive operational excellence, and ensure our platform meets the highest standards of security and compliance.</p>
<p>This is a hands-on leadership role where you'll spend ~20-30% of your time contributing directly to architecture, tooling, and incident response, and the rest driving vision, roadmap, and cross-team execution.</p>
<p><strong>Responsibilities</strong></p>
<ul>
<li>Hire and build a globally-distributed DevOps/SRE engineering team. Recruit, mentor, and manage engineers, and foster a culture of ownership, collaboration, and continuous improvement.</li>
<li>Own the reliability and infrastructure roadmap for our AWS-based production environment, including EKS, RDS, and related AWS services, ensuring scalability, high availability, and cost efficiency.</li>
<li>Lead the organization's security operations (SecOps) practice, including vulnerability management, threat detection, incident response, and remediation, to proactively identify and resolve security issues before they impact customers.</li>
<li>Define and drive engineering OKRs for infrastructure reliability, automation, and security, and track progress against measurable outcomes.</li>
<li>Champion observability and alerting best practices (e.g., Datadog), including automating alert triage and response to reduce mean-time-to-resolution.</li>
<li>Solid understanding of agentic AI infrastructure and how AI agentic workflows apply to SDLC and DevOps processes (e.g., automated investigation, remediation, and PR-generation pipelines).</li>
<li>Drive Infrastructure-as-Code, CI/CD, and automation practices to increase engineering velocity and reduce operational toil.</li>
<li>Work closely with engineering and IT teams to align on infrastructure standards, access controls, tooling, and compliance requirements across the organization.</li>
<li>Ensure the platform meets the highest standards of security, compliance, and data protection; implement and maintain robust security controls and audit-readiness.</li>
<li>Lead and participate in the Incident Management on-call rotation, working with SRE and development teams to meet and exceed availability goals.</li>
<li>Stay current on cloud, DevOps, and security best practices, and provide technical guidance and thought leadership to the broader engineering organization.</li>
</ul>
<p><strong>Required Qualifications</strong></p>
<ul>
<li>8-12 years of experience in DevOps, Site Reliability Engineering, or production infrastructure roles, including people management experience.</li>
<li>Deep hands-on experience running production workloads on AWS, including EKS (Kubernetes), RDS, and other core AWS services (e.g., VPC, IAM, Lambda, S3).</li>
<li>Demonstrated experience running security operations (SecOps) — vulnerability management, incident response, and remediation of production security issues.</li>
<li>5+ years of experience leveraging observability platforms (e.g., Datadog, Prometheus, Grafana) to drive reliability, performance, and alerting improvements.</li>
<li>Strong experience with Infrastructure-as-Code (e.g., Terraform, CloudFormation) and CI/CD automation.</li>
<li>Proficiency in at least one of Go, Python, or Bash, with day-to-day use of Git and test automation pipelines.</li>
<li>Hands-on experience operating Linux/Unix production platforms (Amazon Linux, Ubuntu, RHEL/CentOS).</li>
<li>Proven track record partnering cross-functionally with engineering and IT teams to align on infrastructure, tooling, and security standards.</li>
<li>Demonstrated experience leading incident management and on-call practices for high-availability production systems.</li>
<li>Bachelor's degree in Computer Science, Engineering, or related field required; Master's degree preferred.</li>
</ul>
<p><strong>Preferred Qualifications</strong></p>
<ul>
<li>Experience with major cloud providers beyond AWS, such as Google Cloud Platform or Oracle Cloud Infrastructure (OCI).</li>
<li>Relevant security certifications (e.g., CISSP, AWS Security Specialty, CKS).</li>
<li>Experience with compliance frameworks such as SOC 2, ISO 27001, or HIPAA.</li>
<li>3+ years of experience with Kubernetes or other containerization/orchestration platforms at scale.</li>
<li>Experience with Kubernetes-native delivery tooling, including Argo Workflows and Helm.</li>
<li>5+ years of experience leading teams in an agile/scrum environment.</li>
<li>Experience building or scaling automated investigation and remediation pipelines for production error classes.</li>
</ul>
<p><strong>COMPENSATION &amp; BENEFITS</strong></p>
<p><em>Salary Range: Base Salary Range&nbsp; </em>$220,000 - $265,000 <em>+ Bonus + </em&
Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.
Visa jobbet och ansökUrsprunglig annons: weworkremotely.com