Manager, Site Reliability Engineering
About Delinea:
Delinea is a pioneer in securing human and machine identities through intelligent, centralized authorization, empowering organizations to seamlessly govern their interactions across the modern enterprise. Leveraging AI-powered intelligence, Delinea’s leading cloud-native Identity Security Platform applies context throughout the entire identity lifecycle – across cloud and traditional infrastructure, data, SaaS applications, and AI. It is the only platform that enables you to discover all identities – including workforce, IT administrator, developers, and machines – assign appropriate access levels, detect irregularities, and respond to threats in real-time. With deployment in weeks, not months, 90% fewer resources to manage than the nearest competitor, and a 99.995% uptime, Delinea delivers robust security and operational efficiency without compromise. Learn more about Delinea on Delinea.com http://Delinea.com, LinkedIn https://nam12.safelinks.protection.outlook.com/?url=https%3A%2F%2Fwww.linkedin.com%2Fcompany%2Fdelinea%2F&data=05%7C02%7CMichele.Eaton%40delinea.com%7C766420d58b9a492d94ca08ddd6949f59%7C18d6ed03f0e5486183d2bb7b6c7c2bb2%7C0%7C0%7C638902656186877399%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=b1LCPIWSjxvblUYLlabUnRZ43S8rcspTQcBZsz0UbmU%3D&reserved=0, X https://nam12.safelinks.protection.outlook.com/?url=https%3A%2F%2Fx.com%2Fdelineainc&data=05%7C02%7CMichele.Eaton%40delinea.com%7C766420d58b9a492d94ca08ddd6949f59%7C18d6ed03f0e5486183d2bb7b6c7c2bb2%7C0%7C0%7C638902656186902236%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=goZhva9hWR8F0P5pO6SQUEwigdbvbZiWqmw5%2BHhhf1A%3D&reserved=0, and YouTube https://nam12.safelinks.protection.outlook.com/?url=https%3A%2F%2Fwww.youtube.com%2F%40Delinea&data=05%7C02%7CMichele.Eaton%40delinea.com%7C766420d58b9a492d94ca08ddd6949f59%7C18d6ed03f0e5486183d2bb7b6c7c2bb2%7C0%7C0%7C638902656186928751%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=uXgjuXTFu4kHWBJZ1XdNWN88Mrk5iT9Y24Ejw8fFBi8%3D&reserved=0.
Join our passionate, global team at Delinea and help us make the world a safer and more secure place. Our success is driven by world-class product leadership, outstanding engineers, and strategic investment from TPG. We value diversity, innovation, and a culture of respect and fairness. If you're ready to push boundaries and challenge the status quo in security, we want to hear from you.
Apply today to help us achieve our mission.
Summary:
Delinea is looking for a hands-on Manager of Site Reliability Engineering to lead the SRE and DevOps engineers supporting the Delinea products. This is a working manager role. You will be expected to lead people and lead work: writing and reviewing automation, digging into AKS and Azure telemetry, commanding Sev1 and Sev2 incidents, and improving the observability and deployment practices your team depends on.
The initial scope is the Platform SRE and DevOps team. Over time, the role is expected to expand to cover our FedRAMP High environment and the broader Commercial Platform footprint, so comfort operating in a regulated environment and a willingness to grow scope are essential.
You will lead a blended team of full-time engineers and contractors distributed across multiple time zones. Meeting your team inside their working hours is an expectation of the role. On-call participation is required.
What You Will Do:
- Lead hands-on. Spend a meaningful portion of your week in the environment: reviewing pull requests, validating pipeline changes, tuning monitors and dashboards, running queries in Datadog, and troubleshooting production issues alongside your engineers. This role does not sit above the work.
- Own availability and performance of the Delinea Platform production environments across Azure and AWS, including AKS workloads, ingress and networking, data services, messaging, and CDN or WAF layers.
- Manage a blended team. Hire, onboard, coach, and develop full-time SRE engineers. Direct and manage contractor resources, including scoping work, setting quality expectations, and reviewing deliverables.
- Lead a distributed team. Run one-on-ones, standups, and planning sessions at times that work for engineers in other geographies.
- Participate in on-call. Carry the pager as part of the rotation, act as incident commander for Sev1 and Sev2 events, drive engagement of the right responders, and own communication cadence with support, engineering, and leadership until resolution.
- Command incident response end to end. Own detection, triage, mitigation, customer-facing status communication, and post-incident review. Ensure RCAs are written to a customer-ready standard, preventative actions have owners and target dates, and those actions are driven to closure.
- Raise the observability bar. Improve detection coverage so that issues are found by our monitoring rather than by a customer ticket. Own SLI and SLO definition, alert quality and noise reduction, synthetic coverage, APM instrumentation, log hygiene, and dashboard standards.
- Support FedRAMP and regulated operations. Grow into supporting our FedRAMP High environment, including change control discipline, evidence collection, boundary awareness, and the operational differences between government and commercial environments.
- Reduce toil through automation. Set the expectation that repeat manual work becomes code. Prioritize automation backlog alongside project and reliability work.
- Report on operational health. Produce and present incident metrics, trends, and reliability commitments to leadership, and translate them into a concrete improvement plan.
What You Will Need:
- 6+ years in Site Reliability Engineering, DevOps, or Cloud Operations, with demonstrated
Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.
Visa jobbet och ansökUrsprunglig annons: jobs.ashbyhq.com