Staff Software Engineer - Observability
What you will do: As a Staff Software Engineer working on the observability of the platform, you will play a key role in shaping how teams work with telemetry data, driving strategic decisions, identifying and advocating for best practices for building, running and maintaining observable components. This includes: Define and lead Logging, Metrics, Tracing & Alerting strategies Solve scaling bottlenecks in critical services in our telemetry data pipelines Create advanced tooling to accelerate root-cause analysis, reduce Mean Time to Resolution (MTTR), and eliminate alert fatigue Guide and mentor software engineering and SRE teams on monitoring best practices and instrumenting code. Identify and tackle bottlenecks and single points of failures in the architecture aiming to improve the reliability of the observability platform and the overall user experience. Driving features and products idealization, implementation and adoption. Collaborate on planning and refinement proactively providing input whilst striving for engineering alignment, quality and reducing complexity and dependencies. Respond to alerts and duty responsibilities, like providing support to other engineers and troubleshooting production issues What you bring: You have 5+ years of relevant work experience with highly distributed systems Expertise in designing and implementing APIs and data pipelines for high-throughput, real-time data ingestion Experience improving software reliability between different categories including availability, performance, latency, efficiency, capacity, SLOs and incident management. Proficiency developing and maintaining software and frameworks written in Go and/or Java Advanced understanding of system design and evaluating its tradeoffs Strong stakeholder management, technical communication, and mentorship capability. Highly motivated to learn and continuously develop yourself Experience with containerisation and orchestration technologies (Docker, Kubernetes) and infrastructure as code tools(e.g: Terraform) Observability Stack Expertise : You have hands-on experience operating core telemetry data stores at scale e.g. Elasticsearch/Opensearch/VictoriaLogs/Clickhouse for logging, Prometheus/ VictoriaMetrics for metrics and Grafana Tempo for distributed tracing, Grafana LGTM stack, OpenTelemetry, Alertmanager, Clickhouse Nice to haves: Experience with highly available/fault tolerant, replicated data storage systems, large scale data processing systems is a strong plus Infrastructure and Platform Experience Contributions to open-source observability projects Our Diversity, Equity and Inclusion commitments Our unique approach is a product of our diverse perspectives. This diversity of backgrounds and cultures is essential in helping us maintain our momentum. Our business and technical challenges are unique, and we need as many different voices as possible to join us in solving them - voices like yours. No matter who you are or where you’re from, we welcome you to be your true self at Adyen. Studies show that women and members of underrepresented communities apply for jobs only if they meet 100% of the qualifications. Does this sound like you? If so, Adyen encourages you to reconsider and apply. We look forward to your application! What’s next? Ensuring a smooth and enjoyable candidate experience is critical for us. We aim to get back to you regarding your application within 5 business days. Our interview process tends to take about 4 weeks to complete, but may fluctuate depending on the role. Learn more about our hiring process here . Don’t be afraid to let us know if you need more flexibility. This role is based out of our Amsterdam office. We are an office-first company and value in-person collaboration; we do not offer remote-only roles.
Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.
Visa jobbet och ansökUrsprunglig annons: job-boards.greenhouse.io