Sr. Staff Observability Engineer (GPU Cloud & Telemetry Platform)
Key Responsibilities End-to-End Observability Platform Ownership : Design and scale telemetry pipelines using: Grafana Alloy for metrics collection (Prometheus-compatible pipelines) Datadog Vector for high-throughput log ingestion and transformation Grafana Mimir for scalable time-series storage Grafana Loki for log aggregation and querying Strategic Roadmap : Define the multi-year vision for GPU infrastructure observability, transitioning from reactive monitoring to SLO-driven, predictive, and automated observability . High-Cardinality Telemetry Design : Optimize pipelines for GPU workloads characterized by: High-cardinality labels (GPU IDs, tenants, workloads) Burst-heavy workloads (ML training, inference spikes) Multi-tenant isolation requirements Architect low-latency, high-throughput pipelines capable of ingesting: GPU metrics (utilization, memory, thermals, MIG partitions) Kubernetes and container telemetry Distributed system logs and traces Build and optimize metric pipelines (Alloy → Mimir) ensuring: Efficient remote_write tuning Cost-effective retention strategies Horizontal scalability and compaction tuning Design log pipelines (Vector → Loki) with: Structured logging and enrichment Intelligent filtering/sampling Stream partitioning for high-ingest environments Establish deep observability into: GPU hardware (NVIDIA DCGM, MIG, NVLink, PCIe) Kubernetes GPU operators and scheduling behavior Network fabric (RDMA, InfiniBand, TCP performance) Define GPU-specific SLIs/SLOs such as: GPU utilization efficiency Job scheduling latency Cluster fragmentation Thermal and power anomalies Build rich Grafana dashboards for: Real-time GPU fleet health Tenant-level usage and billing insights Capacity planning and forecasting Standardize dashboard frameworks and reusable panels across teams Enable self-service observability for platform and ML engineering teams Drive adoption of SRE principles : SLIs, SLOs, error budgets tailored to GPU workloads Integrate observability into CI/CD and IaC pipelines (Terraform/Kubernetes) : Automated canary analysis Observability-driven rollbacks Build automation (Go/Python) for: Pipeline health monitoring Dynamic routing and scaling of telemetry workloads Develop tooling and practices for cross-layer correlation : GPU → Node → Kubernetes → Application → Network Lead deep RCA efforts for: GPU contention issues Performance degradation in ML workloads Telemetry pipeline backpressure/failures Enable “needle-in-a-haystack” debugging using unified logs + metrics Mentor engineers and lead design reviews for observability systems Act as a force multiplier across SRE, Infra, and ML platform teams Promote Observability-by-Design in all new GPU cluster deployments Drive adoption and contribution to: Grafana stack (Alloy, Mimir, Loki, Tempo) OpenTelemetry ecosystem Define build vs. buy decisions (Datadog vs OSS vs hybrid approaches) Optimize interoperability between Vector and OTEL pipelines Architect secure telemetry pipelines with: Encryption in transit and at rest Multi-tenant isolation and RBAC Data residency compliance Implement Zero Trust observability patterns Qualifications & Requirements BS/MS in Computer Science or equivalent practical experience Extensive experience in Observability, SRE, or Distributed Infrastructure Proven track record building large-scale telemetry pipelines (metrics/logs) Observability Stack : Grafana Alloy / Prometheus ecosystem Grafana Mimir (or Cortex/Thanos) Grafana Loki Datadog Vector (or similar log pipelines) Programming : Strong in Go or Python Data Systems : TSDBs and log storage at scale Infrastructure : Kubernetes, Linux internals GPU systems (NVIDIA DCGM, CUDA ecosystem) High-performance networking (RDMA, InfiniBand preferred) Cloud & Hybrid: Experience building observability across: Bare-metal GPU clusters Hybrid cloud environments Core Impact Success in this role is measured by: A highly reliable, scalable observability platform powering GPU infrastructure Ability to diagnose complex GPU and distributed system issues in minutes Enabling data-driven optimization of GPU utilization and cost efficiency Building systems that proactively detect and mitigate failures before user impact Recruitment Process and Others Recruitment Process Application Review - 1st Virtual Interview - 2nd Virtual Interview - Offer The exact nature of the recruitment process may vary according to the specific job and may be changed due to scheduling or other circumstances. Interview schedules and the results will be informed to the applicant via the e-mail address submitted at the application stage. Details to Consider This job posting may be closed prior to the stated end date for application if all openings are filled. Coupang has the right to rescind an offer of employment if a candidate is found to have submitted false information as part of the application process. Those eligible for employment protection (recipients of veteran’s benefits, the disabled, etc.) may receive preferential treatment for employment in accordance with applicable laws. J ob titles and responsibilities may be subject to change depending on the candidate's overall experience, etc. this will be communicated to the candidate at the appropriate time before the offer. Hiring may be restricted in case the legal qualifications required for hiring and work performance is not met. This is a full-time regular position and includes 12 weeks of probation period; provided, however, the probationary period may be either skipped, shortened or extended if necessary for business purposes. Privacy Notice Your personal information will be collected and managed by Coupang as stated in the Application Privacy Notice is located below. https://privacy.coupang.com/en/land/jobs/ Document Return Policy This notification is given pursuant to Article 11 (6) of the Fair Hiring Procedure Act. A job applicant, who has applied but not been finally selected for a position at Coupang (the “Company”), may request
Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.
Visa jobbet och ansökUrsprunglig annons: www.coupang.jobs