Staff, Site Reliability Engineer (Tech Infra)
Key Responsibilities: · Serve as a primary point responsible for the reliability, health, and performance of all Coupang customer-facing services. · Gain deep knowledge of Coupang application workflow and dependencies. · Define and track key performance indicators (KPIs) and service-level objectives (SLOs) related to system availability, performance, and reliability. · Build world class incident management process and automation, including fast incident remediation, incident operational reviews and retrospectives. · Develop and implement best practices for creating and maintaining effective monitoring, alerting, and telemetry systems. · Build automation to execute regular Disaster Recovery testing and load testing to stay ahead of expected growth of Coupang services. · Work closely with product development teams to ensure the products are designed with scale and operability in mind. · Build right guardrails and automation for deploying production changes holding the reliability bar. · Participate in a 24x7 rotation for production issue escalations, functions well in a fast-paced environment. · Communicate effectively with people at all levels of the organization. Essential Qualifications: · 5+ years of industry experience building and operating large scale distributed systems. · Deep UNIX/Linux systems knowledge and administration background. · Demonstrated programming skills in one or more of: Python, Java, Golang, Ruby. · Strong problem-solving and analytical skills spanning systems, network (TCP/IP) and code, with a focus on data-driven decision-making. · Experience with cloud-based infrastructure, including AWS, Azure, or Google Cloud Platform. · Strong understanding of DevOps and SRE practices, including continuous integration, continuous delivery, and infrastructure as code (IaC). Experience with Terraform is a plus. · Experience with containerization and orchestration technologies, such as Docker and Kubernetes. · Excellent communication and collaboration skills, with the ability to work with teams across distinct functions and technical domains. · Knowledge of observability ecosystem including metrics, logging, tracing and tools, such as Prometheus, Grafana, Elastic Stack, Datadog, or New Relic. Preferred Qualifications: · Bachelor's degree in computer science, engineering, or a related technical field. · Prior experience working with large scale web-based Java architectures and JVM configuration. · Professional certifications in cloud platforms, monitoring tools, or related technologies. · Previous experience working on a large-scale eCommerce platform. Office: Seoul, Korea
Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.
Visa jobbet och ansökUrsprunglig annons: www.coupang.jobs