> Markdown version of [/jobs/ext/2708329-staff-observability-software-engineer](https://www.wearedevelopers.com/jobs/ext/2708329-staff-observability-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Observability Software Engineer - **Company:** Coupang, Inc. - **Location:** Seattle, United States - **Experience:** Expert - **Salary:** $299,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon Web Services, Data Analysis, Application Performance Management, Microsoft Azure, Continuous Delivery, Continuous Integration, DevOps, Distributed Systems, Elasticsearch, Monitoring of Systems, Python (Programming Language), Site Reliability Engineering Practices, Prometheus, Ruby, Data Logging, Google Cloud, Cloud Platform System, System Availability, Grafana, Infrastructure as Code (IaC), Containerization, Kubernetes, Information Technology, Appdynamics, Dynatrace, Docker, Golang, Programming Languages - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/sr-staff-observability-software-engineer-coupang-8922630 ## About the Role * Bachelor's degree in Computer Science, Electrical Engineering, Math, or a closely related field * 8 years working experience in engineering teams that build large-scale distributed systems * Strong experience in implementing and managing observability solutions in large-scale, complex environments. * Deep knowledge of monitoring, alerting, and logging systems and tools, such as OpenTelemetry, Prometheus, Grafana stack, Elastic stack., * Experience with containerization and orchestration technologies, such as Docker and Kubernetes. * Familiarity with application performance management (APM) tools, such as Dynatrace or AppDynamics. * Professional certifications in cloud platforms, monitoring tools, or related technologies. * Proficiency in at least one programming languages, such as Go, Java, Python, or Ruby. * Familiarity with distributed tracing technologies, such as Jaeger or Zipkin. * Experience with cloud-based infrastructure, including AWS, Azure, or Google Cloud Platform. * Strong understanding of DevOps and SRE practices, including continuous integration, continuous delivery, and infrastructure as code (IaC). * Strong problem-solving and analytical skills, with a focus on data-driven decision-making. * A proven track record of leading and delivering successful observability projects and initiatives. ## Description You will be a part of the Observability Engineering team at Coupang to build and maintain monitoring solutions and application performance solutions. The Observability Engineering team's vision is to enable engineering teams to maintain and improve the quality of their services by providing clear visibility about their systems health, intelligence and actionable insights. We build services and tools to help detect, alert, and troubleshoot systems and applications anomalies, with metrics, logs and traces related to the health and reliability of Coupang services. You will have the opportunity to work with hyper-scale systems, build the next generation Observability Platform based on Kubernetes and other OSS solutions, as well as building software components from scratch. You would work directly with various engineering teams in Coupang, influence them with observability principles and best practices and see your impact directly. What You Will Do * Design, implement, and maintain observability solutions such as monitoring, alerting, logging, and tracing across various platforms, applications, and infrastructure. * Collaborate with cross-functional teams, including software engineers, SREs, and infrastructure teams, to identify and define observability requirements. * Develop and implement best practices for creating and maintaining effective monitoring, alerting, and telemetry systems. * Evaluate and recommend industry-leading observability tools and technologies to improve system visibility and reliability. * Define and track key performance indicators (KPIs) and service-level objectives (SLOs) related to system availability, performance, and reliability. * Assist in the troubleshooting and resolution of complex incidents and problems by analyzing data from observability tools. * Provide guidance and mentorship to other engineers on observability principles, practices, and tools. * Conduct ongoing evaluations of observability systems and identify opportunities for improvements and optimizations. * Drive the standardization and simplification of observability processes, tools, and frameworks across the organization. * Contribute to the development of training materials, documentation, and runbooks for observability systems and practices., * Application Review - Phone Interview - Onsite (or Virtual Onsite) Interview - Offer * The exact nature of the recruitment process may vary according to the specific job and may be changed due to scheduling or other circumstances. * Interview schedules and the results will be informed to the applicant via the e-mail address submitted at the application stage. Details to Consider * This job posting may be closed prior to the stated end date for application if all openings are filled. * Coupang has the right to rescind an offer of employment if a candidate is found to have submitted false information as part of the application process. * Those eligible for employment protection (recipients of veteran's benefits, the disabled, etc.) may receive preferential treatment for employment in accordance with applicable laws. ## Related Videos - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)