> Markdown version of [/jobs/ext/2117100-senior-software-engineer-observability-irm](https://www.wearedevelopers.com/jobs/ext/2117100-senior-software-engineer-observability-irm). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer - Observability & IRM - **Company:** The Trade Desk - **Location:** Boulder, CO, United States - **Experience:** Expert - **Salary:** $124,900.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Software Debugging, Programming Tools, Distributed Systems, Prometheus, Web Applications, Data Logging, Grafana, Kubernetes, Sumo Logic (Software) - **Published:** August 19, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/17994684?backUrl=%2Fcareer%2F17994684%2FSenior-Software-Engineer-Observability-Irm-Colorado-Boulder ## About the Role * Experience building and operating production infrastructure or internal developer tooling * Comfort working across the stack - this role touches distributed systems, Kubernetes, observability pipelines, and web-based tooling * Familiarity with observability concepts: logging, alerting, on-call workflows * Strong debugging instincts: You will be expected to be called on when things break * Clear communication: The team works closely with engineers across the company; you'll need to explain tradeoffs and advocate for solutions Plus skills: * Experience with Grafana, Prometheus, or similar observability tools * Familiarity with Sumo Logic or other log management platforms * Prior work on developer portals or service catalog tooling (Backstage, OpsLevel, etc.) * Experience with Kubernetes at scale * A deep understanding of HunnyPt ## Description The Service Excellence (SE) team owns the tools and infrastructure that help engineers at The Trade Desk understand and operate production systems. The Incident Response Services (IRS) taskforce focuses on the on-call experience. The team is responsible for making incidents easier to detect, manage, and optimize using historical data points information. What you will work on: * Incident management tooling * Build and maintain automation around the incident lifecycle: alerting, escalation, incident channels, retros, and SLA tracking * Help evaluate and migrate our logging stack * Participate in the re-evaluation of our logging vendor and collection architecture * Backstage/Service catalog - Extend our internal developer portal with K8s integrations, maturity models, and SLO adoption tooling * Alert quality tooling - Build the systems that give engineers better signal and less noise - smarter routing, better grouping, tighter feedback loops between alerts and the teams that own them, Variety of technical opportunity is one of the best things about working at The Trade Desk as a software engineer which is why we do not expect you to know every technology we use when you start. What we care about is that you can learn quickly and find solutions to complex problems using the optimum tools for the job. What you know is less important than how well you learn and innovate. We don't need engineers who know all the answers; we need engineers who can invent the answers no one has thought of yet, to the questions yet to be asked. ## Related Videos - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Data binning and understanding histograms](https://www.wearedevelopers.com/videos/2086-data-binning-and-understanding-histograms) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Microservices: how to get started with Spring Boot and Kubernetes](https://www.wearedevelopers.com/videos/242-microservices-how-to-get-started-with-spring-boot-and-kubernetes) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [WeAreDevelopers Dev Digest Issue 116 - The new search wars…](https://www.wearedevelopers.com/magazine/445-wearedevelopers-dev-digest-issue-116-the-new-search-wars) - [Dev Digest 119 - ❤️ === ❤️](https://www.wearedevelopers.com/magazine/454-dev-digest-119) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)