> Markdown version of [/jobs/ext/582848-principal-observability-platform-engineer](https://www.wearedevelopers.com/jobs/ext/582848-principal-observability-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Observability Platform Engineer - **Company:** NSCALE, LLC - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $150,000.0 - $215,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Clusters, Python (Programming Language), Reliability Engineering, Ansible, Prometheus, High Performance Computing, Grafana, Kubernetes, Apache Kafka, Slurm, Machine Learning Operations, Hardware Infrastructure, Vertica, Terraform - **Published:** June 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=b556223566ad98ae ## About the Role Do you have experience in Tooling?, * 8+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles. * You've operated observability infrastructure at serious scale. You know what breaks at 10x and you design for it. * You have a strong bias toward simplicity. You've seen over-engineered observability stacks collapse under their own weight and you build accordingly. * Deep hands-on experience with a significant subset of: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic. * Strong engineering fundamentals, proficient in Python, Go, or similar; comfortable owning complex systems end to end. * Experience with Kubernetes at scale; familiarity with GPU infrastructure or HPC environments (Slurm) is a strong plus. * You can architect systems, write the code, review others' work, and explain the tradeoffs clearly, all in the same week. * Infrastructure-as-Code is default, not optional (Terraform, Ansible, or equivalent). * You influence without authority. Teams want your opinion because it makes their work better. Preferred * Experience with high-volume streaming pipelines for observability data (Kafka, Vector, Fluent Bit, etc.). * Background in AI/ML infrastructure observability: GPU utilisation, training job visibility, inference latency. * Prior experience defining observability strategy at an organisation level. ## Description As a Principal/Staff Observability Platform Engineer, you'll own the technical direction of Nscale's observability platform: the systems that give us deep visibility into GPU clusters, AI workloads, and the infrastructure running them. You treat observability as a product and a discipline, not a tooling exercise. You'll set the architectural roadmap, raise the engineering bar across teams, and ensure our platform scales ahead of the business, not behind it. You understand that complexity is a cost. Solutions that require constant babysitting don't scale, and neither does operational burden. The platforms you build should be simple to operate, easy to understand, and self-evidently correct when something goes wrong. This isn't a "maintain and operate" role. It's a "define, build, and lead" role. What You'll Do * Own the technical strategy and architecture for observability across metrics, logs, traces, and alerting at scale. * Drive platform decisions that have multi-year impact: tooling, data models, ingestion patterns, retention, cardinality management. * Identify systemic gaps before they become incidents; design platforms that make failure visible and fast to diagnose. * Partner with SRE, infrastructure, and AI/ML teams to embed observability natively into how Nscale builds and operates. * Define standards and patterns that other engineers adopt, not by mandate, but because they're clearly better. * Mentor and technically grow the observability team; raise the ceiling on what the team can build and own. * Lead incident postmortems and use them to drive durable platform improvements. * Evaluate and introduce tooling that meaningfully improves signal quality, operational efficiency, or scalability, and retire what doesn't. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Easy Mode Monitoring and Logging with Shiftmon](https://www.wearedevelopers.com/videos/2114-easy-mode-monitoring-and-logging-with-shiftmon) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Effortlessly Scale Prometheus With The Telemetry Data Platform – And Keep your Grafana Dashboards, Too!](https://www.wearedevelopers.com/magazine/3-effortlessly-scale-prometheus-with-the-telemetry-data-platform-and-keep-your-grafana-dashboards-too)