> Markdown version of [/jobs/ext/2706895-observability-platform-engineer](https://www.wearedevelopers.com/jobs/ext/2706895-observability-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Observability Platform Engineer - **Company:** NuScale Power Corporation - **Location:** United States - **Experience:** Expert - **Salary:** $160,000.0 - $230,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, Computer Clusters, Code Review, Software Debugging, Distributed Systems, Python (Programming Language), Reliability Engineering, Ansible, Prometheus, Datadog, Grafana, Kubernetes, Performance Monitor, Apache Kafka, Build Tools, Machine Learning Operations, Vertica, Terraform, Data Pipelines - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-observability-platform-engineer-nscale-9651125 ## About the Role * 5+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles * Experience operating and scaling observability systems in production environments * Strong understanding of monitoring concepts: metrics, logs, traces, alerting, and SLOs * Hands-on experience with several of: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic * Solid programming skills (Python, Go, or similar) with the ability to build and maintain production systems * Experience working with Kubernetes-based infrastructure * Familiarity with Infrastructure-as-Code (Terraform, Ansible, or similar) * Pragmatic mindset with a focus on simplicity, reliability, and maintainability * Strong collaboration skills and ability to work across teams Preferred * Experience with observability data pipelines (Kafka, Vector, Fluent Bit, etc.) * Exposure to AI/ML infrastructure or GPU-based systems * Familiarity with performance monitoring for distributed systems * Experience improving developer experience through observability tooling ## Description As a Senior Observability Platform Engineer, you'll play a key role in designing, building, and scaling Nscale's observability platform. You'll focus on delivering reliable, high-quality visibility into GPU clusters, AI workloads, and the infrastructure that powers them. You approach observability as a product-balancing usability, scalability, and operational efficiency. You build systems that reduce cognitive load for engineers, surface meaningful signals, and enable fast, confident debugging when things go wrong. You'll contribute to platform direction, implement critical systems, and collaborate closely with SRE, infrastructure, and AI/ML teams to ensure observability is embedded into everything we run. This is a hands-on engineering role with meaningful influence over platform design and evolution., * Design, build, and operate scalable observability systems across metrics, logs, traces, and alerting * Contribute to architectural decisions around tooling, data pipelines, storage, and retention strategies * Improve signal quality by reducing noise, managing cardinality, and refining alerting practices * Help identify and address observability gaps before they impact reliability * Partner with SRE, infrastructure, and AI/ML teams to integrate observability into services and platforms * Develop reusable patterns, libraries, and best practices that improve consistency across teams * Participate in incident response and postmortems, driving actionable improvements * Evaluate and adopt tools that improve developer experience, scalability, and operational efficiency * Support and mentor engineers within the team through code reviews and knowledge sharing ## Related Videos - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Easy Mode Monitoring and Logging with Shiftmon](https://www.wearedevelopers.com/videos/2114-easy-mode-monitoring-and-logging-with-shiftmon) - [Better Together: Leveraging Your Observability Tools as a SIEM](https://www.wearedevelopers.com/videos/2118-better-together-leveraging-your-observability-tools-as-a-siem) - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Effortlessly Scale Prometheus With The Telemetry Data Platform – And Keep your Grafana Dashboards, Too!](https://www.wearedevelopers.com/magazine/3-effortlessly-scale-prometheus-with-the-telemetry-data-platform-and-keep-your-grafana-dashboards-too) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)