Principal Observability Platform Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
As a Principal/Staff Observability Platform Engineer, youâll own the technical direction of Nscaleâs observability platform: the systems that give us deep visibility into GPU clusters, AI workloads, and the infrastructure running them. You treat observability as a product and a discipline, not a tooling exercise. Youâll set the architectural roadmap, raise the engineering bar across teams, and ensure our platform scales ahead of the business, not behind it.
You understand that complexity is a cost. Solutions that require constant babysitting donât scale, and neither does operational burden. The platforms you build should be simple to operate, easy to understand, and self-evidently correct when something goes wrong.
This isnât a âmaintain and operateâ role. Itâs a âdefine, build, and leadâ role.
What Youâll Do
- Own the technical strategy and architecture for observability across metrics, logs, traces, and alerting at scale.
- Drive platform decisions that have multi-year impact: tooling, data models, ingestion patterns, retention, cardinality management.
- Identify systemic gaps before they become incidents; design platforms that make failure visible and fast to diagnose.
- Partner with SRE, infrastructure, and AI/ML teams to embed observability natively into how Nscale builds and operates.
- Define standards and patterns that other engineers adopt, not by mandate, but because theyâre clearly better.
- Mentor and technically grow the observability team; raise the ceiling on what the team can build and own.
- Lead incident postmortems and use them to drive durable platform improvements.
- Evaluate and introduce tooling that meaningfully improves signal quality, operational efficiency, or scalability, and retire what doesnât.
Requirements
Do you have experience in Tooling?, * 8+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles.
- Youâve operated observability infrastructure at serious scale. You know what breaks at 10x and you design for it.
- You have a strong bias toward simplicity. Youâve seen over-engineered observability stacks collapse under their own weight and you build accordingly.
- Deep hands-on experience with a significant subset of: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic.
- Strong engineering fundamentals, proficient in Python, Go, or similar; comfortable owning complex systems end to end.
- Experience with Kubernetes at scale; familiarity with GPU infrastructure or HPC environments (Slurm) is a strong plus.
- You can architect systems, write the code, review othersâ work, and explain the tradeoffs clearly, all in the same week.
- Infrastructure-as-Code is default, not optional (Terraform, Ansible, or equivalent).
- You influence without authority. Teams want your opinion because it makes their work better.
Preferred
- Experience with high-volume streaming pipelines for observability data (Kafka, Vector, Fluent Bit, etc.).
- Background in AI/ML infrastructure observability: GPU utilisation, training job visibility, inference latency.
- Prior experience defining observability strategy at an organisation level.
About the company
Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale simplifies AI development while enabling superior results, supporting strategic business outcomes such as cost management, rapid innovation, and environmental responsibility.
We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, youâll build trust through openness and transparency while contributing to the technology that powers the future.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on indeed.comGood distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 120 - Apple and peers
Highest Paying Tech Companies for Developers
MLops â Deploying, Maintaining And Evolving Machine Learning Models in Production
Dev Digest 121 - AI goes offline