> Markdown version of [/jobs/ext/2712463-principal-software-engineer-observability-telemetry-data](https://www.wearedevelopers.com/jobs/ext/2712463-principal-software-engineer-observability-telemetry-data). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Software Engineer - Observability & Telemetry Data - **Company:** Smartsheet Inc. - **Location:** Bellevue, WA, United States (Remote available) - **Experience:** Expert - **Salary:** $222,500.0 - $257,500.0 - **Contract:** Permanent contract - **Skills:** Query Performance, Java (Programming Language), Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Data Analysis, Apache HTTP Server, Big Data, Computer Programming, Data Infrastructure, Data Integrity, Software Debugging, Distributed Systems, Fault Tolerance, Python (Programming Language), Routing, Operational Databases, Prometheus, Data Streaming, Software Technical Review, User-Centered Design, Smartsheet, Datadog, Fluentd, Large Language Models, Snowflake, Grafana, Apache Spark, Backend, Data Lakes, AI Platforms, Kubernetes, Information Technology, AWS Fargate, Apache Kafka, Data Management, Machine Learning Operations, Functional Programming, Cloudwatch, Terraform, Dynatrace, Databricks, Golang - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/principal-software-engineer-observability-telemetry-data-smartsheet-9840385 ## About the Role * 10+ years of experience building and operating large-scale distributed systems, data platforms, or observability infrastructure, including time at Principal or Staff level. * Deep Observability Expertise: Hands-on production ownership of metrics, logs, and distributed tracing at scale, including at least one major backend (Datadog, or comparable) and a working understanding of cardinality and cost mechanics. * Telemetry Data Engineering: Demonstrated depth in large-scale data architecture, including open table formats (Delta Lake, Apache Iceberg), Spark or comparable distributed processing, streaming ingestion, partitioning and schema evolution, and query performance and cost tuning over very large datasets. * Databricks Depth: Practical experience with the Databricks platform (jobs, clusters, Unity Catalog, Delta) and with MLflow for model and agent telemetry. * OpenTelemetry Depth: Practical experience with OTel collectors, semantic conventions, context propagation, and sampling strategy, including tail-based sampling. * Pipeline Engineering: Experience with high-volume log and telemetry pipelines, including FluentBit or Fluentd, streaming transport such as Kinesis or Kafka, and search backend index and mapping design. * Advanced AWS & Kubernetes Expertise: EKS, ECS Fargate, EC2, Lambda, and CloudWatch in production. * 10+ years of programming experience with modern languages such as Go, Java, Python, or Scala, and strong SQL. * Infrastructure as Code: Terraform, and GitOps workflows such as Flux or ArgoCD. * Architectural Influence at Scale: A track record of setting technical direction that multiple teams and organizations adopted, including the written artifacts (architecture decisions, standards, RFCs) that made it durable after you moved on. * Strong incident response instincts, with experience improving mean-time-to-resolution through better instrumentation and better data rather than more heroics. * A degree in Computer Science, Engineering, or a related field, or equivalent practical industry experience. * Legally eligible to work in the U.S. on an ongoing basis, * Experience instrumenting LLM or agentic systems, including OTel GenAI semantic conventions and tracing agent tool-call workflows. * Data reliability engineering practice: freshness, quality, and lineage SLOs for production data pipelines * Telemetry cost engineering or FinOps at scale. * Experience in regulated environments (FedRAMP, GovCloud) and designing telemetry that remains useful under redaction. * Prometheus, Grafana, and Alertmanager. * Snowflake, or experience operating across more than one lakehouse or warehouse platform. * Service catalog or internal developer platform work (Backstage or similar). ## Description * Architect the Telemetry Data Platform: Own the end-to-end design for how telemetry is landed, modeled, and queried, including open table formats, partitioning and schema evolution strategy, separation of storage and compute, tiered retention, and the query interfaces engineers actually use. Make telemetry a durable, portable, Smartsheet-owned dataset rather than a vendor-locked byproduct. * Own the Telemetry Data Model: Define the semantic conventions, shared business identifiers (user, org, plan, tenant), and schema standards that let any signal be correlated with any other, and drive their adoption across every service team at Smartsheet. * Set OpenTelemetry Direction: Lead the migration to OTel-based instrumentation, defining collector architecture, context propagation, and sampling strategy (including tail-based sampling) so that engineers can move from a log line to a trace to a metric without losing the thread. * Integrate AI and Agentic Telemetry with Databricks: Own the architecture connecting our observability platform to Databricks and MLflow, so that agentic and model telemetry (prompt, completion, tool and MCP calls, evaluation results) is captured un-sampled, stays useful under the input/output redaction our governance requires, and reconciles cleanly with the traces and metrics in our primary observability stack. * Instrument the Data Platform Itself: Bring first-class observability to our data estate, including Databricks jobs, pipelines, and warehouses, with meaningful signals for freshness, data quality, lineage, and cost, so data reliability is measured with the same discipline as service reliability. * Build the Analytics Layer on Telemetry: Turn telemetry into decision-grade analytics, covering reliability and incident metrics, telemetry cost and chargeback models, and adoption and coverage reporting that leadership can act on. * Engineer Collection and Routing at Scale: Architect the high-volume collection and routing tier (FluentBit, Kinesis, and OTel collectors) that moves telemetry from every service to its destination across US, EU, AU, and GovCloud regions, and own the migration of these pipelines as we consolidate onto a unified backend. * Own Telemetry Economics: Set the cost architecture for observability data, including ingest governance, cardinality control, and storage tiering, so that teams get the fidelity they need to debug without the spend pressure that causes them to under-instrument. * Carry Architecture Across Org Boundaries: Partner with the Data Platform, AI Platform, and infrastructure organizations to align telemetry architecture with theirs, influence roadmaps you do not own, and represent observability in company-level platform and vendor decisions. * Raise the Technical Bar: Lead design and code reviews, author the architecture decisions and standards others build against, and mentor senior and mid-level engineers on instrumentation, telemetry data modeling, and cost-aware design. * Participate in a production support and on-call rotation, taking ownership of the most complex issue resolution and driving root-cause analysis that improves system resiliency. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Super scaling for the Super Bowl: How to survive 30 million users hitting your backend in 30 minutes](https://www.wearedevelopers.com/videos/100356-super-scaling-for-the-super-bowl-how-to-survive-30-million-users-hitting-your-backend-in-30-minutes) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)