> Markdown version of [/jobs/ext/532354-senior-manager-engineering-observability-platform-remote-eligible](https://www.wearedevelopers.com/jobs/ext/532354-senior-manager-engineering-observability-platform-remote-eligible). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Manager, Engineering - Observability Platform (Remote Eligible) - **Company:** Smartsheet Inc. - **Location:** Bellevue, WA, United States (Remote available) - **Experience:** Expert - **Salary:** $205,000.0 - $275,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Software as a Service, Data Infrastructure, Distributed Systems, Elasticsearch, Smartsheet, Datadog, Cloud Platform System, Large Language Models, Multi-Agent Systems, Mttr, Backend, AI Platforms, AWS Fargate, Machine Learning Operations, Dynatrace - **Published:** June 14, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=3456f2afbb94805c ## About the Role Do you have experience in Vendor relationship management?, * 10+ years of software or platform engineering experience, with strong fundamentals in distributed systems, infrastructure, and backend services. * 3 years of engineering management experience, including direct team building, performance management, and cross-functional delivery ownership. * Deep hands-on expertise with observability tooling: Datadog (APM, metrics, logs, alerting), OpenSearch or Elasticsearch, distributed tracing (OpenTelemetry or equivalent), and SLO/SLA management at scale. * Proven experience operating observability platforms for high-availability, high-throughput production environments. * Experience building and scaling engineering teams in distributed or international focus * Strong execution track record on complex, cross-functional infrastructure programs with high ambiguity. * Clear, direct communication (written and verbal) with both technical and non-technical audiences, including leadership and executive stakeholders. * Proactive risk identification and status communication without prompting. * Experience managing vendors, external delivery partners, and third-party integrations in a platform context. Preferred * Hands-on experience with AI/ML observability: MLflow tracing, LLM evaluation pipelines, or observability for agentic AI systems. * Familiarity with Amazon Bedrock, ECS Fargate, or LangGraph-based multi-agent architectures. * Experience with cloud cost governance and FinOps practices for observability tooling * Exposure to data platform observability and data quality monitoring in a lakehouse context * Experience establishing internal developer platforms, shared libraries, or platform-as-a-service offerings for application teams. * Prior work in SaaS environments with enterprise compliance requirements (SOC 2, FedRAMP, HIPAA). Education & Eligibility * CS, Engineering, or equivalent degree, or commensurate practical experience. * Legally eligible to work in the U.S. on an ongoing basis ## Description * Lead a team of engineers focused on observability platform engineering, driving build-out of a unified observability stack used by all engineering teams at Smartsheet. * Own and evolve the platform's technical roadmap, consolidating multiple tooling platforms, and AI observability tooling into a coherent, scalable capability. * Define platform standards, contribute to architectural direction, and ensure the team operates with engineering rigor and strong operational habits. * Build and scale the team, hiring senior engineers and establishing effective global practices across distributed stakeholders. Observability Engineering * Lead design and delivery of centralized observability infrastructure covering metrics pipelines, distributed tracing, alerting frameworks, and log analytics across Smartsheet services. * Drive SLO/SLA definition and tooling for platform-wide reliability visibility, partnering closely with infrastructure, platform engineering, and on-call teams. * Own governance including instrumentation standards, cost optimization, and rollout of advanced capabilities such as APM, RUM, and custom dashboards. * Lead architecture, scaling, and operational practices for log analytics across high-throughput production workloads. * Establish shared observability libraries, agents, and SDKs that reduce instrumentation burden for application engineering teams. AI Observability * Build and maintain AI/ML observability integrations in partnership with the AI Platform team. * Partner with the Data & AI Platform team to integrate MLflow tracing, Inference Tables, and LLM-as-judge evaluation pipelines into the observability stack. * Develop dashboards and alerting for agentic AI workloads, including latency, token consumption, error rates, and evaluation metric drift. * Contribute to the AI governance and cost observability program, providing telemetry for model usage, cost attribution, and compliance reporting. Cross-Functional Partnership & Execution * Serve as the primary engineering partner for platform consumers across Data & AI, Commerce, Infrastructure, and Security teams, ensuring observability needs are met across workstreams. * Lead complex, cross-functional observability projects with high ambiguity, managing delivery risk, communicating clearly to senior stakeholders, and building alignment across teams. * Partner with delivery partners to coordinate instrumentation across platform modernization and migration workstreams * Contribute to quarterly and annual platform goals, reporting on key reliability and observability metrics to engineering leadership. * Communicate platform status, risks, and roadmap progress to Engineering leadership and above audiences in a clear, executive-ready format. Operational Excellence * Embed on-call culture and incident management discipline into the team, ensuring clear runbooks, fast MTTR, and post-incident learning loops. * Drive cost governance for observability tooling, including spend optimization and efficient resource management. * Champion AI-assisted engineering practices within the team, applying tooling and automation to reduce toil and accelerate delivery. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)