> Markdown version of [/jobs/ext/2154088-manager-engineering-platform-development](https://www.wearedevelopers.com/jobs/ext/2154088-manager-engineering-platform-development). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # manager, engineering - platform development - **Company:** Starbucks - **Location:** Seattle, WA, United States - **Experience:** Expert - **Salary:** $105,000.0 - $120,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Computing Platforms, Microsoft Azure, Software Bug Management, Client Server Models, Cloud Computing, Code Generation, Continuous Integration, Programming Tools, Distributed Systems, Perl (Programming Language), Graph Database, Python (Programming Language), Newrelic, Ruby, Software Engineering, Datadog, Cloud Platform System, Large Language Models, Grafana, Prompt Engineering, Multi-Cloud, Generative AI, Information Technology, Terraform, Splunk, Dynatrace, Golang - **Published:** August 20, 2026 - **Apply:** https://www.careerjet.com/job/us5910de5b2e3a60683637f4193a825988/eaa ## About the Role * 5+ years of professional industry experience with software development. * 2+ years of experience leading engineering teams, including direct people-management responsibilities (hiring, performance, and career development). * Bachelor's degree in Computer Science or related field, or equivalent experience. * Understanding of software development lifecycles, developer workflows, and tools. * Experience building platforms, developer tooling, or shared infrastructure consumed by other engineering teams as internal customers. * Experience in setting up observability tools such as Splunk, Datadog, Newrelic for large enterprises * Working knowledge of infrastructure-as-code (e.g., Terraform) and cloud infrastructure (AWS and/or Azure). * Solid understanding of modern AI architectures and techniques, including prompt engineering, retrieval-augmented generation (RAG), and how LLM-based systems are built and operated. * Experience with, or strong familiarity building, agentic AI systems and the tools that support them., * 5+ years of experience managing medium-to-large-scale projects and cross-functional team leadership. * 5+ years of experience in one or more of the following languages: Python, Go, Java, Perl, and/or Ruby. * 3+ years of experience with large-scale distributed systems and client-server architectures. * 3+ years of experience with service resiliency and observability for applications and computing platforms. * Experience building and operating self-service developer platforms, internal developer portals, or paved-path/golden-path tooling at scale. * Deep expertise with infrastructure-as-code (Terraform), CI/CD, and multi-tenant platform design. * Deep understanding of modern observability architectures and techniques, including OpenTelemetry (OTEL) instrumentation and telemetry pipeline design. * Experience implementing SLO/SLI frameworks, ideally within an SRE or reliability-engineering practice supporting large-scale, mission-critical systems. * Passion for improving developer experiences, and a vision for the future of AI-assisted software development. * Experience leading successful advocacy and training programs to drive the adoption of new technologies within a large engineering organization. * Experience with cloud computing platforms (e.g., Amazon AWS and Microsoft Azure), including multi-cloud environments. * Knowledge and understanding of relevant legal and regulatory requirements, such as SOX, PCI, Data Protection, etc. * Demonstrated experience implementing and managing high-capacity, mission-critical environments. * Expertise in observability platforms and tools such as Datadog, Grafana, and OpenTelemetry (OTEL). ## Description From the beginning, Starbucks set out to be a different kind of company. One that not only celebrated coffee and the rich tradition, but that also brought a feeling of connection. We are known for developing extraordinary leaders who share this passion and are guided by their service to others. The Engineering Manager for our GenAI Platform team, will have a strong focus on observability and audit. This leader contributes to Starbucks' success by leading a team of platform engineers in support of our GenAI Accelerator charter across Starbucks Technology. This leader is responsible for delivering high-quality, reliable, and scalable technology and automation solutions that power the platform - spanning our Agent Development Platform, Observability Platform, MCP gateway, and multi-cloud (AWS and Azure) infrastructure. The Engineering Manager of our Platform team is responsible for delivering a high-quality, reliable and scalable technology and automation solutions that power the platform - driving the enterprise-wide observability roadmap and democratizing access to observability metrics across our multi-cloud infrastructure. The successful candidate enjoys bringing definition to complex technology initiatives, is skilled at navigating ambiguous new frontiers, and is an expert at bringing technology and business value to the fore - while growing and developing a high-performing engineering team. As a manager, engineering, you will... Team Leadership * Identifies and communicates key responsibilities and practices to ensure the immediate team of direct reports promotes a successful attitude, confidence in leadership, and teamwork to achieve business results. * Hires, coaches, and develops engineers; sets clear performance expectations, delivers regular feedback, and grows the team's technical and career trajectory. * Provides coaching and mentoring to engineering partners across the organization. Technical Vision & Delivery * Defines and drives the technical vision and roadmap for observability policies that help deliver on our Back to Starbucks mission by ensuring the platforms behind every partner and customer experience stay reliable, performant, and observable. * Spearheads the design, development, and deployment of observability AI agents and tools to address Starbucks' unique observability engineering challenges. This includes but is not limited to: o Knowledge graphs and knowledge agents. o Advanced, context-aware code generation, completion, and refactoring. o AI-powered assistance for event triage, root cause analysis, and bug fixing. o Smart agents to optimize and augment developer testing and workflow automation. * Establishes and evolves the service-level objective (SLO) framework, defining SLIs, error-budget policy, and reliability targets in partnership with developer teams to balance operational stability with delivery velocity. This includes, but is not limited to: o Defining meaningful SLIs and SLOs across critical services and platforms o Operationalizing error budgets to guide prioritization and release decisions o Driving alerting strategy and signal quality to reduce noise and alert fatigue o Partnering with teams to translate reliability data into actionable engineering outcomes * Owns and evolves the incident management lifecycle, from detection and response to root-cause analysis (RCA) and remediation. * Stays at the forefront of observability engineering, particularly in distributed tracing, telemetry pipelines, and OpenTelemetry (OTEL) standards, and identifies opportunities to apply these advancements to Starbucks' challenges. Evaluates and integrates state-of-the-art observability platforms and tooling. * Develops and implements engineering solutions to improve efficiencies, increase capacity, and reduce costs, driving telemetry retention, sampling, and ingestion strategy to control observability spend at scale. Engineering Excellence * Integrates engineering principles and resolves technical obstacles for specific engineering projects. * Applies logic in problem-solving and analysis of alternatives to assess the financial and operational impact of business initiatives. Leverages critical information to make effective recommendations and decisions. * Supports the development team and performs activities to resolve developer issues in a timely and accurate fashion. * Ensures application and infrastructure architectural solutions are stable, secure, and compliant with Company standards and practices. Partnership * Establishes cross-functional, collaborative relationships with business and technology partners. * Communicates highly complex ideas and concepts to non-technical peers and customers. * Partners with external vendors to support and advance the tools and platforms we use. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) - [Empowering Thousands of Developers: Our Journey to an Internal Developer Platform](https://www.wearedevelopers.com/videos/1519-empowering-thousands-of-developers-our-journey-to-an-internal-developer-platform) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)