> Markdown version of [/jobs/ext/2588930-principal-software-engineer-observability](https://www.wearedevelopers.com/jobs/ext/2588930-principal-software-engineer-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Software Engineer - Observability - **Company:** Walt Disney Studios - **Location:** New York, NY, United States - **Experience:** Expert - **Salary:** $184,300.0 - $247,100.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Big Data, Software as a Service, Code Review, Continuous Integration, Database Applications, Cursor (Graphical User Interface Elements), Distributed Computing Environment, Github, Machine Learning, Open Source Technology, Performance Tuning, Systems Development Life Cycle, Software Engineering, Systems Integration, Datadog, Flask (Web Framework), Large Language Models, Snowflake, Grafana, Multi-Agent Systems, Prompt Engineering, Model Validation, Caching, Fastapi, Pandas, Build Management, Containerization, AI Platforms, Pyspark, Information Technology, Production Code, Codebase, Api Design, Restful APIs, GPT, Software Version Control, Docker, Databricks, Microservices - **Published:** August 14, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/27931461/Principal-Software-Engineer-Observability-New-York-New-York-1015 ## About the Role * Bachelor's degree in computer science, Engineering, or equivalent experience. * 10+ years of software engineering experience, including building AI-powered or data-driven applications and scalable APIs (e.g. FastAPI, Flask) and deploying production systems at scale. * Proven ability to ramp quickly on unfamiliar codebases and rebuild or modernize systems with AI at significantly higher productivity than a typical engineer, while maintaining quality. Tangible examples where your application of AI has led to 2x to 4x times productivity gains are a strong plus. * Demonstrated expertise in model prompting, context engineering and harnessing, and the design of orchestration patterns and fully autonomous agentic workflows. * Strong AI/ML engineering experience, including orchestrating foundation models (e.g. Claude, OpenAI, Qwen) using frameworks like LangChain or LangGraph. * Track record building and deploying high-quality systems across both front end and back end on hyperscaler cloud platforms (AWS, Azure, or GCP), and integrating frontier model APIs (Anthropic, OpenAI). * Fluency with AI-assisted development tools (e.g., Cursor, Claude Code) to accelerate engineering velocity, and demonstrated experience optimizing AI usage for cost and performance through prompt optimization, caching, and smart model routing. * Experience with modern development practices, including version control (GitHub), containerization (Docker), cloud-native deployments (AWS/EKS), and mature CI/CD pipelines. * Strong understanding of API design, microservices architecture, and standard SDLC workflows. * Track record of working with extreme independence and innovation, scoping and delivering high-impact work that maps directly to business outcomes, and the ability to set technical direction and mentor other engineers. * Strong analytical and technical skills to troubleshoot issues, iterate rapidly, and quickly arrive at viable solutions. * Strong collaboration and communication skills, with the ability to work cross-functionally and clearly explain complex technical concepts to technical and non-technical stakeholders., * Experience building SaaS solutions in the observability space and handling high-volume telemetry data (e.g., Datadog, Grafana, Conviva). Familiarity with OpenTelemetry (OTel) is a bonus. * Experience integrating with enterprise AI services such as AWS Bedrock, including model invocation, routing, and governance integration. * Experience deploying and serving open-source models hosted locally or in private infrastructure, including inference optimization and cost/performance tuning. * Familiarity with large-scale data platforms and distributed data processing tools (e.g., PySpark, Pandas, Databricks, Snowflake). * Knowledge of prompt design, model evaluation, and fine-tuning foundation models (e.g., Claude, GPT). * Experience implementing production-grade systems at scale within a fast-paced, distributed environment. ## Description As a Principal Software Engineer, you are a forward-thinking, highly technical, hands-on-keyboard builder. You work across unfamiliar codebases, teams, and subject matters to accelerate software engineering velocity with AI, shipping production code for our products and applications across the enterprise in a fast-paced, AI-native engineering environment, with a strong understanding of application observability. You will design and build intelligent systems that improve the reliability and performance of Disney's large-scale streaming ecosystem, including fully autonomous agentic systems and real-time pipelines that turn telemetry, logs, and user signals into automated detection, root cause analysis, and proactive insights across Disney+, Hulu, and ESPN. You operate with independence and a strong innovation bias, scoping your own work, making sound tradeoff decisions, and tying everything you build to clear business value and security approvals. You use frontier AI models as a force multiplier to deliver well above typical engineering velocity while keeping token and compute costs rational through optimization and smart model routing. You have a strong desire to work as a technical thought leader setting technical direction and raising the engineering bar through design and code reviews. Responsibilities * Design and operate intelligent, production-grade systems that use real-time signals and AI-driven detection to improve the health of streaming platforms, critical services, and customer experience. * Build and scale fully autonomous agentic systems powered by modern frontier models (e.g. Anthropic, OpenAI, Google, Meta) that reason over complex system behavior, decide and act with minimal human intervention, and drive faster detection and resolution. * Design intelligent harnesses, memory, and multi-agent orchestration patterns - including prompting and context strategies, task decomposition, retrieval, tool routing, and error recovery - that maximize model performance and accuracy and reduce hallucinations across AI-generated context and actions. * Develop end-to-end systems across front end and back end, including data and decisioning pipelines, and scalable APIs that deliver predictive signals, explainability, and insights to engineering teams and product stakeholders. * Keep AI usage cost-effective through prompt and context optimization, caching, batching, and smart model routing across frontier APIs and self-hosted open-source models. * Apply software engineering best practices end to end: clean, well-tested code, thorough code reviews, and mature CI/CD, owning components of production systems. * Drop into brand-new or unfamiliar codebases, rapidly build a working mental model, and modernize them with AI, partnering cross-functionally to embed intelligence into workflows such as incident response, release validation, and customer insights and to drive innovation in observability, reliability, and developer productivity. ## Related Videos - [Developer Experience, Platform Engineering and AI powered Apps](https://www.wearedevelopers.com/videos/990-developer-experience-platform-engineering-and-ai-powered-apps) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [The AI-Ready Stack: Rethinking the Engineering Org of the Future](https://www.wearedevelopers.com/videos/1706-the-ai-ready-stack-rethinking-the-engineering-org-of-the-future) - [Streaming AI Responses in Real-Time with SSE in Next.js & NestJS](https://www.wearedevelopers.com/videos/1630-streaming-ai-responses-in-real-time-with-sse-in-next-js-nestjs) ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)