> Markdown version of [/jobs/ext/2697982-principal-engineer-software-dev-production-infrastructure](https://www.wearedevelopers.com/jobs/ext/2697982-principal-engineer-software-dev-production-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Engineer Software, Dev & Production Infrastructure - **Company:** Palo Alto Networks - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $147,000.0 - $237,500.0 - **Contract:** Temporary contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, C++ (Programming Language), Cloud Engineering, Software Debugging, Programming Tools, Distributed Systems, OSI Models, Python (Programming Language), Network Protocols, Open Source Technology, Systems Development Life Cycle, Software Systems, Systems Architecture, Google Cloud, Load Balancing, Large Language Models, Multi-Agent Systems, Reliability of Systems, Technical Debt, Infrastructure as Code (IaC), Backend, Kubernetes, Palo Alto Networks, Terraform, Dynatrace, Golang, Microservices - **Published:** September 3, 2026 - **Apply:** https://www.themuse.com/jobs/paloaltonetworks/principal-engineer-software-dev-production-infrastructure-chronosphere ## About the Role * Experience: 8-10+ years of relevant experience in high-scale infrastructure, production engineering, systems architecture, or developer platform roles. * Language Agnosticism: 7 + years of experience in at least one backend language (e.g., Go, Python, Java, C++, or Rust). We value modular, testable code over specific syntax knowledge, though Go is our primary language. * Cloud-Native & Distributed Systems: 5 + years of deep expertise building and debugging systems that deal with CAP theorem trade-offs, eventual consistency, and distributed tracing. Robust experience working with AWS or GCP and navigating Kubernetes clusters. * Systems Expertise: Deep knowledge of Linux internals, process management, resource isolation, and networking protocols (OSI model, load balancing, service meshes). * Execution & Quality: A proven track record of estimating complex work effectively, debugging distributed architectures via logs/traces, and proactively identifying design edge cases., * Commitment to AI-Native Engineering: Experience or active interest in using next-gen AI coding assistants and agentic platforms to accelerate engineering velocity and eliminate boilerplate work. * Experience with AI-enabled agentic development workflows (e.g., building, deploying, or developing alongside autonomous AI agents, multi-agent frameworks, and LLM-driven orchestration tools). * Active engagement or experience with open-source communities. ## Description Chronosphere, a Palo Alto Networks company, is the observability platform built for control in the modern, containerized world. Chronosphere empowers customers to focus on the data and insights that matter by reducing complexity, optimizing costs, and remediating issues faster. Chronosphere reduces data volumes and associated costs by 84% on average while saving developers thousands of hours. Recognized as a leader by major analyst firms, Chronosphere is trusted by the world's most innovative brands, including DoorDash, Affirm, and Zillow. One Org, Two Impact Areas: Our Infrastructure organization is scaling rapidly, and we are looking for world-class Principal Engineers to drive the future of our platform. We are actively hiring for two distinct pillars within the same org and will consider all applicants for both tracks. During our unified interview process, we will work closely with you to determine which team-or combination of responsibilities-aligns best with your technical strengths and career aspirations. The Two Areas You Will Be Considered For: Production Engineering (Reliability & Scale) * The Core Focus: Ensuring all live components of our global platform are optimized, highly reliable, and battle-tested. This role focuses on pushing our massive distributed cloud resources to their absolute limits., + Solve complex distributed systems problems at the largest scales in the world. + Scale our production infrastructure and architecture globally using Kubernetes, GCP, and AWS. + Use Go to build and operate highly scalable, resilient backend services. + Develop the internal tools necessary to monitor, alert, and optimize our live production footprint. Dev Infrastructure (Developer Velocity & Platforms) * The Core Focus:Building the platform and tooling that enables developer velocity and software reliability for the entire Chronosphere engineering organization. This team owns the full end-to-end SDLC., + Architect & Build:Design high-scale developer tooling, dynamic testing environments, and CI/CD pipelines in a 100% modern, containerized microservices ecosystem. + Infrastructure as Code (IaC): Treat infrastructure as a first-class citizen by defining and managing entire ephemeral environments using declarative IaC (Terraform). + Drive Systemic Quality: Identify and eliminate systemic bottlenecks across the development lifecycle through architectural changes, advanced tooling, and near-real-time telemetry processing. Combined Principal-Level Leadership Regardless of which pillar you align with most, as a Principal Engineer you will: * Own massive technical initiatives from inception to delivery, balancing feature velocity with long-term technical debt. * Define platform standards and reference architectures that span a 1-3 year horizon. * Act as the "glue" across the organization, consulting on infrastructure best practices and up-leveling the team through dedicated mentorship. ## Related Videos - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [The AI-Native Engineering Org: What’s Real, What’s Hype, What’s Next](https://www.wearedevelopers.com/videos/100004-the-ai-native-engineering-org-what-s-real-what-s-hype-what-s-next) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai)