Principal Software Engineer (Cloud Infrastructure and Platform Engineering

Palo Alto Networks
Santa Clara, CA, United States
4 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$156,400.0 - $253,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Software Quality Data Security Machine Learning Regression Testing Software Safety Software Deployment Software Engineering Google Cloud
+18 more
Normalized Discounted Cumulative Gain Retrieval-Augmented Generation Large Language Models Multi-Agent Systems Prompt Engineering Control Structures Multi-Cloud Agentic-AI Build Management AI Platforms Information Technology Palo Alto Networks Prompt Injection Virtual Agents Invoking Functions Model Context Protocol Oracle Cloud Infrastructure Microservices

Job description

We are seeking a rare hybrid practitioner: a Principal AI Systems Engineer & Strategist who bridges the gap between high-level AI strategy and hands-on, production-grade autonomous systems engineering to join our Cloud Infrastructure and Platform Engineering (CIPE) organization. You will be the key architect of our strategy to embed intelligence into every stage of the developer lifecycle-from ideation and documentation to coding, testing, deployment, and observability.

Most enterprise AI initiatives fail not because of model quality, but due to poor use-case selection, weak execution loop design, fragile context management, and an absence of robust governance. This role exists to change that ratio.

In this dual-impact position, you will design the strategic roadmap for autonomous AI investments, architect production-ready harness environments (control loops, state evaluation, dynamic context management, and fallback mechanics), and drive high-value AI agent workflows from early pilot to resilient, production-grade execution.

Your work will directly accelerate developer velocity, elevate our standards for software quality, and unlock new business opportunities by enabling the rapid integration of agentic AI into our products. This role carries executive-level visibility and the autonomy to solve our most complex engineering challenges. If you are a recognized expert in developer platforms and are passionate about leveraging AI to redefine engineering efficiency at a global scale, we want to hear from you. This role is located at our dynamic Santa Clara California headquarters campus, and in office 3 days a week.

Your Impact

  1. Strategic Roadmap & Opportunity Selection * AI Opportunity Mapping: Identify, score, and prioritize enterprise use cases across core business functions, tying each directly to measurable P&L metrics and ROI baselines. * Build vs. Buy vs. Partner Strategy: Evaluate foundation models, framework architectures, and third-party AI platforms to publish clear, defensible architectural and procurement recommendations. * Business Case & Value Realization: Construct multi-year ROI models (acquisition cost, latency overhead, inference cost, payback period) and continuously measure post-deployment business lift. * Risk, Compliance & Governance: Establish guardrails aligned with NIST AI RMF, ISO/IEC 42001, and global regulatory frameworks (e.g., EU AI Act), ensuring data security, model alignment, and threat defense against prompt injection or logic escalation. * Adoption & Change Leadership: Partner with cross-functional leadership to guide organizational change, ensuring pilots transition into core production tools.

  2. Harness Engineering, Control Loops & Agents * Harness & Loopback Architecture: Design and build execution environments (“harnesses”) that wrap foundational reasoning models in closed loopback systems, enabling self-correction, state monitoring, and bounded autonomy. * Context Engineering & MCP: Architect dynamic context windows (retrieved artifacts, short/long-term memory, system instructions, few-shot examples) leveraging protocols like the Model Context Protocol (MCP) to maximize precision while minimizing context rot and token overhead. * Autonomous Agents & Tooling: Build agentic frameworks capable of structured tool selection, multi-agent orchestration, function calling, and deterministic recovery when loops stall or drift. * Advanced RAG Pipelines: Implement hybrid search, vector embeddings, chunking strategies, multi-stage reranking, and agentic retrieval to ground models in proprietary enterprise data.

  3. Evaluation, Production Deployment & Reliability * Continuous Evaluation (Evals): Construct labeled test suites, LLM-as-judge scoring pipelines, retrieval accuracy metrics (MRR, NDCG), and continuous regression testing to measure quality objectively. * Production Operations (LLMOps): Manage latency, multi-tier caching, streaming responses, guardrail enforcement, and fallback routes to maintain SLA targets. * Observability & Diagnostics: Track agent trace logs, tool invocation paths, loop behavior, and cost drivers in real time to catch edge cases before users do.

  4. Other Opportunities for Impact * Drive Organization-Wide Initiatives: You are a builder, so you won’t just stop at ideation. Beyond concepts, ensure your builds show step-change improvements in key engineering metrics like including code velocity, review cycle time, test effectiveness, incident reduction, and overall feature launches. * Lead Cross-Functional Initiatives: Spearhead complex, cross-functional projects that require influencing and aligning multiple engineering organizations and their leadership. * Enable Secure Innovation: Develop foundational AI platforms that empower teams to prototype, deploy, and scale threat-intelligent cloud features, embedding Palo Alto Networks’ security natively. * Innovate at Enterprise Scale: Address intricate challenges in multi-cloud environments (AWS, Azure, GCP, and OCI) supporting thousands of microservices, secure workloads, and global threat detection pipelines.

Requirements

  • 7+ years in Software Engineering / ML Engineering, with 2+ years dedicated specifically to Applied AI, RAG architectures, LLM orchestration, and AI strategy execution.
  • Bachelor’s or Master’s degree in Computer Science, Data Science, Engineering, or an equivalent background of practical industry experience.
  • Proven history of bringing agentic AI solutions into real-world production that achieved clear ROI and survived real usage at scale.
  • System Design over Algorithmic Puzzles: Demonstrated experience solving real-world AI engineering challenges-preventing agent infinite loops, handling context degradation, designing dynamic retrieval, and securing LLM system boundaries.
  • Harnessing & Loopback Expertise: Strong hands-on understanding of autonomous loop mechanics, state verification, step-wise evaluation, and error-recovery harnesses.
  • Tooling & Protocols: Deep practical knowledge of Model Context Protocol (MCP), function calling, agent frameworks, hybrid search systems, and vector databases.
  • Evals & Quality Engineering: Experience replacing anecdotal quality checks with automated, statistically grounded evaluation suites and production monitoring.
  • Business & ROI Fluency: Ability to translate complex model behaviors and inference economics into clear executive business cases.
  • Governance & Security Literacy: Practical knowledge of AI safety, prompt injection defenses, data privacy constraints, and compliance frameworks.
  • Stakeholder Execution: Experience acting as a forward-deployed/applied AI leader, driving alignment across engineering, product, legal, and executive leadership.

Benefits & conditions

The compensation offered for this position will depend on qualifications, experience, and work location. For candidates who receive an offer at the posted level, the starting base salary (for non-sales roles) or base salary + commission target (for sales/com-missioned roles) is expected to be the annual range listed below. The offered compensation may also include restricted stock units and a bonus. A description of our employee benefits may be found here.

$156,400.00 - $253,000.00/yr

Our Commitment

We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together.

About the company

At Palo Alto Networks®, we’re united by a shared mission-to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.

Who We Are

In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!

We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:45 min

Fusing developer experience and platform engineering for agentic SDLC

Julia Kordick Julia Kordick · World Congress 2026 Europe

2:36 min

Choosing between managed AI platforms and custom governance

Péter Farkas Péter Farkas · Europe 2026 Virtual

1:22 min

Overcoming developer challenges in multi-cloud environments

Sandeep Pal Sandeep Pal · Coffee With Developers

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:20 min

Introduction to multi-cloud development infrastructure

Sandeep Pal Sandeep Pal · Coffee With Developers

1:49 min

Augmenting junior and principal engineering roles with AI

Neel Sundaresan Neel Sundaresan +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all