Staff Software Engineer(Cloudops)

Palo Alto Networks
Santa Clara, CA, United States
1 day ago
Apply on www.techcareers.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$124,000.0 - $201,500.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Computing Platforms Automation of Tests Microsoft Azure Cloud Computing Cloud Engineering Code Review Continuous Integration Custom Software Distributed Systems Github
+27 more
Identity and Access Management Python (Programming Language) Linux Kernel Machine Learning Node.Js Performance Tuning Reliability Engineering Prometheus Software Engineering TypeScript AWS Cdk Pulumi Google Cloud Cloud Platform System Autoscaling Delivery Pipeline Large Language Models Multi-Agent Systems Prompt Engineering Multi-Cloud Generative AI Gitlab-ci Kubernetes Virtual Agents Asynchronous Programming Api Design Terraform

Job description

  • Multi-Cloud Generative IaC & Software-Defined Infrastructure: Architect and maintain scalable cloud systems across AWS, Azure, and GCP using Pulumi, AWS CDK, or Terraform. Integrate AI development workflows and custom LLM agents to accelerate safe infrastructure compilation, drift detection, and automated cross-cloud refactoring.
  • Intelligent Automation & Agentic Workflows: Engineer custom software utilities, internal services, and autonomous agents using (TypeScript/Node.js, Go, or Python, alongside frameworks like LangChain or CrewAI) to orchestrate complex provisioning, predictive auto-scaling, and closed-loop self-healing systems.
  • AI-Driven Cloud Governance & Economics: Leverage predictive machine learning models to analyze multi-cloud spend patterns, autonomously executing real-time resource-optimization strategies via API-driven software actions (e.g., dynamic spot-instance bidding, intelligent right-sizing across AWS, Azure, and GCP).
  • Cognitive Observability & Infrastructure Security: Implement next-gen observability frameworks (OpenTelemetry, Prometheus) coupled with AI anomaly detection. Embed security directly into the deployment pipeline, utilizing LLMs to automatically audit Cloud IAM policies, scan for vulnerabilities, and generate contextual patches. Intelligent
  • Container Orchestration: Manage production-grade Kubernetes clusters (EKS, AKS, GKE). Optimize resource allocation, cluster auto-scaling, and service meshes using AI-driven traffic routing and predictive capacity planning.
  • Autonomous Incident Response: Act as a tier-3 software escalation engineer for complex distributed systems anomalies. Help design and train our internal “On-Call AI Agent” to ingest logs, perform automated Root Cause Analysis (RCA), and submit pre-validated Pull Requests to resolve underlying system defects.

Requirements

  • Software Engineering & AI Orchestration: Strong software engineering fundamentals in TypeScript (Node.js), Go, or Python. Experience interfacing with LLM APIs (OpenAI, Anthropic, Google Vertex AI, AWS Bedrock), vector databases, and prompt engineering for systems-level orchestration.
  • Multi-Cloud & Containers: Deep proficiency in at least two major cloud platforms (AWS, Azure, GCP) with a strong architectural understanding of the third. Expert-level knowledge of Kubernetes (CKA preferred) and cloud-native networking.
  • Next-Gen CI/CD: Experience building intelligent delivery pipelines using GitHub Actions or GitLab CI, featuring integrated automated testing, security gates, and AI-assisted code reviews.
  • Systems Mastery: Deep understanding of Linux internals, distributed systems architecture, asynchronous programming patterns, and performance tuning., * 6+ years of experience in Cloud Software Engineering, Site Reliability Engineering (SRE), or Distributed Systems Infrastructure.
  • 2+ years of hands-on experience integrating AI tools, LLMs, or predictive analytics into deployment workflows, pipelines, or software platforms.
  • Proven track record of architecting and operating large-scale, high-throughput distributed systems.
  • Soft Skills & Mindset Agentic Problem-Solving: A mindset that moves past “how do I automate this task?” to “how do I build an autonomous system that solves this permanently?”
  • Collaborative AI-First Culture: Ability to partner with Core AI/ML teams to bridge the gap between model deployment and high-availability cloud infrastructure.

Benefits & conditions

The compensation offered for this position will depend on qualifications, experience, and work location. For candidates who receive an offer at the posted level, the starting base salary (for non-sales roles) or base salary + commission target (for sales/com-missioned roles) is expected to be the annual range listed below. The offered compensation may also include restricted stock units and a bonus. A description of our employee benefits may be found here (https://benefits.paloaltonetworks.com/) .

$124,000.00 - $201,500.00/yr

Our Commitment

We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together.

About the company

At Palo Alto Networks, we’re united by a shared mission-to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.

Who We Are

In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!

We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.techcareers.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:35 min

Defining a serverless architecture using AWS CDK

Raphael Manke Raphael Manke · World Congress 2023

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

3:20 min

Overview of infrastructure as code tools

Alexander Bubeck · World Congress 2023

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all