> Markdown version of [/jobs/ext/2677126-staff-software-engineer-cloudops](https://www.wearedevelopers.com/jobs/ext/2677126-staff-software-engineer-cloudops). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer(Cloudops) - **Company:** Palo Alto Networks - **Location:** Santa Clara, CA, United States - **Experience:** Experienced - **Salary:** $124,000.0 - $201,500.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Computing Platforms, Automation of Tests, Microsoft Azure, Cloud Computing, Cloud Engineering, Code Review, Continuous Integration, Custom Software, Distributed Systems, Github, Identity and Access Management, Python (Programming Language), Linux Kernel, Machine Learning, Node.Js, Performance Tuning, Reliability Engineering, Prometheus, Software Engineering, TypeScript, AWS Cdk, Pulumi, Google Cloud, Cloud Platform System, Autoscaling, Delivery Pipeline, Large Language Models, Multi-Agent Systems, Prompt Engineering, Multi-Cloud, Generative AI, Gitlab-ci, Kubernetes, Virtual Agents, Asynchronous Programming, Api Design, Terraform - **Published:** September 2, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3373890416&tx=ZT2423TTI&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * Software Engineering & AI Orchestration: Strong software engineering fundamentals in TypeScript (Node.js), Go, or Python. Experience interfacing with LLM APIs (OpenAI, Anthropic, Google Vertex AI, AWS Bedrock), vector databases, and prompt engineering for systems-level orchestration. * Multi-Cloud & Containers: Deep proficiency in at least two major cloud platforms (AWS, Azure, GCP) with a strong architectural understanding of the third. Expert-level knowledge of Kubernetes (CKA preferred) and cloud-native networking. * Next-Gen CI/CD: Experience building intelligent delivery pipelines using GitHub Actions or GitLab CI, featuring integrated automated testing, security gates, and AI-assisted code reviews. * Systems Mastery: Deep understanding of Linux internals, distributed systems architecture, asynchronous programming patterns, and performance tuning., * 6+ years of experience in Cloud Software Engineering, Site Reliability Engineering (SRE), or Distributed Systems Infrastructure. * 2+ years of hands-on experience integrating AI tools, LLMs, or predictive analytics into deployment workflows, pipelines, or software platforms. * Proven track record of architecting and operating large-scale, high-throughput distributed systems. * Soft Skills & Mindset Agentic Problem-Solving: A mindset that moves past "how do I automate this task?" to "how do I build an autonomous system that solves this permanently?" * Collaborative AI-First Culture: Ability to partner with Core AI/ML teams to bridge the gap between model deployment and high-availability cloud infrastructure. ## Description * Multi-Cloud Generative IaC & Software-Defined Infrastructure: Architect and maintain scalable cloud systems across AWS, Azure, and GCP using Pulumi, AWS CDK, or Terraform. Integrate AI development workflows and custom LLM agents to accelerate safe infrastructure compilation, drift detection, and automated cross-cloud refactoring. * Intelligent Automation & Agentic Workflows: Engineer custom software utilities, internal services, and autonomous agents using (TypeScript/Node.js, Go, or Python, alongside frameworks like LangChain or CrewAI) to orchestrate complex provisioning, predictive auto-scaling, and closed-loop self-healing systems. * AI-Driven Cloud Governance & Economics: Leverage predictive machine learning models to analyze multi-cloud spend patterns, autonomously executing real-time resource-optimization strategies via API-driven software actions (e.g., dynamic spot-instance bidding, intelligent right-sizing across AWS, Azure, and GCP). * Cognitive Observability & Infrastructure Security: Implement next-gen observability frameworks (OpenTelemetry, Prometheus) coupled with AI anomaly detection. Embed security directly into the deployment pipeline, utilizing LLMs to automatically audit Cloud IAM policies, scan for vulnerabilities, and generate contextual patches. Intelligent * Container Orchestration: Manage production-grade Kubernetes clusters (EKS, AKS, GKE). Optimize resource allocation, cluster auto-scaling, and service meshes using AI-driven traffic routing and predictive capacity planning. * Autonomous Incident Response: Act as a tier-3 software escalation engineer for complex distributed systems anomalies. Help design and train our internal "On-Call AI Agent" to ingest logs, perform automated Root Cause Analysis (RCA), and submit pre-validated Pull Requests to resolve underlying system defects. ## Related Videos - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Building Reliable Serverless Applications with AWS CDK and Testing](https://www.wearedevelopers.com/videos/812-building-reliable-serverless-applications-with-aws-cdk-and-testing) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [The Software Engineer 2030: From Coder To AI Orchestrator? - Patrick Schnell](https://www.wearedevelopers.com/videos/1825-the-software-engineer-2030-from-coder-to-ai-orchestrator-patrick-schnell) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)