> Markdown version of [/jobs/ext/2000189-principal-kafka-site-reliability-engineer-devops](https://www.wearedevelopers.com/jobs/ext/2000189-principal-kafka-site-reliability-engineer-devops). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Kafka Site Reliability Engineer (DevOps) - **Company:** Palo Alto Networks - **Location:** Santa Clara, CA, United States - **Contract:** Permanent contract - **Skills:** Application Frameworks, Computing Platforms, Automation of Tests, Big Data, Cloud Computing, DevOps, Distributed Systems, Hadoop Distributed File System, Python (Programming Language), Performance Tuning, Reliability Engineering, Apache Zookeeper, Palo Alto Networks, Data Analytics, Apache Kafka, Build Tools, Malware Detection - **Published:** August 9, 2026 - **Apply:** https://www.wayup.com/i-j-palo-alto-networks-215343898837745/ ## About the Role + Hands on experience with managing production Kafka clusters. + Strong development/automation skills. Must be very comfortable with reading and writing Python. Commits to Kafka source code would be a big plus. + In-depth understanding of the internals of Kafka cluster management, Zookeeper, partitioning, topic replication and mirroring. + Very good grasp of monitoring and metrics collection, performance tuning, and troubleshooting complicated situations with distributed systems. + Tools-first mindset. You build tools for yourself and others to increase efficiency and to make hard or repetitive tasks easy and quick. + Organized, focused on building, improving, resolving and delivering. Good communicator in and across teams, great teamwork, and a character of taking ownership. ## Description We are reshaping the cybersecurity market through our cloud-delivered security services, and our cloud infrastructure is quickly and massively growing with a global footprint. We're looking for great SREs, as well as software engineers interested in production engineering, to help us scale the largest enterprise security cloud infrastructure in the world. Description Palo Alto Networks reinvented the enterprise firewall, growing from a start-up to a multi-billion-dollar company. Our Application Framework, the latest offering in our cloud-delivered security services, ingests security events from hundreds of thousands of firewalls deployed across the globe to provide a massive data analytics platform for deep inspection, anomaly detection, and actionable security automation. Our cloud infrastructure is home to a series of massive and complicated distributed systems and virtualization software platforms which enable big data processing around security services, sandboxing and malware detection, URL categorization and malicious site/domain identification, and security research/response., + You will be responsible for maintaining and scaling production Kafka clusters with very high ingestion rates, Zookeeper clusters, as well as other big data pipeline systems such as Kafka and HDFS. + You will improve scalability, service reliability, capacity, and performance. + You will write automation code for managing, monitoring, measuring, expanding, and healing clusters. + You are not an operator, you're an experienced software engineer focused on operations. + You will do Kafka tuning, capacity planning, and deep dive troubleshooting. + You will participate in the occasional on-call rotation supporting the infrastructure. + You will roll up the sleeves to troubleshoot incidents, formulate theories and test your hypothesis, and narrow down possibilities to find the root cause. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Let's Get Started With Apache Kafka® for Python Developers](https://www.wearedevelopers.com/videos/565-let-s-get-started-with-apache-kafka-for-python-developers) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)