Principal Machine Learning Engineer

Palo Alto Networks
Santa Clara, CA, United States
1 day ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$163,200.0 - $264,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Data Analysis Software Applications Automated Storage and Retrieval Systems Microsoft Azure Big Data Cyber Security Continuous Integration Software Debugging Distributed Systems
+28 more
Python (Programming Language) Knowledge-Based Systems Machine Learning Open Source Technology Regression Testing Azure Machine Learning Software Deployment Software Engineering SQL Databases Management of Software Versions AI Infrastructure Retrieval-Augmented Generation Large Language Models Multi-Agent Systems Prompt Engineering Software Application Programming Generative AI Backend Agentic-AI Data Layers Build Management Information Technology Low Latency Prompt Injection Virtual Agents Evaluation of Large Language Models Model Inference Docker

Job description

Your Career The AI Canvas team is building our next-generation data exploration experience for cybersecurity. AI Canvas is designed to transform how security teams explore, understand, and act on complex security data by combining natural language interaction, real-time visualization, dashboards, and intelligent agents in a unified experience.

We are building beyond traditional dashboards and static user interfaces. AI Canvas is based on a modern agent- and skill-driven architecture, bringing together AI, data exploration, visualization, and collaboration. We are looking for engineers who are comfortable operating in ambiguity, excited by hard technical problems, and motivated to build reliable AI systems that security teams can trust.

As a Principal Machine Learning Engineer on the AI Canvas team, you will take significant technical ownership of the AI layer powering the platform. You will help define the architecture, build production-grade AI systems, and shape how intelligence is integrated throughout the product.

You will work across large language models (LLMs), retrieval-augmented generation (RAG), AI agents and assistants, agent harnesses, natural-language-to-query generation, evaluation systems, guardrails, and cloud-based AI infrastructure.

This is a highly technical and hands-on role. You will work closely with ML, backend, UI, product, and design teams to solve challenging AI problems in cybersecurity, where reliability, accuracy, scalability, latency, observability, and trust are critical.

Your Impact

  • Provide technical leadership for the architecture and development of the AI capabilities powering AI Canvas, from early design through production deployment and continuous improvement.
  • Design and build scalable, production-grade systems using LLMs, RAG, AI agents, assistants, tools, skills, and agent harnesses.
  • Architect intelligent workflows that enable users to explore complex security data through natural language, including natural-language-to-query generation, follow-up interactions, clarification, investigation, and troubleshooting.
  • Define and evolve the architecture for agentic AI systems, including orchestration, context management, tool invocation, memory, reasoning workflows, and multi-step task execution.
  • Build robust evaluation frameworks and harnesses for measuring AI quality, including correctness, relevance, reliability, regression detection, and end-to-end product behavior.
  • Establish evaluation methodologies using automated metrics, LLM-based evaluators, human evaluation, and representative production datasets.
  • Design and implement guardrails and safety mechanisms to improve reliability, reduce hallucinations, enforce system constraints, and ensure responsible behavior of AI-powered features.
  • Drive improvements in model and system quality through prompt engineering, retrieval strategies, model selection, fine-tuning where appropriate, and systematic experimentation.
  • Build scalable RAG and knowledge-retrieval systems capable of grounding AI responses in large, complex, and evolving security datasets.
  • Partner closely with backend and platform engineers to build reliable APIs, services, and infrastructure supporting AI workloads at production scale.
  • Establish engineering best practices around observability, debugging, tracing, versioning, reproducibility, testing, and monitoring of LLM and agent-based systems.
  • Evaluate emerging models, frameworks, and AI infrastructure and determine when and how they should be incorporated into the product.
  • Balance rapid experimentation with the engineering rigor required to operate mission-critical AI systems in production.
  • Mentor engineers, influence technical direction across teams, and raise the engineering bar for applied AI development.

Requirements

  • Bachelor’s degree in Computer Science, Machine Learning, Engineering, or a related technical field, or equivalent practical experience.
  • 7+ years of software engineering, machine learning engineering, or related industry experience, including significant experience building production systems.
  • Strong experience designing and building machine learning or AI-powered applications at scale.
  • Hands-on experience building applications using large language models and generative AI technologies.
  • Experience with one or more areas such as RAG, AI agents, AI assistants, tool-calling systems, agent orchestration, or LLM-based workflows.
  • Strong proficiency in Python and experience developing production-grade backend or ML services.
  • Experience designing evaluation systems for AI/ML applications, including offline evaluation, regression testing, quality measurement, and production monitoring.
  • Strong understanding of modern ML and AI concepts, including embeddings, retrieval, ranking, prompt engineering, model inference, and experimentation.
  • Experience building scalable systems on public cloud platforms such as GCP, AWS, or Azure.
  • Strong software engineering fundamentals, including system design, distributed systems, APIs, testing, CI/CD, and observability.
  • Ability to lead complex technical initiatives across multiple teams while remaining deeply hands-on.
  • Strong communication skills and the ability to translate ambiguous product problems into clear technical architectures and execution plans.

Preferred Experience

  • Master’s or PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related technical field.
  • Deep experience building and operating LLM-powered products or agentic AI platforms in production.
  • Experience with AI development frameworks and tooling for model orchestration, agent systems, evaluation, tracing, or observability.
  • Experience building natural-language-to-SQL or natural-language-to-query systems, semantic layers, or AI-powered data exploration products.
  • Experience designing LLM evaluation harnesses, synthetic test generation, LLM-as-a-judge approaches, or human-in-the-loop evaluation workflows.
  • Experience with vector databases, search and retrieval infrastructure, embeddings, and large-scale knowledge systems.
  • Experience with containerization and orchestration technologies such as Docker and Kubernetes.
  • Experience designing highly available, low-latency AI inference and backend services.
  • Familiarity with AI security, prompt injection defenses, data privacy, access control, and guardrails for enterprise AI applications.
  • Experience in cybersecurity, security analytics, observability, or large-scale data platforms.
  • Contributions to open-source AI, ML, agent, or infrastructure projects.

Benefits & conditions

The compensation offered for this position will depend on qualifications, experience, and work location. For candidates who receive an offer at the posted level, the starting base salary (for non-sales roles) or base salary + commission target (for sales/com-missioned roles) is expected to be the annual range listed below. The offered compensation may also include restricted stock units and a bonus.

$163,200.00 - $264,000.00/yr

Our Commitment

We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together.

About the company

At Palo Alto Networks®, we’re united by a shared mission-to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.

Who We Are

In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!

We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

7:10 min

Exploring pathways into the machine learning engineering field

Jose Luis Latorre Millas · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:32 min

Refactoring bulk frontend operations into scalable backend methods

Noam Honig · LIVE

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all