> Markdown version of [/jobs/ext/3570595-principal-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/3570595-principal-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Machine Learning Engineer - **Company:** Palo Alto Networks - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $163,200.0 - $264,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Data Analysis, Software Applications, Automated Storage and Retrieval Systems, Microsoft Azure, Big Data, Cyber Security, Continuous Integration, Software Debugging, Distributed Systems, Python (Programming Language), Knowledge-Based Systems, Machine Learning, Open Source Technology, Regression Testing, Azure Machine Learning, Software Deployment, Software Engineering, SQL Databases, Management of Software Versions, AI Infrastructure, Retrieval-Augmented Generation, Large Language Models, Multi-Agent Systems, Prompt Engineering, Software Application Programming, Generative AI, Backend, Agentic-AI, Data Layers, Build Management, Information Technology, Low Latency, Prompt Injection, Virtual Agents, Evaluation of Large Language Models, Model Inference, Docker - **Published:** October 3, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=7a04988ac8f29104 ## About the Role * Bachelor's degree in Computer Science, Machine Learning, Engineering, or a related technical field, or equivalent practical experience. * 7+ years of software engineering, machine learning engineering, or related industry experience, including significant experience building production systems. * Strong experience designing and building machine learning or AI-powered applications at scale. * Hands-on experience building applications using large language models and generative AI technologies. * Experience with one or more areas such as RAG, AI agents, AI assistants, tool-calling systems, agent orchestration, or LLM-based workflows. * Strong proficiency in Python and experience developing production-grade backend or ML services. * Experience designing evaluation systems for AI/ML applications, including offline evaluation, regression testing, quality measurement, and production monitoring. * Strong understanding of modern ML and AI concepts, including embeddings, retrieval, ranking, prompt engineering, model inference, and experimentation. * Experience building scalable systems on public cloud platforms such as GCP, AWS, or Azure. * Strong software engineering fundamentals, including system design, distributed systems, APIs, testing, CI/CD, and observability. * Ability to lead complex technical initiatives across multiple teams while remaining deeply hands-on. * Strong communication skills and the ability to translate ambiguous product problems into clear technical architectures and execution plans. Preferred Experience * Master's or PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related technical field. * Deep experience building and operating LLM-powered products or agentic AI platforms in production. * Experience with AI development frameworks and tooling for model orchestration, agent systems, evaluation, tracing, or observability. * Experience building natural-language-to-SQL or natural-language-to-query systems, semantic layers, or AI-powered data exploration products. * Experience designing LLM evaluation harnesses, synthetic test generation, LLM-as-a-judge approaches, or human-in-the-loop evaluation workflows. * Experience with vector databases, search and retrieval infrastructure, embeddings, and large-scale knowledge systems. * Experience with containerization and orchestration technologies such as Docker and Kubernetes. * Experience designing highly available, low-latency AI inference and backend services. * Familiarity with AI security, prompt injection defenses, data privacy, access control, and guardrails for enterprise AI applications. * Experience in cybersecurity, security analytics, observability, or large-scale data platforms. * Contributions to open-source AI, ML, agent, or infrastructure projects. ## Description Your Career The AI Canvas team is building our next-generation data exploration experience for cybersecurity. AI Canvas is designed to transform how security teams explore, understand, and act on complex security data by combining natural language interaction, real-time visualization, dashboards, and intelligent agents in a unified experience. We are building beyond traditional dashboards and static user interfaces. AI Canvas is based on a modern agent- and skill-driven architecture, bringing together AI, data exploration, visualization, and collaboration. We are looking for engineers who are comfortable operating in ambiguity, excited by hard technical problems, and motivated to build reliable AI systems that security teams can trust. As a Principal Machine Learning Engineer on the AI Canvas team, you will take significant technical ownership of the AI layer powering the platform. You will help define the architecture, build production-grade AI systems, and shape how intelligence is integrated throughout the product. You will work across large language models (LLMs), retrieval-augmented generation (RAG), AI agents and assistants, agent harnesses, natural-language-to-query generation, evaluation systems, guardrails, and cloud-based AI infrastructure. This is a highly technical and hands-on role. You will work closely with ML, backend, UI, product, and design teams to solve challenging AI problems in cybersecurity, where reliability, accuracy, scalability, latency, observability, and trust are critical. Your Impact * Provide technical leadership for the architecture and development of the AI capabilities powering AI Canvas, from early design through production deployment and continuous improvement. * Design and build scalable, production-grade systems using LLMs, RAG, AI agents, assistants, tools, skills, and agent harnesses. * Architect intelligent workflows that enable users to explore complex security data through natural language, including natural-language-to-query generation, follow-up interactions, clarification, investigation, and troubleshooting. * Define and evolve the architecture for agentic AI systems, including orchestration, context management, tool invocation, memory, reasoning workflows, and multi-step task execution. * Build robust evaluation frameworks and harnesses for measuring AI quality, including correctness, relevance, reliability, regression detection, and end-to-end product behavior. * Establish evaluation methodologies using automated metrics, LLM-based evaluators, human evaluation, and representative production datasets. * Design and implement guardrails and safety mechanisms to improve reliability, reduce hallucinations, enforce system constraints, and ensure responsible behavior of AI-powered features. * Drive improvements in model and system quality through prompt engineering, retrieval strategies, model selection, fine-tuning where appropriate, and systematic experimentation. * Build scalable RAG and knowledge-retrieval systems capable of grounding AI responses in large, complex, and evolving security datasets. * Partner closely with backend and platform engineers to build reliable APIs, services, and infrastructure supporting AI workloads at production scale. * Establish engineering best practices around observability, debugging, tracing, versioning, reproducibility, testing, and monitoring of LLM and agent-based systems. * Evaluate emerging models, frameworks, and AI infrastructure and determine when and how they should be incorporated into the product. * Balance rapid experimentation with the engineering rigor required to operate mission-critical AI systems in production. * Mentor engineers, influence technical direction across teams, and raise the engineering bar for applied AI development. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)