> Markdown version of [/jobs/ext/3020658-principal-data-engineer-llm-ai-platforms-remote](https://www.wearedevelopers.com/jobs/ext/3020658-principal-data-engineer-llm-ai-platforms-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Data Engineer, LLM/AI Platforms (Remote) - **Company:** CrowdStrike - **Location:** Sunnyvale, CA, United States (Remote available) - **Experience:** Experienced - **Salary:** $195,000.0 - $290,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, BigQuery, Data as a Services, Information Engineering, Data Infrastructure, Data Systems, Data Warehousing, Shard (Database Architecture), Distributed Computing Environment, Distributed Systems, Fault Tolerance, Java Virtual Machine (JVM), Python (Programming Language), Open Source Technology, Queueing Systems, DataOps, Large Language Models, Snowflake, Prompt Engineering, Apache Spark, Generative AI, AI Platforms, Kubernetes, Information Technology, Code Testing, Apache Flink, Production Code, Data Analytics, Dask, Apache Kafka, Data Management, Machine Learning Operations, Video Streaming, Oracle Cloud Infrastructure, Data Pipelines, Devsecops, Docker - **Published:** September 21, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/plrdbjlcfa ## About the Role * MLOps Tools (MLflow, Sagemaker, Vertex AI) * Experience with common agentic workflow frameworks (e.g., LangChain, LlamaIndex). * Expert-level proficiency in a high-level coding language (Python, or JVM technologies). * Deep experience with distributed data processing frameworks (e.g., Spark, Dask, Flink). * Strong expertise with cloud platforms (AWS, GCP, or OCI) and related data services. * Containerization and orchestration mastery (Docker, Kubernetes). * Message queuing and streaming technologies (Kafka, Pulsar). * Data Warehousing (Snowflake, BigQuery) and Data Orchestration (Airflow, Kubeflow)., * Master's degree or PhD in Computer Science, Data Engineering, or a related STEM field, or equivalent practical experience. * 10+ years of progressive experience in Data Engineering/Platform Engineering, with at least 3 years focused on architecting and building platforms for AI/ML or Data Science at massive scale. * Demonstrable hands-on experience in LLM engineering (fine-tuning, prompt engineering, deployment), RAG, and developing agentic workflows. * Proven track record of designing and delivering large-scale distributed systems (sharding, partitioning, concurrency). * Exceptional ability to write clean, elegant, performant, and well-tested code, coupled with a proactive mindset for delivering results quickly. * A thorough understanding of engineering practices, including effective peer code reviews, resilient architecture design, and comprehensive testing paradigms. * Prior experience in a Principal or Staff level engineering role, demonstrating technical leadership and mentorship capabilities. * Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes. Bonus Points: * Direct experience building, deploying, and managing LLMs in a production environment. * Prior experience in the cybersecurity, intelligence, or high-compliance industries. * Contributions to open-source projects related to data or AI/ML., * Direct experience building, deploying, and managing LLMs in a production environment. * Prior experience in the cybersecurity, intelligence, or high-compliance industries. * Contributions to open-source projects related to data or AI/ML. ## Description CrowdStrike is looking for a Principal Data Engineer with deep expertise in Large Language Models (LLMs) and AI platforms to join our growing Data Science Platform Engineering Team. You will be a key leader, responsible for designing, building, and deploying cutting-edge data infrastructure that powers our next generation of AI-driven security products. This role requires significant hands-on experience in LLM integration, agentic workflows, and agent harnessing to deliver high-impact, scalable solutions. You will champion engineering excellence, focusing on shipping fast, writing elegant, high-quality code, and actively mentoring and strengthening the team's technical knowledge and capabilities. The scale of our systems and data are approaching Exabytes in size. Experience with extremely large-scale systems, including DevSecOps patterns, practices, and standards are important for this work. What You'll Do: * Architect, implement, and optimize data platforms and pipelines specifically designed to support LLMs, Retrieval-Augmented Generation (RAG), and sophisticated AI agentic systems at Exabyte scale. * Drive the adoption and deployment of agentic workflows and agent harnessing techniques to create autonomous, data-driven security features. * Design and implement highly scalable, fault-tolerant, and cost-effective data solutions, emphasizing rapid iteration and high-quality deployment. * Write elegant, production-ready code with a focus on performance, maintainability, and testing rigor, ensuring the ability to ship fast without compromising quality. * Provide technical leadership and deep expertise in data modeling, normalization, and semantic cataloging for AI/ML workloads. * Establish best practices for MLOps/DataOps surrounding LLMs, including monitoring, observability, and zero-touch recovery mechanisms for AI services. * Actively mentor engineers, conducting technical workshops, leading design reviews, and strengthening the team's knowledge in cutting-edge AI platform technologies. * Collaborate across the organization with Data Scientists, Product Managers, and other engineering teams to transform research prototypes into robust, production-grade services. * Own the end-to-end lifecycle of critical data services: development, testing, deployment, and monitoring. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [13 AI Tools You Have to Try](https://www.wearedevelopers.com/magazine/219-13-ai-tools-you-have-to-try)