> Markdown version of [/jobs/ext/2502358-principal-scientist-data-pipeline-engineer](https://www.wearedevelopers.com/jobs/ext/2502358-principal-scientist-data-pipeline-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Scientist - Data Pipeline Engineer - **Company:** Adobe Inc. - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $268,000.0 - **Contract:** Permanent contract - **Skills:** Training Data, Java (Programming Language), C++ (Programming Language), Databases, Information Engineering, Data Infrastructure, Software Debugging, Distributed Data Store, Distributed Systems, Job Scheduling, Python (Programming Language), Machine Learning, Software Engineering, Large Language Models, Apache Spark, Indexer, Data Strategy, Adobe, Data Lakes, Information Technology, Machine Learning Operations - **Published:** August 31, 2026 - **Apply:** https://www.dice.com/job-detail/ee543fce-0f45-40ed-a344-e799aa27b3ed ## About the Role * 10+ years of experience in data engineering, ML infrastructure, or distributed systems, including work at large scale (billions of records or assets) * Strong software engineering background, with hands-on expertise in distributed systems and frameworks such as Ray, Spark, or equivalent large-scale data processing frameworks * Proficiency in Python, plus strong experience in a systems-level language (C++, Rust, Go, or Java) with strong debugging skills across distributed and ML-centric runtime environments. * Deep knowledge of databases and storage systems at scale such as data lakes, indexing, and retrieval across billions of data points * Strong ML background, particularly expertise in optimizing GPU inference pipelines for VLMs, LLMs, or other large models (batching, quantization, serving, throughput/latency tradeoffs) * Experience with data curation for model training: understanding what makes data valuable for training generative or multimodal models, not just how to move it efficiently * Comfort operating across the full stack, from low-level systems and GPU optimization to higher-level data strategy and curation decisions * Ability to communicate clearly and partner effectively across data, infrastructure, and modeling teams * Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Machine Learning, or a related field ## Description We're looking for a Principal ML Engineer to architect and scale the multimodal data processing pipelines and infrastructure behind Adobe Firefly's multimodal foundation models (image, video, audio). In this role, you'll sit at the intersection of data engineering and applied ML building distributed, GPU-accelerated systems that turn billions of raw assets into training-ready data at scale., * Architect and optimize large-scale distributed pipelines that process billions of images, video, and audio assets through ML workflows into training-ready data * Scale up inference throughput across the pipeline (batching, parallelism, hardware utilization) to turn raw collected data into training data faster and more cheaply * Identify and eliminate bottlenecks across ingestion, processing, and delivery, from storage and I/O to compute scheduling ARCHITECT SCALABLE DATA INFRASTRUCTURE * Design systems that reliably store, index, and serve billions of data points, each requiring substantial processing spanning large-scale databases, distributed storage, and high-throughput compute * Apply deep expertise in distributed systems and frameworks such as Ray (or equivalent) to orchestrate large-scale, GPU/CPU-heavy data workloads * Own architecture decisions including database and storage choices, job scheduling, GPU cluster utilization that let the platform scale alongside data and model growth DRIVE DATA CURATION FOR MODEL TRAINING * Bring a strong ML background, especially inference optimization for VLMs and LLMs and data curation for training * Partner closely with modeling teams to understand what data improves training outcomes, and translate that into pipeline and curation requirements * Operate as a hands-on technical leader who bridges data engineering and applied ML ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Optimizing Discovery: PostgreSQL's Role in Transforming GetYourGuide's Search](https://www.wearedevelopers.com/videos/1647-optimizing-discovery-postgresql-s-role-in-transforming-getyourguide-s-search) - [3x Performance: A Humbling Journey](https://www.wearedevelopers.com/videos/100165-3x-performance-a-humbling-journey) - [Developer Experience, Platform Engineering and AI powered Apps](https://www.wearedevelopers.com/videos/990-developer-experience-platform-engineering-and-ai-powered-apps) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1146-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)