> Markdown version of [/jobs/ext/2718635-principal-data-processing-engineer-oss](https://www.wearedevelopers.com/jobs/ext/2718635-principal-data-processing-engineer-oss). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Data Processing Engineer-OSS - **Company:** DATAPELAGO, INC. - **Location:** Mountain View, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** C (Programming Language), Java (Programming Language), Adobe InDesign, Apache HTTP Server, C++ (Programming Language), Code Review, Computer Programming, Continuous Integration, Linux System Administration, Open Source Technology, Performance Tuning, Software Deployment, Data Processing, Apache Spark, Information Technology, Apache Flink, Free and Open-Source Software, Presto - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/principal-data-processing-engineer-oss-datapelago-8256738 ## About the Role * BS/MS in Computer Science (or a related field) with 6+ years of relevant experience * 3+ years of deep technical experience in instrumenting, analyzing, and optimizing the performance of data processing engine components on benchmark and customer workloads. * Sound knowledge of the architecture and internal operation of one or more of Apache Spark, Apache Flink, Presto/Trino. * Demonstrated experience in the design, development, and successful release of high-performance data processing engines for large production deployments. * Exceptional programming skills in C, C++, and Java. * Extensive development experience in Linux environments. * Excellent communication and collaboration skills, with the ability to articulate complex technical concepts to both technical and non-technical audiences. * Strong analytical and problem-solving skills with a passion for performance optimization. ## Description As a Principal Data Processing Engineer (OSS), you will be a key individual contributor in adopting and advancing the capabilities of open-source software (OSS) platforms such as Apache Gluten, Velox, Apache Spark, and Apache Flink in the context of DataPelago's data processing engine. You will enhance the functional breadth, performance, scale, and reliability of the DataPelago engine through downstream and upstream contributions. You will have the opportunity to engage with the community working on these platforms. This is a unique opportunity to make a significant impact on a category-defining product and work with a talented team of engineers. What You'll Do: * Influence the architecture of how our data processing engine interfaces with open-source platforms and engines. * Lead the design of functional and performance enhancements to open source platforms such as Apache Gluten and Velox, and their integration with our data processing engine. * Individually design, implement, test, optimize, and maintain components of the data processing engine. * Analyze the technology roadmap of Apache Gluten, Velox, and equivalent platforms and identify opportunities for our engine to enhance technology and product leadership. * Collaboration: Partner with engineering, product management, the open-source community and customer success teams. * Foster best practices in design and code reviews, testing, CI/CD, and issue resolution to maintain the highest product quality, security, efficiency, and productivity. ## Related Videos - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [Python-Based Data Streaming Pipelines Within Minutes](https://www.wearedevelopers.com/videos/1233-python-based-data-streaming-pipelines-within-minutes) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Building the platform for providing ML predictions based on real-time player activity](https://www.wearedevelopers.com/videos/944-building-the-platform-for-providing-ml-predictions-based-on-real-time-player-activity) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)