ML Data Platform Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+1 more
Job description
We are looking for an ML Data Platform Engineer to make the data behind our models and products dependable, understandable, and easy to use. This role sits where data engineering meets machine learning. You will turn messy, changing real-world sources into durable datasets and interfaces that researchers, product engineers, and customer-facing technical teams can trust. The goal is not to build a large platform for its own sake. It is to make each new model, product surface, and authorized data source faster to bring online without compromising correctness. What You Will Own
- Build and improve ingestion, backfill, validation, and observability for high-volume, time-dependent data.
- Define clear data contracts and point-in-time semantics for model training, evaluation, and product use.
- Create reusable workflows for bringing new public and customer-authorized sources into the system.
- Build quality, lineage, freshness, and access controls that make data trustworthy in repeated use.
- Develop efficient datasets and query interfaces for machine-learning and product workloads.
- Diagnose whether failures originate in source data, transformations, model inputs, or serving systems.
- Work closely with research and product engineers so data requirements become reliable software, not recurring manual projects.
-
Exercise judgment about which abstractions should become durable infrastructure and which should remain purpose-built. First 90 Days
- 30 days: Understand the data lifecycle behind Ask The Grid and our ML work; identify the most consequential reliability and usability gaps.
- 60 days: Ship a reusable ingestion, backfill, validation, or dataset primitive used in active product or research work.
- 90 days: Own a dependable end-to-end data workflow, with documented contracts, quality checks, and clear operational visibility., Position Title: Data Platform Engineer Databricks Location Seattle / NYC (hybrid work) Responsibilities Design, build, and enhance platform capabilities within Databricks an…
- 16 hours ago
- Apply easily
Requirements
You May Be A Fit If
- You enjoy making difficult real-world data useful, not merely moving it between systems.
- You understand how warehouse or streaming data becomes training, evaluation, and product data.
- You care about temporal correctness, reproducibility, leakage, lineage, and source rights.
- You can design practical systems without reaching immediately for a large-company platform.
- You are comfortable debugging incomplete APIs, changing schemas, and surprising data behavior.
- You communicate clearly with researchers, product engineers, and customer-facing teammates.
-
You want substantial ownership on a small team and can make progress without a mature data organization around you. Helpful Background
- Strong Python and SQL experience.
- Experience with data engineering, ML data systems, dataset engineering, platform engineering, or high-quality analytics engineering.
- Experience with object storage, warehouses, streaming or workflow systems, columnar formats, APIs, and data-quality tooling.
- Experience preparing data for model training, evaluation, or scientific computing.
- Familiarity with time-series, geospatial, weather, market, event, or other temporally sensitive data is useful.
- Startup or small-team experience is helpful, but evidence of unusually strong ownership matters more than a particular company background.
Benefits & conditions
Location And Working Style New York City or Boston/Cambridge. We work together in person regularly and will choose the home office based on the strongest candidate and their closest collaborators. Compensation And Benefits Base salary range: $175K-$245K, plus meaningful early-stage equity. Parisi Labs provides medical and dental benefits. Final compensation depends on level, location, experience, and role scope. Interview Process
- Founder conversation with the CTO.
- Technical working session on a representative data-platform problem.
- Collaboration conversation with a research or product teammate.
- In-person final.
- Offer review.
About the company
Parisi Labs is an AI company building learning systems for complex physical environments. We combine historical and live data with real operational context to help people understand the present, evaluate possible futures, and make better decisions. Energy is our first proving ground. Ask The Grid (https://askthegrid.com) is our public product for exploring the systems, markets, and assets that make up the power grid. We are a small technical team working across machine learning, data infrastructure, software, and real-world operations.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are Large Language Models?
How to Become an AI Engineer
Making Data Warehouses Fast: A Developer’s Story
MLOps – What’s the deal behind it?