> Markdown version of [/jobs/ext/48407-data-engineering](https://www.wearedevelopers.com/jobs/ext/48407-data-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineering - **Company:** Lever, Inc. - **Location:** New York, NY, United States - **Experience:** Expert - **Salary:** $170,000.0 - $190,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, JIRA, Big Data, Computer Programming, Information Engineering, Data Infrastructure, Database Queries, Distributed Systems, Job Scheduling, Python (Programming Language), SQL Databases, Data Streaming, Workflow Management Systems, Circleci, Cloud Platform System, GitHub Copilot, Apache Spark, Backend, Git, Pyspark, Kubernetes, Apache Kafka, Data Pipelines, Amazon Elastic Mapreduce (EMR) - **Published:** May 21, 2026 - **Apply:** https://jobs.lever.co/h1/92413fc3-0cd1-40fd-b835-d589b6ea6ead ## About the Role You're a hands-on Staff IC and technical leader who thrives in complex data environments. You bring clarity to ambiguity, turn messy problems into reliable systems, and operate with a strong sense of ownership and impact. You collaborate effectively across functions, help others move faster, and are comfortable working across the full data and infrastructure stack. - You have a proven track record of leading large, complex technical projects from concept to production. - You bring deep experience building and evolving large-scale data architectures, pipelines, or distributed systems. - You operate well in high-ambiguity environments, make pragmatic trade-offs, and keep execution moving. - You communicate clearly with both technical and non-technical partners and influence direction without needing formal authority. - You raise the bar for engineering excellence through thoughtful design, high-quality code, and strong documentation. - You invest in others through mentorship, pairing, and constructive feedback. REQUIREMENTS -8+ years as a software, data, or backend engineer building and operating scalable, production-grade systems. - Experience with large-scale data processing (e.g., Spark/PySpark on EMR or similar) or scalable distributed backend systems, with the ability to quickly deepen expertise in our data stack (PySpark, EMR, Hudi/Delta). - Strong proficiency in SQL, including writing and optimizing complex queries over large datasets. - Strong programming experience in Python (or a modern language with the ability to quickly ramp up in Python). - Experience designing systems or large-scale datasets/pipelines with attention to performance, reliability, and maintainability. - Hands-on experience with modern engineering workflows and tooling such as Git, JIRA, and CI/CD systems (e.g., CircleCI). - Comfort deploying and troubleshooting distributed workloads in cloud environments such as AWS EMR or Kubernetes. - Experience with workflow orchestration or job scheduling tools (e.g., Airflow, Argo). - Demonstrated ability to independently drive complex, cross-team technical initiatives and influence stakeholders without formal authority. - Experience with streaming/messaging technologies (e.g., Kafka, Kinesis) nice to have - Background in RWE, healthcare data, or other complex/regulated data domains is preferred - Experience using AI-assisted coding tools (e.g., GitHub Copilot, Claude Code) to accelerate development while maintaining quality is encouraged ## Description Data Engineering is responsible for the development and delivery of our most important asset, our data. Across thousands of data sources globally, the team ensures that only accurate, normalized data flows to our customers, at the speed required to match real-world changes. As we expand the markets we serve and increase the breadth and depth of data we capture, we need senior technical leaders who can drive execution, scalability, and architectural excellence. WHAT YOU'LL DO AT H1 As a Staff Data Engineer on the Real World Evidence (RWE) team, you'll be one of the most senior individual contributors and a key technical leader for our largest datasets and pipelines. You'll drive some of H1's most visible data initiatives and help reduce bottlenecks across teams, providing critical technical leadership support during US hours. You will: - Act as a self-starter who drives execution independently, taking ownership and initiative with minimal need for day-to-day direction. - Lead high-visibility RWE projects, starting with claims data, and keep multiple initiatives moving by proactively unblocking teams. - Own the end-to-end architecture for critical data assets, ensuring solutions are scalable, reliable, and aligned with H1's long-term vision. - Design, build, and optimize large-scale data pipelines (hundreds of TBs) for performance, reliability, and cost efficiency. - Partner with Product, Data Science, and downstream engineering teams to align priorities, manage dependencies, and deliver high-value outcomes. - Represent engineering in cross-functional forums, shaping roadmaps and reducing reliance on senior leadership for day-to-day decisions. - Develop deep domain expertise and mentor other engineers, helping raise the technical bar and influence the evolution of our data products. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Integrate your Cognitive Assistant with 3rd-party DBs and software](https://www.wearedevelopers.com/videos/249-integrate-your-cognitive-assistant-with-3rd-party-dbs-and-software) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)