> Markdown version of [/jobs/ext/2599326-senior-data-engineer](https://www.wearedevelopers.com/jobs/ext/2599326-senior-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer - **Company:** Benchmark - **Location:** Chicago, IL, United States (Remote available) - **Experience:** Expert - **Salary:** $135,000.0 - $160,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Data Analysis, Automation of Tests, Big Data, Cloud Computing, Code Review, Databases, Continuous Integration, Information Engineering, Data Infrastructure, Extract Transform Load (ETL), Data Systems, Distributed Computing Environment, Django Web Framework, Amazon DynamoDB, Fault Tolerance, Identity and Access Management, Python (Programming Language), PostgreSQL, Operational Databases, Software Architecture, SQL Databases, Unstructured Data, Data Logging, Data Processing, Sql Optimization, Large Language Models, Apache Spark, Electronic Medical Records, Git, Kubernetes, Infrastructure Automation Frameworks, Data Management, Functional Programming, Api Design, Amazon Simple Queue Service (SQS), Data Pipelines, Amazon Elastic Mapreduce (EMR), Docker - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=5a715f3ec85a8b5e ## About the Role * Bachelor's degree in a STEM field or equivalent professional experience * 8+ years of professional experience building and operating production data systems * Demonstrated independent ownership of major data systems, platform components, architectural decisions, or complex modernization initiatives * Advanced Python software-engineering experience, including modular architecture, type annotations, automated testing, packaging, dependency management, API design, and reusable library or framework development * Advanced SQL and data-modeling skills across operational and analytical workloads * Experience architecting fault-tolerant ETL/ELT systems that support replay, backfills, schema evolution, and failure recovery * Strong experience designing and operating cloud-based production architectures, with AWS preferred * Production experience with Docker, Kubernetes, and orchestration frameworks such as Airflow or an equivalent * Practical experience with CI/CD, infrastructure as code, automated testing, and production observability * Demonstrated ownership of a legacy modernization or platform migration effort, including dependency analysis, migration sequencing, validation, cutover, and operational transition * Significant experience diagnosing production incidents, complex data failures, and performance bottlenecks and implementing durable corrective actions * Ability to evaluate and clearly communicate architectural tradeoffs involving scalability, reliability, maintainability, security, cost, delivery speed, and team capability * Ability to independently convert incomplete or ambiguous requirements into pragmatic technical direction and executable delivery plans * Demonstrated technical influence through design reviews, engineering standards, mentorship, or shared platform development Preferred Skills: * Expert-level Python engineering for maintainable production systems, not only standalone scripts or notebooks * Expert-level SQL and strong relational and analytical data-modeling knowledge * Experience handling large data volumes, schema evolution, data-quality enforcement, and complex transformation workflows * Experience with AWS services such as S3, Lambda, SQS/SNS, IAM, and managed data-processing services * Strong experience with Git, Docker, Kubernetes, and automated software-delivery practices * Experience designing reusable platform abstractions and determining when functionality belongs in a shared component versus an individual pipeline * Strong troubleshooting skills across application code, infrastructure, orchestration, databases, and data behavior * Ability to lead through technical credibility and influence without relying on formal authority * Relevant technologies include SQL, Python, AWS, PostgreSQL, Spark/EMR, Git, Docker, Kubernetes, Airflow, Django, and DynamoDB, * Experience owning major components of a greenfield data platform or internal developer platform * Experience replacing a commercial or legacy ETL platform with code-first, cloud-native tooling * Experience developing shared Python libraries, frameworks, templates, or platform abstractions used by other engineers * Hands-on experience with Spark and AWS EMR for distributed data processing * Experience working in regulated or security-sensitive environments, including AWS GovCloud * Experience implementing LLM- or agent-based workflows with structured outputs, tool integration, validation, observability, security boundaries, and appropriate human review * Experience serving as a technical subject matter expert in client-facing or cross-functional architecture discussions ## Description This is a senior individual contributor role for a deeply experienced Data Engineer who can independently lead complex platform-level work from discovery through production operation. The successful candidate will operate with limited day-to-day technical oversight, translate ambiguous objectives into executable technical plans, own consequential architecture and design decisions, and build reusable capabilities that increase the effectiveness of the broader data engineering team. This role will be a critical technical contributor to Benchmark's transformation from legacy ETL systems to a modern, cloud-native architecture built on Python, Kubernetes, AWS, and AI-assisted and agentic workflows., * Owning the technical design, implementation, rollout, and operational support of major components of a greenfield data platform leveraging Python, Kubernetes, and AWS * Leading complex platform initiatives from discovery and architecture through incremental delivery and production adoption * Independently translating ambiguous business and technical objectives into sound designs, implementation plans, and production-ready systems * Designing, developing, and maintaining scalable, fault-tolerant ETL/ELT pipelines across structured and unstructured data * Building reusable platform capabilities for configuration, execution, logging, metrics, error handling, testing, deployment, and operational support * Designing data workflows for idempotency, replayability, backfills, schema evolution, partial-failure recovery, and safe production rollout * Assessing legacy data processes and leading incremental modernization strategies that preserve production continuity while reducing operational risk * Establishing engineering patterns and standards, leading technical design reviews, and challenging unnecessary complexity or weak architectural assumptions * Diagnosing complex production, performance, and data-quality issues and driving durable corrective actions * Improving platform scalability, reliability, observability, maintainability, security, and cost efficiency * Collaborating with application engineering, data science, QA, product, analytics, and client-facing teams to deliver clean, reliable, and production-ready data capabilities * Evaluating and integrating AI-assisted or agentic workflows where they provide measurable improvements to data processing, engineering productivity, or system interaction * Providing technical mentorship through architecture guidance, code reviews, reusable patterns, documentation, and direct engineering feedback * Acting as a primary technical subject matter expert in internal, cross-functional, and client-facing discussions ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)