> Markdown version of [/jobs/ext/2120139-lead-data-engineer](https://www.wearedevelopers.com/jobs/ext/2120139-lead-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Data Engineer - **Company:** The Doyle Group - **Location:** Shelton, CT, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Query Performance, Artificial Intelligence, Amazon Web Services, Amazon S3, Architectural Patterns, Code Review, Data Architecture, Information Engineering, Data Integration, Extract Transform Load (ETL), Relational Databases, Dimensional Modeling, Django Web Framework, Identity and Access Management, Python (Programming Language), PostgreSQL, OAuth, Operational Databases, Query Optimization, Raw Data, Role-Based Access Control, Power BI, Software Requirements Analysis, SQL Databases, Data Logging, Google Data Studio, Flask (Web Framework), Large Language Models, Snowflake, Virtual Environment, Backend, Cloudformation, Fastapi, Build Management, Amazon Relational Database Service, AWS Glue, AWS Fargate, Functional Programming, Cloudwatch, Restful APIs, Amazon Simple Queue Service (SQS), Terraform, Pagination, Data Pipelines, Call Tracing - **Published:** August 19, 2026 - **Apply:** https://www.dice.com/job-detail/cf820953-00a7-40f7-9164-e0a2eb49a56c ## About the Role This is a full-time, direct-hire position. Candidates must be authorized to work in the United States without current or future visa sponsorship., * 6+ years of professional experience in data engineering, analytics engineering, or a backend data-focused role, including time spent setting technical direction or acting as a de facto technical lead. * Expert-level Python for production data work, including: + Clean, modular, testable code; virtual environments and dependency management; error handling and logging + Building and consuming REST APIs (FastAPI, Flask, or similar) + Extending existing frameworks (e.g., Django) or writing custom connectors rather than building everything from scratch * Expert-level SQL, including CTEs, window functions, and query optimization. * Deep, hands-on Snowflake expertise - a core requirement, not a nice-to-have: + Schema design, dimensional modeling, partitioning and clustering + Warehouse sizing and credit/cost management + Role-based access control and query performance tuning + Comfort working in a raw-to-curated (medallion-style) data architecture * Comprehensive hands-on AWS experience - a requirement, not a preference: + Core comfort with S3, Lambda, and IAM + Working experience with several of: Glue, Step Functions, EventBridge, ECS/Fargate, Secrets Manager, CloudWatch, RDS, and SQS/SNS + Infrastructure as code via Terraform or CloudFormation, with a clear point of view on structuring state, modules, and environments * Experience with a managed ELT/ETL / data integration platform such as Airbyte or Funnel.io. * Pipeline orchestration experience with at least one of Temporal, Prefect, or AWS Step Functions. * Meaningful big-data / cloud-warehouse experience - a background limited to a single relational database (e.g., Postgres) is not sufficient depth for this role. * Hands-on experience building with LLMs and AI agents in production - not just using AI coding assistants day to day: + MCP-compatible interfaces, structured tool schemas, or agent orchestration workflows * Demonstrated technical leadership: + Setting architectural direction, running design reviews, and mentoring engineers through code review and pairing (this is a lead role measured by technical influence, not headcount managed) * Clear communication with both technical and non-technical stakeholders, with strong attention to detail, project management, and organizational skills. * Bachelor''s degree or equivalent professional experience. * Must be legally authorized to work in the United States without sponsorship now or in the future. ADDITIONAL PLUS * Background in media, advertising, or direct-response marketing (channels, sources, conversions, performance signals), or an adjacent domain such as CRM, sales enablement, or a DSP/media platform * Experience managing a high-volume connector environment (30-50+ integrations) with a monitoring and cost-management discipline * Prior experience mentoring or growing a data engineering team toward a formal management track * Experience with Power BI, Looker Studio, or similar BI tools consuming the pipelines you build * Comfort operating in a fast-paced, deliverable-driven, daily-standup team culture * Product-building experience - having shipped something from scratch, not just extended existing systems ## Description * Own the full ETL/ELT lifecycle across dozens of media, CRM, and vendor data sources: extraction, transformation, loading into Snowflake, orchestration, scheduling, retries, alerting, and backfills, and define the patterns the rest of the team builds to. * Design and build scalable, reliable, idempotent data pipelines that move raw data through to modeled, analysis-ready tables in a raw-to-curated (medallion-style) data architecture. * Set the architectural direction for how new pipelines, connectors, and integrations get designed, reviewed, and shipped across the team. * Integrate external systems via REST APIs, including ad platforms, CRMs, call tracking, and vendor feeds, handling OAuth 2.0 and token refresh, pagination, rate limits, exponential backoff, schema drift, and partial failures without losing or duplicating data. * Configure and manage ELT connectors such as Airbyte, Funnel.io, and AWS Glue, including sync scheduling, schema-change handling, monitoring, and row-volume cost management. * Manage AWS infrastructure as code using Terraform or CloudFormation so environments stay versioned, peer-reviewed, and reproducible, and serve as the final technical call on infrastructure decisions. * Own core Snowflake functionality: schema design, dimensional and analytics modeling, partitioning and clustering, warehouse sizing, role-based access control, and query performance tuning. * Build data quality and observability into every pipeline: freshness and volume checks, row-level validation, reconciliation against source-of-truth platforms, lineage, and alerting that surfaces problems before stakeholders do. * Treat platform cost, including Snowflake credit consumption and AWS spend, as a design constraint rather than an afterthought, and set the cost guardrails other engineers design against. * Design LLM-readable, MCP-compatible interfaces with clear action-based endpoints and strongly typed schemas so AI agents and automation can reliably query and act on the data platform, and define how those agents access and maintain context. * Serve as the primary technical point of contact for analytics, media buying, account, and finance stakeholders, translating their needs into system requirements and resolving data and application issues. * Lead code review, pair with other engineers on hard problems, unblock the team, and help set the shared technical standards the rest of the group codes to. * Mentor engineers, help define coding standards and architectural patterns, and weigh in on technical hiring decisions, without owning formal people-management or performance responsibilities. * Help keep the team building reusable, leverageable systems rather than one-off, ad hoc fixes that don''t compound over time. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Intro to FastAPI](https://www.wearedevelopers.com/videos/462-intro-to-fastapi) - [Keeping applications secure by evolving OAuth 2.0 and OpenID Connect](https://www.wearedevelopers.com/videos/100152-keeping-applications-secure-by-evolving-oauth-2-0-and-openid-connect) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Building and Deploying Multi-Agent Systems with ADK and Vertex AI](https://www.wearedevelopers.com/videos/1918-building-and-deploying-multi-agent-systems-with-adk-and-vertex-ai) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [The Prompt Engineer ✍️](https://www.wearedevelopers.com/magazine/216-the-prompt-engineer) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)