> Markdown version of [/jobs/ext/2166041-data-engineer](https://www.wearedevelopers.com/jobs/ext/2166041-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** CoStar Group - **Location:** Richmond, VA, United States (Remote available) - **Experience:** Expert - **Salary:** $133,000.0 - $172,000.0 - **Contract:** Permanent contract - **Skills:** Sql Data Warehouse, Java (Programming Language), .NET Framework, Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Apache HTTP Server, Microsoft Azure, Big Data, BigQuery, C Sharp (Programming Language), Code Generation, Continuous Integration, Data as a Services, Data Deduplication, Information Engineering, Extract Transform Load (ETL), Data Mapping, Data Migration, Data Vault Modeling, Data Warehousing, DevOps, Distributed Computing Environment, Python (Programming Language), Scala (Programming Language), Data Streaming, Google Cloud, File Transfer Protocol (FTP), Feature Engineering, Data Ingestion, Azure Data Factory, Sql Optimization, Snowflake, Apache Spark, Backend, Data Lakes, Debezium, Collibra, Apache Flink, AWS Glue, Qlikview, Star Schema, AWS Data Analytics, Apache Kafka, Graphql, Spark Streaming, Data Management, Database Replication, Terraform, Azure Synapse Analytics, Software Version Control, Data Pipelines, Amazon Redshift, Databricks - **Published:** August 21, 2026 - **Apply:** https://www.dice.com/job-detail/00fba233-6943-462b-99d8-f90cb69264e7 ## About the Role * Bachelor's degree required from an accredited, not-for-profit, in-person college/university * Track record of commitment to prior employers * 5+ years of data engineering experience, including experience setting technical direction on projects and influencing design across a team * Hands-on experience building batch and incremental ingestion pipelines from multiple heterogeneous external sources (REST/GraphQL APIs, flat-file and SFTP feeds, third-party data vendors, and database replication/CDC) * Demonstrated experience normalizing inconsistent third-party schemas into standardized enterprise data models, including source-to-target mapping, deduplication, and entity resolution * Advanced SQL and strong data modeling skills (dimensional/star schema, normalized, or Data Vault), plus proficiency in Python, Scala, Java, or C#/.NET * Production experience with a distributed processing engine and a cloud data warehouse or lakehouse platform (e.g., Databricks/Spark, Snowflake, BigQuery, Redshift, Synapse/Fabric), including pipeline orchestration tooling (e.g., Airflow, Dagster, Databricks Workflows, Azure Data Factory) * Experience implementing data quality, validation, and reconciliation controls for externally sourced data, with version control and CI/CD in a major cloud environment (AWS, Azure, or Google Cloud Platform) * Experience with agentic engineering with AI-assisted development tools (Claude Code or similar) to accelerate software delivery, * Production-scale expertise with Databricks (Delta Lake, Unity Catalog, Delta Live Tables / declarative pipelines, Databricks SQL, Workflows) or an equivalent lakehouse platform, including open table formats (Delta Lake, Apache Iceberg, Apache Hudi) * Experience with a modular transformation framework and tested, version-controlled models (dbt, SQLMesh, or equivalent) * Experience with managed ingestion and replication tooling (Fivetran, Airbyte, Qlik Replicate, AWS DMS, Debezium, Kafka Connect) and streaming or near-real-time ingestion (Kafka, Kinesis, Spark Structured Streaming, Flink) * Data quality and observability frameworks (Great Expectations, Soda, dbt tests, Monte Carlo) and data cataloging, lineage, and governance tooling (Unity Catalog, AWS Glue Data Catalog, Collibra, Alation) * Master data management, entity resolution, or record-linkage tooling, including address and geospatial standardization * AWS data services (S3, Glue, EMR, Lambda, Step Functions, Redshift), Terraform, and backend data services in C# / .NET * AI/ML and feature-engineering pipelines and ability to apply AI assist or agentic workflows for code generation, review and root-cause analysis ## Description We are seeking a Senior Data Engineer to join a new team building our next generation of data warehousing pipelines and ecosystem in the finance tech organization. Initially this role will focus on building a data migration and transformation process for external systems for use in our enterprise contracting system. This will include receiving data from multiple data sources of varying format, frequency and type. You will create applications to process it in a repeatable, customizable and recurring state to simplify the conversion process of our financial data across systems and acquisitions both today and in the future. You will also help standardize our data environment across our billing systems to allow for greater visibility into patterns, errors and optimizations to our business, consisting of billions of dollars of payments a year. This position is in Richmond, VA and has a work schedule of Monday through Thursday in office and Friday work from home. Responsibilities: * Design, build, and maintain scalable data pipelines and data platforms * Develop and optimize ETL/ELT processes for large, complex datasets * Collaborate with engineering, product, DevOps and DBA teams to deliver production systems * Contribute to technical design decisions and mentor other engineers on the team ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Meet Your New BFF: Backend to Frontend without the Duct Tape](https://www.wearedevelopers.com/videos/682-meet-your-new-bff-backend-to-frontend-without-the-duct-tape) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)