Data Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+39 more
Job description
- Provide data engineering expertise in the development, implementation, integration, and sustainment of data architectures, data hubs, data lakes, and data warehouse solutions.
- Support the design and development of data models, data structures, and data acquisition processes that enable data-driven decision making across the HC/HR Data Domain.
- Develop engineering and implementation plans for data hubs, data acquisition, data modeling, data integration and related data engineering activities.
- Research and evaluate existing data sources within the data lake and enterprise data environment to identify authoritative, reliable, and appropriate sources for data hub development.
- Plan, create, and maintain data architectures, ensuring alignment with business requirements.
- Design, develop, maintain, and optimize ETL/ELT data pipelines and data transformation processes supporting enterprise data integration and data hub activities.
- Develop and maintain batch and streaming data pipelines using technologies such as Spark, Python, Databricks, Palantir Foundry, Kafka, and related data engineering tools.
- Develop and maintain data acquisition processes to ingest, transform, validate, and integrate data from diverse enterprise sources.
- Implement incremental data loading strategies to optimize data processing and data freshness.
- Define and implement approaches for handling late-arriving data, processing windows, data freshness, and other data lifecycle considerations.
- Identify opportunities to automate manual data processes and improve the efficiency, reliability, and scalability of data engineering workflows.
- Develop, maintain, and optimize data engineering solutions within Databricks and enterprise data lake environments.
- Configure, monitor, and manage Databricks clusters in accordance with DON policies, standards, security requirements, and technical guidelines.
- Document Databricks cluster configurations, parameters, and operational standards for internal and external stakeholders.
- Develop and maintain Spark-based data processing solutions using PySpark, Spark SQL, Spark Data Frame, Data Sets, and related technologies.
- Implement and maintain Delta Lake solutions, including Delta Live Tables where applicable, to support scalable and reliable data processing.
- Optimize data processing, storage, and query performance within data lake environments.
- Support data engineering and application development activities within Palantir Foundry, including development and maintenance of ontologies, ETL/ELT pipelines, applications, and data-driven user interfaces.
- Support the integration of Palantir Foundry capabilities with enterprise data engineering and analytics environments.
- Develop and maintain streaming data solutions using Kafka, Kafka Streams, ksqlDB, and related technologies.
- Configure and manage Kafka topics and associated components, including Schema Registry, to support reliable and scalable data processing.
- Develop and maintain Python-based data processing applications and AWS Lambda functions.
- Support data integration and processing using AWS services such as S3, Kinesis, Lambda, and DynamoDB.
- Monitor, troubleshoot, and optimize cloud-based data processing solutions for performance and scalability.
- Develop and maintain data quality controls, validation processes, and data quality gates to ensure data accuracy, completeness, consistency, and reliability.
- Develop and execute data-driven testing and unit testing for Spark, Python, and other data processing solutions.
- Establish and maintain data lifecycle policies and processes, including retention, backup, recovery, and data management requirements.
- Monitor and troubleshoot data pipelines and processing jobs to identify and resolve data quality, performance, and integration issues.
- Identify opportunities to improve data processing performance, reliability, scalability, and maintainability.
Requirements
Years’ Experience: 5+ years professional data engineering related experience
Education: Bachelor’s degree in computer science, Mathematics, Statistics or IT related field
Clearance: Applicants must be able to obtain and maintain a Secret security clearance. United States Citizenship is required as part of the eligibility criteria to be able to obtain this type of security clearance.
Certifications:
· Active CompTIA Security+ or CompTIA Network+ certification, * Professional experience in data architecture, data engineering, data hub, data lake, and/or data warehouse development.
- Experience with Databricks and/or Palantir Foundry required
- Active CompTIA Security+ or CompTIA Network+ certification preferred. If selected, the candidate must be able to obtain a CompTIA Security+ certification prior to being eligible to begin supporting the program., * 5+ years of IT experience focusing on enterprise data engineering to include data modeling, data quality, data mapping tools, and other technical documentation
- Must be eligible to obtain and maintain a Secret security clearance
- Bachelor’s degree in computer science, Engineering, Mathematics, Statistics, or IT related field required
- Active CompTIA Security+ or CompTIA Network+ certification preferred. If selected, the candidate must be able to obtain a CompTIA Security+ certification prior to being eligible to begin supporting the program.
- Professional experience in data architecture, data engineering, data hub, data lake, and/or data warehouse development.
- Experience designing, developing, implementing, and supporting enterprise-scale data engineering solutions.
- Experience with data modeling, data mapping, data quality, data integration, and data management.
- Experience supporting large-scale, high-performance enterprise data applications.
-
Experience developing and supporting ETL/ELT data pipelines and data transformation processes.
- This includes data integration tools such as SSIS, Pentaho, AWS Data Migration Service, etc.
Hands-on experience with Databricks required
Experience with Python and SQL for data engineering and data processing.
Experience with Apache Spark and/or PySpark, include Spark SQL, Data Frames, and/or Data Sets preferred.
Experience supporting data integration, migration, transformation, data warehouse, data hub, and/or data lake/delta lake implementations.
Experience with batch and streaming data processing architectures.
Experience with Kafka and/or other distributed steaming technologies.
Experience developing and supporting Kafka Streams
Experience working in AWS cloud environment with familiarity in AWS services such as S3, Lambda, Kinesis, and/or DynamoDB.
Experience with Git based version control, including branching, merging, pull requests, and repository management.
Experience supporting CI/CD pipelines and cloud monitoring with tools such as Jenkins, Docker, and/or CloudWatch preferred
· Experience with Palantir Foundry, including building and maintaining ontologies, ETL/ELT pipelines, applications, and data-driven user interfaces preferred
· Experience with the Jupiter data and analytics environment, particularly in support of the HC/HR Data Domain preferred but not required
· Debug, troubleshooting, design and implement solutions to complex technical issues
· Ability to thrive in a team-based environment
- Experience briefing the benefits and constraints of technology solutions to technology partners, stakeholders, team members, and senior level of management
- Must have strong written and verbal communication
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Making Data Warehouses Fast: A Developer’s Story
Top Big Data Technologies That You Need to Know
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again