Senior Data Engineer
Sparibis Llc
Washington, DC, United States
8 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.indeed.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source
Tech stack
Query Performance
Amazon Web Services
Amazon S3
Unit Testing
Cloud Database
CompTIA Network+
CompTIA Security+
Data Architecture
Information Engineering
Data Files
Data Hub
Data Integration
+38 more
Extract Transform Load (ETL)
Data Mapping
Data Transformation
Data Structures
Data Warehousing
Software Debugging
Decision Support Systems
Amazon DynamoDB
Apache Hive
Python (Programming Language)
Pentaho Data Integration
Software Tools
Data Driven Tests
Software Configuration Management
Software Engineering
SQL Databases
SQL Server Integration Services
Data Streaming
Enterprise Data Management
Data Processing
Cloud Monitoring
Apache Spark
AWS Lambda
Git
Data Lakes
Pyspark
Information Technology
Data Analytics
AWS Data Analytics
Apache Kafka
Data Management
Functional Programming
Cloudwatch
Software Version Control
Data Pipelines
Docker
Jenkins
Databricks
Job description
- Provide data engineering expertise in the development, implementation, integration, and sustainment of data architectures, data hubs, data lakes, and data warehouse solutions.
- Support the design and development of data models, data structures, and data acquisition processes that enable data-driven decision making across the HC/HR Data Domain.
- Develop engineering and implementation plans for data hubs, data acquisition, data modeling, data integration and related data engineering activities.
- Research and evaluate existing data sources within the data lake and enterprise data environment to identify authoritative, reliable, and appropriate sources for data hub development.
- Plan, create, and maintain data architectures, ensuring alignment with business requirements.
- Design, develop, maintain, and optimize ETL/ELT data pipelines and data transformation processes supporting enterprise data integration and data hub activities.
- Develop and maintain batch and streaming data pipelines using technologies such as Spark, Python, Databricks, Palantir Foundry, Kafka, and related data engineering tools.
- Develop and maintain data acquisition processes to ingest, transform, validate, and integrate data from diverse enterprise sources.
- Implement incremental data loading strategies to optimize data processing and data freshness.
- Define and implement approaches for handling late-arriving data, processing windows, data freshness, and other data lifecycle considerations.
- Identify opportunities to automate manual data processes and improve the efficiency, reliability, and scalability of data engineering workflows.
- Develop, maintain, and optimize data engineering solutions within Databricks and enterprise data lake environments.
- Configure, monitor, and manage Databricks clusters in accordance with DON policies, standards, security requirements, and technical guidelines.
- Document Databricks cluster configurations, parameters, and operational standards for internal and external stakeholders.
- Develop and maintain Spark-based data processing solutions using PySpark, Spark SQL, Spark Data Frame, Data Sets, and related technologies.
- Implement and maintain Delta Lake solutions, including Delta Live Tables where applicable, to support scalable and reliable data processing.
- Optimize data processing, storage, and query performance within data lake environments.
- Support data engineering and application development activities within Palantir Foundry, including development and maintenance of ontologies, ETL/ELT pipelines, applications, and data-driven user interfaces.
- Support the integration of Palantir Foundry capabilities with enterprise data engineering and analytics environments.
- Develop and maintain streaming data solutions using Kafka, Kafka Streams, ksqlDB, and related technologies.
- Configure and manage Kafka topics and associated components, including Schema Registry, to support reliable and scalable data processing.
- Develop and maintain Python-based data processing applications and AWS Lambda functions.
- Support data integration and processing using AWS services such as S3, Kinesis, Lambda, and DynamoDB.
- Monitor, troubleshoot, and optimize cloud-based data processing solutions for performance and scalability.
- Develop and maintain data quality controls, validation processes, and data quality gates to ensure data accuracy, completeness, consistency, and reliability.
- Develop and execute data-driven testing and unit testing for Spark, Python, and other data processing solutions.
- Establish and maintain data lifecycle policies and processes, including retention, backup, recovery, and data management requirements.
- Monitor and troubleshoot data pipelines and processing jobs to identify and resolve data quality, performance, and integration issues.
- Identify opportunities to improve data processing performance, reliability, scalability, and maintainability.
Requirements
Years’ Experience: 5+ years professional data engineering related experience.
Education: Bachelor’s degree in computer science, Mathematics, Statistics or IT related field.
Clearance: Applicants must be able to obtain and maintain a Secret security clearance. United States Citizenship is required as part of the eligibility criteria to be able to obtain this type of security clearance.
Certifications:
- Active CompTIA Security+ or CompTIA Network+ certification., * Professional experience in data architecture, data engineering, data hub, data lake, and/or data warehouse development.
- Experience with Databricks and/or Palantir Foundry required.
- Active CompTIA Security+ or CompTIA Network+ certification preferred. If selected, the candidate must be able to obtain a CompTIA Security+ certification prior to being eligible to begin supporting the program., * 5+ years of IT experience focusing on enterprise data engineering to include data modeling, data quality, data mapping tools, and other technical documentation.
- Must be eligible to obtain and maintain a Secret security clearance.
- Bachelor’s degree in computer science, Engineering, Mathematics, Statistics, or IT related field required.
- Active CompTIA Security+ or CompTIA Network+ certification preferred. If selected, the candidate must be able to obtain a CompTIA Security+ certification prior to being eligible to begin supporting the program.
- Professional experience in data architecture, data engineering, data hub, data lake, and/or data warehouse development.
- Experience designing, developing, implementing, and supporting enterprise-scale data engineering solutions.
- Experience with data modeling, data mapping, data quality, data integration, and data management.
- Experience supporting large-scale, high-performance enterprise data applications.
- Experience developing and supporting ETL/ELT data pipelines and data transformation processes.
- This includes data integration tools such as SSIS, Pentaho, AWS Data Migration Service, etc.
- Hands-on experience with Databricks required.
- Experience with Python and SQL for data engineering and data processing.
- Experience with Apache Spark and/or PySpark, include Spark SQL, Data Frames, and/or Data Sets preferred.
- Experience supporting data integration, migration, transformation, data warehouse, data hub, and/or data lake/delta lake implementations.
- Experience with batch and streaming data processing architectures.
- Experience with Kafka and/or other distributed steaming technologies.
- Experience developing and supporting Kafka Streams.
- Experience working in AWS cloud environment with familiarity in AWS services such as S3, Lambda, Kinesis, and/or DynamoDB.
- Experience with Git based version control, including branching, merging, pull requests, and repository management.
- Experience supporting CI/CD pipelines and cloud monitoring with tools such as Jenkins, Docker, and/or CloudWatch preferred.
- Experience with Palantir Foundry, including building and maintaining ontologies, ETL/ELT pipelines, applications, and data-driven user interfaces preferred.
- Experience with the Jupiter data and analytics environment, particularly in support of the HC/HR Data Domain preferred but not required.
- Debug, troubleshooting, design and implement solutions to complex technical issues.
- Ability to thrive in a team-based environment.
- Experience briefing the benefits and constraints of technology solutions to technology partners, stakeholders, team members, and senior level of management.
- Must have strong written and verbal communication.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
EM
Eli McGarvie
over 3 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
DS
Dhannush Subramani
Top Big Data Technologies That You Need to Know
about 4 years ago
BB
Benedikt Bischof
Making Data Warehouses Fast: A Developer’s Story
about 4 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
about 1 month ago
KD
Krissy Davis
Best Coding Boot Camps in Germany
over 3 years ago