Data Automation Engineer (Public Trust)
System One
Washington, DC, United States
1 day ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dcjobsite.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$119,912.0
Working hours
Regular working hours
Job source
Tech stack
Agile Methodology
Artificial Intelligence
Amazon Web Services
Amazon S3
Data Analysis
JIRA
Microsoft Azure
Continuous Integration
Information Engineering
Data Integration
Extract Transform Load (ETL)
Data Mining
+53 more
Data Systems
Amazon DynamoDB
Github
Identity and Access Management
Python (Programming Language)
Machine Learning
Microsoft SQL Server
SQL Azure
Open Source Technology
Operational Databases
PCI Data Security Standards
Performance Tuning
Role-Based Access Control
Search Technologies
Apache Solr
SQL Databases
Systems Integration
Software Vulnerability Management
Enterprise Data Management
Amazon Connect
Data Processing
Data Ingestion
Azure Data Factory
Retrieval-Augmented Generation
Large Language Models
Apache Spark
Software Troubleshooting
Multi-Cloud
Generative AI
AWS Lambda
Amazon Virtual Private Cloud (VPC)
Cloudformation
Containerization
Apache Flume
Kubernetes
Infrastructure Automation Frameworks
Information Technology
Data Lineage
Deployment Automation
HuggingFace
AWS Glue
Bicep
AWS Fargate
AWS Data Analytics
Apache Kafka
Firewall Services Module
Restful APIs
Terraform
Data Pipelines
Amazon Elastic Mapreduce (EMR)
Docker
Jenkins
Databricks
Job description
- Design and implement scalable data automation workflows using AWS services, with integration to selected Azure data platforms where required. Develop ETL/ELT processes to ingest, transform, and move data across Amazon DynamoDB, SQL Server hosted on AWS, Azure SQL, and other enterprise data sources.
- Design, develop, and support batch and near-real-time ingestion pipelines using Apache Spark and technologies such as Kafka or Flume, and collaborate with the search engineering team to integrate those pipelines with the existing Apache Solr platform.
- Evaluate and apply Generative AI services and frameworks, such as Amazon Bedrock, Azure OpenAI, Hugging Face, and LangChain, to prototype and evaluate selected GenAI-assisted capabilities, such as metadata enrichment, data-quality analysis, structured data extraction, anomaly identification, and natural-language access to enterprise data.
- Recommend suitable use cases for future implementation.
- Develop scalable data-processing solutions using Amazon EMR and containerized deployment environments such as AWS Fargate or Kubernetes.
- Integrate Amazon Connect customer-interaction data into analytical data stores for operational reporting and analytics.
- Apply source-control, build, containerization, and CI/CD practices using tools such as GitHub, Azure DevOps, Jenkins, and Docker.
- Implement data solutions in accordance with established security and compliance controls, including identity and access management, KMS encryption, VPC isolation, role-based access control, and firewall policies.
- Support Agile DevOps processes with sprint-based delivery of pipeline and AI-enabled features., System One, and its subsidiaries including Joulé, ALTA IT Services, CM Access, TPGS, and MOUNTAIN, LTD., are leaders in delivering workforce solutions and integrated services across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible full-time employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan.
Requirements
- Bachelor’s degree in Computer Science or a related field and 5+ years of experience in data engineering, data automation, or a related discipline.
- Ability to independently design, develop, test, and troubleshoot Python- and SQL-based data pipelines in AWS environments, including integrations with Azure services where required, and clearly explain personal contributions to production implementations.
- Strong hands-on experience with Apache Spark and working knowledge of at least one streaming or ingestion technology, such as Apache Kafka or Apache Flume.
- Hands-on experience with multiple AWS data and integration services, including several of the following: Amazon S3, AWS Glue, AWS Lambda, Amazon EMR, AWS Step Functions, and at least one AWS database service.
- Practical experience integrating at least one LLM platform or model service, such as Amazon Bedrock, Azure OpenAI Service, or an open-source model, into a Python-based workflow.
- Experience integrating REST APIs and external services into Python-based data pipelines and automated workflows.
- Experience using Jira and one or more source-control, build, or CI/CD platforms, such as GitHub, Azure DevOps, or Jenkins.
- Strong troubleshooting and performance-optimization skills across SQL, Spark, batch pipelines, and near-real-time ingestion workflows.
- Experience supporting production data platforms, including SLA monitoring, incident resolution, root-cause analysis, data reconciliation, performance troubleshooting, vulnerability remediation, and recurring maintenance.
- Good communication and presentation skills.
- US Citizenship and ability to obtain Federal government Public Trust clearance., * Relevant certifications, such as AWS Certified Data Engineer - Associate, AWS Certified Machine Learning - Specialty, Microsoft Certified: Azure AI Engineer Associate, or Databricks Certified Data Engineer.
- Familiarity with retrieval-augmented generation pipelines, embeddings, and vector-search technologies such as Apache Solr, Amazon OpenSearch Service, pgvector, or similar platforms.
- Experience with multi-cloud data integration (AWS and Azure).
- Experience with Docker and Kubernetes for containerized deployment, scalable data processing, and orchestration.
- Experience operationalizing Generative AI workflows, including prompt and model configuration, evaluation, observability, monitoring, and lifecycle management.
- Knowledge of data lineage/governance tools (Purview, Unity Catalog, AWS Glue Catalog).
- Familiarity with infrastructure-as-code tools, such as Terraform, AWS CloudFormation, or Bicep, for automated deployments.
- Experience with compliance frameworks (FedRAMP, PCI-DSS, HIPAA).
About the company
System One, and its subsidiaries including Joulé and Mountain Ltd., are leaders in delivering outsourced services and workforce solutions across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dcjobsite.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
almost 3 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
EM
Eli McGarvie
Data Engineer Salary UK
over 3 years ago
LM
Luis Minvielle
The Most Popular IT Jobs on the Market
over 2 years ago
DS
Dhannush Subramani
Top Big Data Technologies That You Need to Know
about 4 years ago
LM
Luis Minvielle
Top-Paying Tech Jobs (with Salaries)
over 2 years ago