> Markdown version of [/jobs/ext/493380-etl-data-engineer](https://www.wearedevelopers.com/jobs/ext/493380-etl-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ETL Data Engineer - **Company:** Eliassen Group - **Location:** Tysons, VA, United States - **Experience:** Experienced - **Salary:** $145,600.0 - $183,040.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Business Analytics Applications, Confluence, Build Automation, Automation of Tests, Big Data, Configuration Management, Code Generation, Information Systems, Databases, Extract Transform Load (ETL), Software Debugging, Memory Management, Apache Hive, Python (Programming Language), Object-Oriented Software Development, Pair Programming, Performance Tuning, Query Optimization, Cloud Services, Software Engineering, SQL Databases, Data Logging, Google Cloud, GitHub Copilot, Large Language Models, Multi-Agent Systems, Concurrency, Prompt Engineering, Apache Spark, Data Lakes, AI Platforms, Pyspark, Information Technology, Wikis, Virtual Agents, Dataiku, Cloudwatch, Code Restructuring, GPT, Data Pipelines - **Published:** June 5, 2026 - **Apply:** https://dejobs.org/x/x/304424BCFC014AED955485B31109309F/job/ ## About the Role * Experience building data pipelines with Apache Spark (PySpark preferred) and SQL. * Experience with SQL engines such as Hive or Trino and cloud data platforms including AWS S3, EMR, and Lambda. * Understanding of data skew, large-volume processing, and troubleshooting job failures due to resources, data quality, and scalability. * Hands-on debugging and mitigation experience. * Practical experience building LLM-powered agent systems that use tools and produce structured outputs. * Experience with agent frameworks such as LangChain, LangGraph, or AWS Strands. * Knowledge of prompt engineering, RAG architectures, and context or memory management. * Experience with foundation model APIs such as Anthropic Claude, Amazon Nova, or OpenAI. * Understanding of agent memory tiers and strategies for persistence, pruning, and retrieval. * Familiarity with harness patterns including deterministic guardrails, tool routing, and verification loops. * Hands-on experience with AI development tools such as GitHub Copilot, Q Developer, ChatGPT, or Claude. * Experience with spec-driven development for AI-assisted code generation and validation. * Ability to leverage AI pair programming for suggestions, debugging, refactoring, and automated test generation. * Experience with AWS services including S3, EMR, EMR on EKS, Lambda, Bedrock, and Step Functions. * Hands-on experience using S3 with Spark and related file format or consistency considerations. * Familiarity with AWS Bedrock guardrails, knowledge bases, and agent orchestration. * Exposure to Google Cloud Vertex AI or equivalent managed AI platforms. * Familiarity with AWS monitoring and logging tools such as CloudWatch and CloudTrail. * Proficiency in Python with clean, modular, and performant code and understanding of functional concepts. * Strong understanding of collections, concurrency, and memory management. * Proficiency with SQL window functions, joins, aggregations, and complex query optimization including edge cases., * Bachelor's degree in Computer Science, Data Science, Information Systems, or related discipline with at least two years of related experience, or equivalent training and work experience. Financial services experience preferred. * Demonstrated expertise in object-oriented and database technologies resulting in enterprise-quality solutions. * Knowledge of software engineering approaches including test automation, build automation, and configuration management. * Strong written and verbal technical communication skills and effective cross-team collaboration. * Ability to learn new skills rapidly and operate in a fast-paced environment. ## Description * Build and maintain ETL/ELT pipelines using Apache Spark, Hive, and Trino across S3-based data lakes. * Develop and optimize SQL for large-scale surveillance datasets using window functions, joins, and complex aggregations. * Engineer big data systems on EMR-on-EC2 and EMR-on-EKS and deliver solutions on analytical platforms such as SageMaker, Domino, or Dataiku. * Participate in data quality monitoring, anomaly detection, and production incident investigation. * Develop AI agent systems using AWS Bedrock and agent frameworks such as Strands Agents SDK or LangChain/LangGraph. * Design agent harnesses that combine LLM reasoning with deterministic execution including skill or RAG-based SQL generation and structured output validation. * Implement agent memory, context management, and tool integration including MCP servers, API connectors, and data catalog lookups. * Build evaluation frameworks for agent accuracy covering paraphrase robustness, routing precision, and structural consistency. * Stay informed on advances in LLM frameworks and emerging AI capabilities. * Write clean, well-tested code and contribute to CI/CD pipelines and infrastructure-as-code on AWS. * Ensure secure handling of sensitive regulatory data with auditable execution traces. * Adhere to secure development practices and technology policies. * Partner across teams, communicate at the appropriate technical level, and maintain documentation on Confluence or Wiki. * Learn from senior team members and contribute to process improvement. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [The Impact of AI on Game Development and the Industry](https://www.wearedevelopers.com/videos/100297-the-impact-of-ai-on-game-development-and-the-industry) - [Streaming AI Responses in Real-Time with SSE in Next.js & NestJS](https://www.wearedevelopers.com/videos/1630-streaming-ai-responses-in-real-time-with-sse-in-next-js-nestjs) - [Stack Overflow: Community and AI](https://www.wearedevelopers.com/videos/600-stack-overflow-community-and-ai) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)