Big Data Engineer (AI-Forward)

SGA Inc.
Tysons, VA, United States
3 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Experience level
Starter
Experience required
1 year minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Agile Methodology Artificial Intelligence Amazon Web Services Amazon S3 Big Data Code Review Information Systems Continuous Integration Information Engineering Data Migration Cursor (Graphical User Interface Elements)
+17 more
Software Debugging Apache Hive Python (Programming Language) Scrum Methodology SQL Databases GitHub Copilot Large Language Models Prompt Engineering Apache Spark Kaggle Git Integration Tests Information Technology Software Coding GPT Data Pipelines Databricks

Job description

Software Guidance & Assistance, Inc., (SGA), is searching for a Junior Big Data Engineer (AI-Forward) for a CONTRACT assignment with one of our premier Regulatory clients in Tysons, VA., * Build and maintain data processing pipelines using Spark and Python, with mentorship from senior engineers

  • Use AI coding assistants as your default working mode for scaffolding pipelines, generating tests, understanding unfamiliar code, and drafting documentation, then review and validate everything before it ships. You own the output, not the tool
  • Write automated unit and integration tests for data quality, and debug pipeline failures using AI to speed up root-cause analysis
  • Prototype quickly and build small tools that remove repetitive work for the team, turning rough ideas into something runnable in hours rather than weeks
  • Document what you build and share what’s working, including prompts, workflows, and AI patterns that made you faster

Requirements

We are seeking an early-career Big Data Engineer who works AI-first. This role is for someone 1-2 years into their career who already treats AI tools as a core part of how they build, not an occasional autocomplete, and who has real projects to show for it.

You will build and maintain data pipelines alongside senior engineers using Spark, Python, and AWS. SQL and coding fundamentals matter, but we’re not screening for deep expertise. We’ll teach the data engineering. What we can’t teach is the instinct to reach for AI, get a working prototype fast, and then have the judgment to know when the output is wrong., * Bachelor’s degree in Computer Science, Information Systems, Engineering, or related discipline with 1-2 years of relevant experience. Internships, co-ops, research work, open-source contributions, and substantial personal projects all count. Master’s degree may substitute for experience.

AI Skills (Core Requirement)

  • Daily hands-on use of AI development tools such as GitHub Copilot, Q Developer, ChatGPT, Claude, Cursor, or Kiro
  • Prompt engineering: you can get useful output from a vague starting point, and you iterate rather than accepting the first answer
  • Critical validation: you can spot when AI code is subtly wrong, including bad edge-case handling, hallucinated APIs, and plausible-looking logic errors. This is the single most important skill in the role
  • Prototyping speed and workflow design: you’ve used AI to go from idea to working demo fast, and you’ve changed how you work to take advantage of it
  • Bonus: experience with AI agents, MCP servers, RAG, or LLM APIs in something you actually built

Technical Skills

  • Solid programming ability in Python, writing clean, readable code you can explain line by line, with comfort using Git and code review
  • Working SQL: joins, aggregations, filtering, and awareness of NULLs and duplicates. Basic is fine
  • Some exposure to Spark and AWS (S3 at minimum) in any setting, including coursework, internship, personal project, or free-tier Databricks, plus a basic grasp of why moving data between machines is expensive
  • Strong communication and fast learning, with the ability to explain a technical problem to someone who wasn’t in your head and the instinct to ask for help early

The X Factor

  • At least one project you can walk us through, covering what you built, how you used AI to build it, what broke, and what you’d do differently
  • Something you built or taught yourself because you wanted to, not because it was assigned
  • Comfortable pushing back respectfully, on people and on AI output

Preferred Skills :

  • Exposure to Hive, Trino, or another query engine
  • Experience with CI/CD
  • Prior Agile team experience (Scrum or Kanban)
  • Hackathons, Kaggle, technical blog posts, or AWS certifications
  • Financial Services exposure

About the company

By applying for a job with SGA, you agree to allow SGA to process your application for this and future opportunities in accordance with our Privacy Policy. Also, to ensure timely processing, you agree to be contacted by our AI recruiter via email, text, or phone. Message frequency varies and data rates may apply, but you can reply STOP to any SMS message to opt-out of texts and may contact SGA at to opt-out of AI communications. The choice not to engage with AI will not adversely impact your consideration for placement. AI is not used to make any hiring determinations.

SGA is a technology and resource solutions provider driven to stand out. We are a women-owned business. Our mission: to solve big IT problems with a more personal, boutique approach. Each year, we match consultants like you to more than 1,000 engagements. When we say let’s work better together, we mean it. You’ll join a diverse team built on these core values: customer service, employee development, and quality and integrity in everything we do. Be yourself, love what you do and find your passion at work. Please find us at .

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:31 min

Essential AI and human skills for future teams

Alexander Weißhaupt Alexander Weißhaupt +1 · World Congress 2025

1:22 min

Downloading and inspecting data frames via the Kaggle API

Lutske van der Meer Lutske van der Meer · World Congress 2024

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · World Congress 2024

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:05 min

Introducing the sample Kaggle recipe dataset

Olena Kutsenko · World Congress 2022

Videos

See all

Related articles

See all