Big Data/Spark Engineer

Software Guidance & Assistance Inc
Derwood, MD, United States
3 days ago
Apply on www.disabledperson.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Java (Programming Language) Agile Methodology Artificial Intelligence Amazon Web Services Amazon S3 Data Analysis Build Automation Automation of Tests Big Data Cloud Computing Configuration Management Information Systems
+35 more
Databases Continuous Integration Data Architecture Data Integration Extract Transform Load (ETL) Database Queries Software Debugging Distributed Systems Memory Management Apache Hadoop Apache Hive Python (Programming Language) Object-Oriented Software Development Operational Databases Performance Tuning Scrum Methodology Standard Sql Scala (Programming Language) Simple Data Format Software Engineering SQL Databases System Testing Test Case Workflow Management Systems Data Processing Data Ingestion GitHub Copilot Concurrency Prompt Engineering Apache Spark Information Technology Functional Programming GPT Data Pipelines Serverless Computing

Job description

Seeking a Big Data Engineer to design, develop, and optimize large-scale data processing systems. Will work closely with cross-functional teams to architect data pipelines, implement data integration solutions, and ensure the performance, scalability, and reliability of big data platforms. Ideally will have deep expertise in distributed systems, cloud platforms, and modern big data technologies such as Hadoop, Spark etc., * Design, develop, and maintain large-scale data processing pipelines using Big Data technologies (e.g., Hadoop, Spark, Python, Scala).

  • Implement data ingestion, storage, transformation, and analysis of solutions that are scalable, efficient, and reliable.
  • Stay current with industry trends and emerging Big Data technologies to continuously improve the data architecture.
  • Collaborate with cross-functional teams to understand business requirements and translate them into technical solutions.
  • Optimize and enhance existing data pipelines for performance, scalability, and reliability.
  • Develop automated testing frameworks and implement continuous testing for data quality assurance.
  • Conduct unit, integration, and system testing to ensure the robustness and accuracy of data pipelines.
  • Work with data scientists and analysts to support data-driven decision-making across the organization.
  • Ability to write and maintain automated unit, integration, and end-to-end tests.
  • Monitor and troubleshoot data pipelines in production environments to identify and resolve issues.

Requirements

  • Bachelor’s degree in Computer Science, Information Systems or related discipline with at least five (5) years of related experience, or equivalent training and/or work experience; Master’s degree and past Financial Services industry experience preferred.
  • Demonstrated technical expertise in Object Oriented and database technologies/concepts which resulted in deployment of enterprise quality solutions.
  • Past experience with developing enterprise quality solutions in an iterative or Agile environment.
  • Extensive knowledge of industry leading software engineering approaches including Test Automation, Build Automation and Configuration Management frameworks.
  • Experience with object oriented programming languages such as Java, Scala or Python.

Essential Technical Skills:

  • AI Tool Proficiency: Hands-on experience with AI development tools (GitHub Copilot, Q Developer, ChatGPT, Claude, etc.).
  • Technical Background: Strong software development background with ability to contribute to technical discussions.
  • Agile Methodology: Extensive experience with Scrum, Kanban, and continuous improvement practices.

Big Data technologies

  • Experience with Big data technologies such as Hadoop, Spark, Hive & Trino.
  • Understanding of common issues like:

  • Data skew and strategies to mitigate it.
  • Working with massive data volumes in PetaBytes.
  • Troublehsooting job failures due to resource limitations, bad data, scalability challenged.

Real-world debugging and mitigation stories. AI Skills

  • Prompt Engineering: Proficiency in crafting effective prompts for AI coding assistants and analysis tools.
  • AI Workflow Design: Experience redesigning development processes to leverage AI capabilities.
  • Data Analysis: Ability to interpret AI-generated insights and translate them into actionable team improvements.
  • Change Management: Experience leading teams through AI adoption and workflow transformation.

SQL Skills (Window Functions, Joins, Complex Queries)

  • Comfort with SQL window functions, multi-table joins, aggregations.
  • Provide examples of write/optimize SQL queries.
  • Handle edge cases like NULLs, duplicates, ordering, etc.

Apache Spark (Development, Internals & Tuning)

  • Understanding of Spark’s core architecture - executors, tasks, stages, DAG.
  • Spark performance tuning techniques: partitioning, caching, broadcast joins, etc.
  • Troubleshooting slow running/stuck jobs or resource issues in Spark.
  • Experience optimizing Spark jobs for large-scale datasets.

Cloud Technologies

  • Exposure to AWS services like S3, EMR, Glue, Lambda, Athena, etc.
  • Use S3 with Spark (e.g., dealing with file formats, consistency issues).
  • EKS, Serverless knowledge, etc.

Programming - Python or Scala

  • Ability to write clean, modular, and performant code.
  • Functional programming concepts (e.g., immutability, higher-order functions).
  • Real-world use cases where they wrote scalable data processing code.
  • Understanding of collections, concurrency, and memory management.

Good to have:

  • Experience with managing production data pipelines/ETL systems.
  • Experience with CI/CD.
  • Experience writing test cases.
  • AWS certifications.

About the company

Software Guidance & Assistance, Inc., (SGA), is searching for a Big Data/Spark Engineer for a contract assignment with one of our premier Regulatory clients. Must be local to one of these office locations: Rockville MD or Tysons VA., SGA is a technology and resource solutions provider driven to stand out. We are a women-owned business. Our mission: to solve big IT problems with a more personal, boutique approach. Each year, we match consultants like you to more than 1,000 engagements. When we say let’s work better together, we mean it. You’ll join a diverse team built on these core values: customer service, employee development, and quality and integrity in everything we do. Be yourself, love what you do and find your passion at work. Please find us at https://sgainc.com/ .

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.disabledperson.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · World Congress 2024

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · World Congress 2025

Videos

See all

Related articles

See all