> Markdown version of [/jobs/ext/1968756-data-engineer-w-spark-python-and-aws-onsite-work-bethesda-md-no-remote](https://www.wearedevelopers.com/jobs/ext/1968756-data-engineer-w-spark-python-and-aws-onsite-work-bethesda-md-no-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer -w/ Spark, Python and AWS - Onsite work Bethesda, MD - No Remote - **Company:** Yoh Services LLC - **Location:** Bethesda, MD, United States - **Salary:** $106,288.0 - $151,840.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon S3, Big Data, Cloud Computing Security, Continuous Integration, Couchbase Servers, Information Engineering, Data Infrastructure, Extract Transform Load (ETL), Software Debugging, Disaster Recovery, Distributed Computing Environment, Distributed Systems, Apache Hadoop, Python (Programming Language), PostgreSQL, NoSQL, Apache Oozie, Performance Tuning, Software Deployment, Enterprise Data Management, Data Processing, Spring Cloud, Autoscaling, System Availability, Apache Spark, Software Troubleshooting, Reliability of Systems, Cloudformation, Containerization, Data Lakes, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Apache Kafka, Apache Nifi, Data Management, Cloudwatch, Restful APIs, Terraform, Stream Processing, Data Pipelines, Dynatrace, Jenkins, Microservices - **Published:** August 7, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/87928133/1 ## About the Role 12+ years of IT experience with 5+ years of experince in Data Engineer, * 6+ years of overall IT experience. * 4+ years in Data Engineering and Big Data technologies. * Strong experience developing Spark applications using Scala/Python. * Experience building cloud-native applications on Kubernetes/EKS. * Hands-on experience with Kafka-based streaming architecture. * Experience designing scalable ETL/data ingestion pipelines using Apache NiFi. * Good understanding of distributed systems, Hadoop ecosystem, and data lake architecture. * Experience with relational and NoSQL databases. * Experience in production support, performance optimization, and troubleshooting. * Knowledge of CI/CD, Infrastructure as Code, and cloud security best practices. Preferred Skills * Experience with microservices architecture and REST APIs. * Experience with Kubernetes deployments and containerized applications. * Knowledge of EMR cluster administration and performance tuning. * Exposure to Infrastructure as Code (Terraform/CloudFormation). * AWS Certification (Associate or Professional) is preferred. * Excellent analytical, debugging, and communication skills. ## Description + Spark + Hadoop + Scala/Python + EMR + Distributed data processing and large-scale batch workloads As a Senior Associate L2 - Full Stack Data Platform Support Engineer, you will be responsible for supporting and maintaining business-critical data platforms and cloud-native applications running on AWS. You will ensure the availability, stability, and performance of batch and real-time data processing systems by proactively monitoring production environments, troubleshooting complex issues, performing root cause analysis, and driving timely incident resolution. You will collaborate closely with SRE, Infrastructure, Development, and Business teams to support production deployments, optimize platform performance, improve operational efficiency, and implement automation to enhance system reliability and customer experience., * Provide L2 production support for enterprise data platforms running on AWS, ensuring high availability and reliability of business-critical applications. * Monitor and support Big Data applications built on EMR (Spark/Hadoop), Kafka, Apache NiFi, EKS, Aurora PostgreSQL, DocumentDB, Couchbase, and S3. * Investigate, troubleshoot, and resolve production incidents, application failures, performance bottlenecks, and infrastructure issues within defined SLAs. * Perform root cause analysis (RCA) for recurring issues and implement permanent fixes in collaboration with engineering teams. * Monitor batch and streaming data pipelines, ensuring timely completion of critical workflows and resolving data processing failures. * Support workflow orchestration using Oozie and assist in maintaining scheduling dependencies and operational stability. * Monitor and troubleshoot Kafka producers, consumers, topics, partitions, and message processing, including Kafka lag analysis. * Support Kubernetes (Amazon EKS) workloads, monitor pod health, node utilization, and application deployments. * Monitor EMR cluster health, auto-scaling, resource utilization, Spark job execution, and optimize performance when required. * Support database administration activities for Aurora PostgreSQL, Amazon DocumentDB, and Couchbase, including query analysis and performance troubleshooting. * Execute production deployments, configuration changes, and release validations through Jenkins and Harness CI/CD pipelines. * Develop and maintain operational dashboards, monitoring alerts, and health checks using enterprise monitoring tools such as Dynatrace and CloudWatch. * Perform capacity planning, performance analysis, and recommend infrastructure optimization opportunities to improve system stability and reduce cloud costs. * Work closely with cross-functional teams including SRE, Infrastructure, Development, Product, and Business teams during incident resolution and planned maintenance activities. * Participate in on-call support, production releases, disaster recovery exercises, and major incident management. * Identify automation opportunities to reduce manual operational effort and improve platform reliability. * Ensure compliance with operational procedures, security standards, and change management processes ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Our GitOps approach for deploying an Identity Provider and an API Gateway in a SaaS company](https://www.wearedevelopers.com/videos/776-our-gitops-approach-for-deploying-an-identity-provider-and-an-api-gateway-in-a-saas-company) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)