Data Engineer -w/ Spark, Python and AWS - Onsite work Bethesda, MD - No Remote
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+33 more
Job description
- Spark
- Hadoop
- Scala/Python
- EMR
- Distributed data processing and large-scale batch workloads
As a Senior Associate L2 - Full Stack Data Platform Support Engineer, you will be responsible for supporting and maintaining business-critical data platforms and cloud-native applications running on AWS. You will ensure the availability, stability, and performance of batch and real-time data processing systems by proactively monitoring production environments, troubleshooting complex issues, performing root cause analysis, and driving timely incident resolution. You will collaborate closely with SRE, Infrastructure, Development, and Business teams to support production deployments, optimize platform performance, improve operational efficiency, and implement automation to enhance system reliability and customer experience., * Provide L2 production support for enterprise data platforms running on AWS, ensuring high availability and reliability of business-critical applications.
- Monitor and support Big Data applications built on EMR (Spark/Hadoop), Kafka, Apache NiFi, EKS, Aurora PostgreSQL, DocumentDB, Couchbase, and S3.
- Investigate, troubleshoot, and resolve production incidents, application failures, performance bottlenecks, and infrastructure issues within defined SLAs.
- Perform root cause analysis (RCA) for recurring issues and implement permanent fixes in collaboration with engineering teams.
- Monitor batch and streaming data pipelines, ensuring timely completion of critical workflows and resolving data processing failures.
- Support workflow orchestration using Oozie and assist in maintaining scheduling dependencies and operational stability.
- Monitor and troubleshoot Kafka producers, consumers, topics, partitions, and message processing, including Kafka lag analysis.
- Support Kubernetes (Amazon EKS) workloads, monitor pod health, node utilization, and application deployments.
- Monitor EMR cluster health, auto-scaling, resource utilization, Spark job execution, and optimize performance when required.
- Support database administration activities for Aurora PostgreSQL, Amazon DocumentDB, and Couchbase, including query analysis and performance troubleshooting.
- Execute production deployments, configuration changes, and release validations through Jenkins and Harness CI/CD pipelines.
- Develop and maintain operational dashboards, monitoring alerts, and health checks using enterprise monitoring tools such as Dynatrace and CloudWatch.
-
Perform capacity planning, performance analysis, and recommend infrastructure optimization opportunities to improve system stability and reduce cloud costs.
- Work closely with cross-functional teams including SRE, Infrastructure, Development, Product, and Business teams during incident resolution and planned maintenance activities.
- Participate in on-call support, production releases, disaster recovery exercises, and major incident management.
- Identify automation opportunities to reduce manual operational effort and improve platform reliability.
- Ensure compliance with operational procedures, security standards, and change management processes
Requirements
12+ years of IT experience with 5+ years of experince in Data Engineer, * 6+ years of overall IT experience.
- 4+ years in Data Engineering and Big Data technologies.
- Strong experience developing Spark applications using Scala/Python.
- Experience building cloud-native applications on Kubernetes/EKS.
- Hands-on experience with Kafka-based streaming architecture.
- Experience designing scalable ETL/data ingestion pipelines using Apache NiFi.
- Good understanding of distributed systems, Hadoop ecosystem, and data lake architecture.
- Experience with relational and NoSQL databases.
- Experience in production support, performance optimization, and troubleshooting.
- Knowledge of CI/CD, Infrastructure as Code, and cloud security best practices.
Preferred Skills
- Experience with microservices architecture and REST APIs.
- Experience with Kubernetes deployments and containerized applications.
- Knowledge of EMR cluster administration and performance tuning.
- Exposure to Infrastructure as Code (Terraform/CloudFormation).
- AWS Certification (Associate or Professional) is preferred.
- Excellent analytical, debugging, and communication skills.
Benefits & conditions
Estimated Min Rate: $51.10 Estimated Max Rate: $73.00
What’s In It for You? We welcome you to be a part of the largest and legendary global staffing companies to meet your career aspirations. Yoh’s network of client companies has been employing professionals like you for over 65 years in the U.S., UK and Canada. Join Yoh’s extensive talent community that will provide you with access to Yoh’s vast network of opportunities and gain access to this exclusive opportunity available to you. Benefit eligibility is in accordance with applicable laws and client requirements. Benefits include:
- Medical, Prescription, Dental & Vision Benefits (for employees working 20+ hours per week)
- Health Savings Account (HSA) (for employees working 20+ hours per week)
- Life & Disability Insurance (for employees working 20+ hours per week)
- MetLife Voluntary Benefits
- Employee Assistance Program (EAP)
- 401K Retirement Savings Plan
- Direct Deposit & weekly epayroll
- Referral Bonus Programs
- Certification and training opportunities
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on jobs.localjobnetwork.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Highest Paying Tech Companies for Developers
Top Big Data Technologies That You Need to Know
Making Data Warehouses Fast: A Developer’s Story