Senior AWS Data Engineer (Real-Time / Streaming)

Infinite Computer Solutions Inc
Atlanta, GA, United States
23 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Amazon Web Services Amazon S3 Cloud Computing Continuous Integration Data Integrity Extract Transform Load (ETL) Fraud Prevention and Detection Monitoring of Systems Python (Programming Language) Enterprise Messaging Systems Data Streaming Systems Architecture
+15 more
Systems Integration Amazon Connect Parquet Data Storage Technologies Delivery Pipeline Servicebus Event Driven Architecture Data Lakes Pyspark AWS Glue AWS Data Analytics Apache Kafka Terraform Stream Processing Data Pipelines

Job description

  • Develop a comprehensive plan for migrating near real-time fraud detection campaigns from on-premises systems to AWS.
  • Design and implement event-driven architectures to process inbound dialer data (fraud events) using services such as Amazon EventBridge, Kafka, Kinesis Data Streams, and Kinesis Firehose.
  • Build and manage scalable data pipelines using AWS Glue (ETL jobs, Crawlers), PySpark, and Python for data ingestion, transformation, and processing.
  • Configure and manage Glue Crawlers to automatically discover schemas and update the Data Catalog.
  • Store and optimize data using Parquet format and enable analytics through Amazon Athena for efficient querying.
  • Develop integrations between Customer Profiles and messaging platforms to automatically trigger profile updates and downstream processes.
  • Implement automation to trigger fraud-related outbound calls based on updates in customer profiles.
  • Design and orchestrate workflows using AWS Step Functions to manage complex processing pipelines.
  • Provision and manage cloud infrastructure using Terraform (Infrastructure as Code).
  • Optimize system architecture for scalability, reliability, cost-efficiency, and ensure data integrity and security.
  • Conduct end-to-end testing of the entire framework to validate functionality, performance, and reliability.
  • Deploy, automate, and manage resources using CI/CD pipelines.
  • Continuously monitor system performance and implement optimizations post-deployment.
  • Maintain detailed documentation of architecture, workflows, and operational processes.

Technical Skills

  • Strong expertise in AWS services including:

  • Lambda
  • S3
  • EventBridge
  • Kinesis (Data Streams & Firehose)
  • Glue (ETL + Crawlers)
  • Step Functions
  • Amazon Connect
  • Athena
  • Macie
  • Proficient in:

  • Python
  • PySpark

Requirements

  • Experience in:

  • Glue Crawlers for schema discovery and cataloging
  • Parquet-based data storage
  • Building scalable data pipelines
  • Strong understanding of event-driven architectures (Pub/Sub model)
  • Hands-on experience with Terraform (Infrastructure as Code)
  • Familiarity with CI/CD tools and automation pipelines

Preferred Qualifications

  • AWS Certifications:

  • AWS Certified Developer
  • AWS Solutions Architect
  • Experience with:

  • Messaging platforms like Kafka and Amazon EventBridge
  • Designing real-time data processing systems
  • Using Glue Crawlers + Athena for data lake architectures

About the company

Cargill is committed to providing food and agricultural solutions to nourish the world in a safe, responsible, and sustainable way. Sitting at the heart of the supply chain, we par…

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

55 sec

Validating data processing architectures via containerized events

Modood Alvi · World Congress 2025

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

1:43 min

AWS infrastructure stack and data flow pipeline overview

Artem Volk Artem Volk +1 · World Congress 2024

2:03 min

Introduction to open table formats built on Parquet

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

1:09 min

Evaluating mature stream processing frameworks for production systems

Soroosh Khodami Soroosh Khodami · World Congress 2024

Videos

See all

Related articles

See all