AWS Data Engineer

Visionary Innovative Technology Solutions LLC
United States
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Airflow Amazon Web Services Amazon S3 Microsoft Azure Cloud Computing Security Continuous Integration Information Engineering Extract Transform Load (ETL) Data Warehousing Cursor (Graphical User Interface Elements)
+25 more
Database Queries Database Theory Elasticsearch Identity and Access Management Python (Programming Language) Machine Learning NumPy SQL Databases Systems Integration Unstructured Data Data Processing Data Ingestion Sql Optimization GitHub Copilot Large Language Models Prompt Engineering Apache Spark Generative AI Git Pandas Data Lakes Pyspark Scikit Learn AWS Data Analytics Data Pipelines

Job description

  • Design and develop scalable data pipelines and data engineering solutions on AWS.
  • Build and maintain robust ETL/ELT pipelines using Python and AWS services.
  • Develop data ingestion, transformation, validation, and processing frameworks.
  • Design and implement enterprise AWS Data Lake and Data Warehouse architectures.
  • Work with AWS services such as S3, Glue, Lambda, EMR, Redshift, Athena, Kinesis, Step Functions, and ECS/EKS.
  • Develop complex data processing solutions using Python, PySpark, and SQL.
  • Design data models for structured, semi-structured, and unstructured datasets.
  • Optimize data pipelines for performance, scalability, reliability, and cost.
  • Implement data quality, governance, lineage, validation, and monitoring processes.
  • Develop orchestration workflows using Apache Airflow / AWS MWAA.
  • Build CI/CD pipelines for data engineering applications and infrastructure.
  • Collaborate with architects, data scientists, ML engineers, application developers, and business stakeholders., * Build data pipelines supporting Machine Learning and Generative AI use cases.
  • Prepare, clean, transform, and engineer datasets for AI/ML models.
  • Work with LLMs, embeddings, vector databases, and RAG pipelines.
  • Experience with Amazon Bedrock and foundation models is highly desirable.
  • Develop AI-enabled data processing and automation solutions.
  • Integrate pre-trained AI/ML models through APIs.
  • Support model training, evaluation, deployment, and monitoring workflows.
  • Use Python, Pandas, NumPy, Scikit-learn, and other AI/ML libraries as needed.
  • Work with AI coding assistants such as GitHub Copilot, Amazon Q Developer, or Cursor.

Requirements

  • 14+ years of overall IT experience with strong Data Engineering background.
  • Extensive hands-on experience with AWS Data Engineering.
  • Strong Python programming skills.
  • Advanced SQL skills including complex queries, CTEs, window functions, and optimization.
  • Strong experience with AWS S3, Glue, Redshift, Lambda, EMR, and Athena.
  • Strong experience designing ETL/ELT pipelines.
  • Hands-on experience with PySpark / Apache Spark.
  • Experience with data lake and data warehouse architecture.
  • Experience with Airflow / AWS MWAA.
  • Strong understanding of data modeling and database concepts.
  • Experience with Git and CI/CD.
  • Strong understanding of cloud security, IAM, encryption, and AWS best practices.

AI/GenAI Skills - Preferred

  • Generative AI / LLM experience.
  • Amazon Bedrock.
  • RAG architecture.
  • Prompt Engineering.
  • Embeddings and Vector Databases.
  • LangChain / LlamaIndex.
  • OpenAI / Azure OpenAI / Anthropic.
  • Pinecone / OpenSearch / Elasticsearch / FAISS.
  • Machine Learning and model integration.
  • AI/ML API integration.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · WWC 2024

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:43 min

AWS infrastructure stack and data flow pipeline overview

Artem Volk Artem Volk +1 · WWC 2024

Videos

See all

Related articles

See all