Cloud Data Engineer(Healthcare)

Helishores Inc
Austin, United States of America
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Austin, United States of America

Tech stack

Agile Methodologies
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Azure
Big Data
Unix
Data Mining
Data Warehousing
Database Development
DevOps
Hive
Identity and Access Management
Korn Shell
Performance Tuning
Scrum
Standard Sql
Shell Script
Amazon Web Services (AWS)
Software Deployment
SQL Stored Procedures
SQL Databases
Teradata
Data Processing
Spark
Ab Initio
GIT
Data Lake
PySpark
Atlassian Tools
Functional Programming
Cloudwatch
Amazon Web Services (AWS)
Jenkins
Databricks
Artifactory

Job description

  • 3 days a week from Austin, TX office
  • Work with business and technical leadership to understand requirements.
  • Design to the requirements and document the designs.
  • Ability to write product-grade performant code for data extraction, transformations and loading using Spark, PySpark.
  • Do data modeling as needed for the requirements.
  • Write performant queries using Teradata SQL, Hive SQL and Spark SQL against Teradata and DataBricks Unity Catalog.
  • Implementing dev-ops pipelines to deploy code artifacts on to the designated platform/servers like AWS or DataBricks.
  • Implement DataBricks job orchestration using DataBricks.
  • Troubleshooting the issues, providing effective solutions and jobs monitoring in the production environment.
  • Participate in sprint planning sessions, refinement/story-grooming sessions, daily scrums, demos and retrospectives.

Requirements

  • Strong development experience in Spark, PySpark, Shell scripting, Teradata, and DataBricks.

  • Experience with Ab Initio is required.

  • Strong experience in writing complex and effective SQLs (using Teradata SQL, Hive SQL and Spark SQL) and Stored Procedures.

  • Proficiency in Shell Scripting for automation, system administration, and data processing tasks.

  • Excellent work experience on DataBricks as data warehouse/Data Lake implementations.

  • Experience in Agile and working knowledge on DevOps tools (Git, Jenkins, Artifactory) Unix/Linux Shell scripting (KSH) and basic administration of Unix servers CA7 Enterprise Scheduler AWS (S3, EC2, SNS, SQS, Lambda, ECS, Glue, IAM, and CloudWatch) Databricks (Delta lake, Notebooks, Pipelines, cluster management, Azure/AWS integration).

  • Hands-on experience working with Teradata, including data modeling, SQL development, performance tuning, and large-scale data management.

  • Experience in Jira and Confluence Exercises considerable creativity, foresight, and judgment in conceiving, planning, and delivering initiatives.

  • Strong experience developing large-scale data processing solutions using Apache Spark and PySpark. Desired:

  • Health care domain knowledge.

Apply for this position