Data Engineer

Tata Consultancy Services Limited
Atlanta, GA, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$100,000.0 - $130,000.0
Working hours
Regular working hours
Job source

Tech stack

Tibco Ems Amazon Web Services Amazon S3 Microsoft Azure Big Data BigQuery Cloud Computing Information Engineering Data Files Data Infrastructure Extract Transform Load (ETL) Data Systems
+32 more
IBM DB2 Database Development Database Queries IBM InfoSphere DataStage Apache Hadoop Hadoop Distributed File System Apache Hive IBM WebSphere MQ Python (Programming Language) Machine Learning Mainframes Enterprise Messaging Systems Scrum Methodology Power BI Sqoop Tableau (Software) Teradata SQL Unstructured Data Apache Spark Git Data Lakes Pyspark Data Analytics Apache Kafka Spark Streaming Data Management Spotfire Stream Processing Data Pipelines Serverless Computing TIBCO (Software) Databricks

Job description

  • Defines data requirements, gather, and wrangle large scale of structured and unstructured data, and validate data by running various data tools in the Data Environment.
  • Supports the standardization, customization, and ad-hoc data analysis, and will develop the mechanisms to ingest, analyze, validate, normalize and clean data.
  • Creates data policy and develop interfaces and retention models which requires synthesizing or anonymizing data.
  • Implements statistical data quality procedures on new data sources, and by applying rigorous iterative data analytics, supports Data Scientists and analytics and insights creation in data sourcing and preparation to visualize data and synthesize insights of commercial value.
  • Develops and maintains data engineering best practices and contributes to Insights on data analytics and visualization concepts, methods and techniques.
  • Works closely with the data science and business intelligence teams to develop data models and pipelines for research, reporting, and machine learning.
  • Design, implement, and support scalable data infrastructure solutions to integrate with multi-heterogeneous data sources, aggregate and retrieve Big Data in a fast and safe mode, curate data that can be used in BI reporting, analysis, machine learning models and ad-hoc data requests.
  • Build data pipelines that clean, transform, and aggregate data from disparate sources.
  • Engages with business teams to gather requirements and design data solutions.
  • Mentors team of more Junior Data Engineers.
  • Collaborates across multiple projects to provide data engineering expertise across teams.
  • Analyzes most relevant insights and shares with leadership to provide strategic recommendations for the business
  • Lead a team of data engineers and act as a key senior contributor to a data engineering project.

Requirements

Do you have experience in Research?, * 7+ years of overall IT experience

  • 5+ years of experience in a data engineering/ETL role with a track record of manipulating, processing, and extracting value from large datasets
  • 3+ years of experience with Big Data tools/technologies like Hadoop, Spark, Spark SQL, Kafka, Sqoop, Hive, S3, HDFS, or Cloud platforms e.g. AWS, GCP, etc.
  • 3+ years building, testing, and optimizing data ingestion pipelines, architectures, and data sets with Tibco, IBM or others.
  • Databricks UI, Managing Databricks Notebooks, Delta Lake with Python, Delta Lake with Spark SQL, Delta Live Tables, Unity Catalog.
  • High-velocity high-volume stream processing with Apache Kafka and Spark Streaming.
  • Strong SQL skills with ability to write intermediate complexity queries.
  • ETL experience with PySpark, Spark SQL , IBM Data Stage or similar.
  • Agile Scrum, Kanban or SAFe experience.

Skills Desired

  • Databricks, Python (and/or Scala) and PySpark/Scala-Spark.
  • Database solutions like Databricks, Teradata, Mainframe, DB2 or BigQuery.
  • BI Solutions like Spotfire, OAC, Tableau or PowerBI
  • Azure, AWS Serverless technologies, like, S3, Kinesis/MSK, lambda, and Glue.
  • Messaging Platforms like Kafka, Amazon MSK & TIBCO EMS or IBM MQ Series.
  • Strong SQL skills with ability to write intermediate complexity queries
  • Experience with GIT code versioning software

Benefits & conditions

(part of Tata group) 3.93.9 out of 5 stars Atlanta, GA $100,000 - $130,000 a year, Pulled from the full job description

  • Pet insurance
  • Health insurance
  • Vision insurance
  • Dental insurance
  • Commuter assistance, Salary Range-$100,000-$130,000 a year #LI-KR3 TCS Employee Benefits Summary: Discretionary Annual Incentive. Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans. Family Support: Maternal & Parental Leaves. Insurance Options: Auto & Home Insurance, Identity Theft Protection. Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimbursement. Time Off: Vacation, Time Off, Sick Leave & Holidays. Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all