Data Engineer
Library & Archives Commission, Texas State
Dallas, United States of America
2 days ago
Role details
Contract type
Permanent contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
EnglishJob location
Dallas, United States of America
Tech stack
Java
Big Data
Computer Programming
Databases
Data Infrastructure
ETL
Query Languages
Linux
DevOps
Github
Hadoop
HBase
Hive
Python
NoSQL
Cloudera
Scala
SQL Databases
Teradata
Google Cloud Platform
Spark
PySpark
Requirements
- Linux
- Hadoop
- Hive
- HQL (Hive Query Language)
- Python
- Spark
- Pyspark
- Big data processing
- Google Cloud Platform (Google cloud platform)
Responsibilities:
- Software resource technical skill, such as; Pyspark,Python, Scala, Java, Hadoop hive, big data processing experience.
- Awareness and good to have on hands-on of Cloudera Data Platform, ETL, (ok if not all )
- Understanding of databases like Teradata, SQL, Hive, HBase, NoSQL, etc..
- Hands-on Pyspark/Python programming exposure and Good knowledge of Spark SQL.
- Strong in Software Programming/Engineering with a good understanding of DevOps, GitHub etc.
- Must know the Hadoop concept and open to learning new technology/toolset
- Learning attitude and flexible with project and timings
- Good communication and presentation skills.