Big Data Engineer

Vision Technologies, LLC
Sunnyvale, CA, United States
12 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$90,000.0 - $110,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Airflow Amazon Web Services Apache HTTP Server Microsoft Azure Big Data Cloud Computing Continuous Delivery Continuous Integration Information Engineering Data Governance
+33 more
Data Stores Database Queries Software Debugging Distributed Systems Fault Tolerance Apache Hadoop Hadoop Distributed File System Apache HBase Apache Hive Python (Programming Language) Machine Learning NoSQL Apache Oozie Software Engineering SQL Databases Sqoop Data Streaming Unstructured Data Data Processing Scripting Apache Spark Hdinsight Electronic Medical Records Kubernetes Information Technology Collibra Apache Flink Apache Kafka Spark Streaming Data Management Amazon Elastic Mapreduce (EMR) Databricks Programming Languages

Requirements

ingesting, transforming, and analyzing massive volumes of structured and unstructured data to support enterprise analytics, machine learning, and reporting workloads. The ideal candidate will combine deep technical expertise across the Hadoop ecosystem with strong software engineering fundamentals and a clear understanding of how to deliver reliable, performant, and cost-effective data platforms in production environments.Required Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
  • Five or more years of professional experience designing and operating big-data pipelines on Hadoop.
  • Strong hands-on expertise with Apache Spark (Scala, Python, or Java) in production environments.
  • Solid experience with Hive, HDFS, Sqoop, HBase, and the broader Hadoop ecosystem.
  • Hands-on experience with streaming data platforms such as Kafka, Spark Streaming, or Flink.
  • Strong SQL skills and experience working with both relational and NoSQL data stores.
  • Experience with workflow orchestration tools such as Airflow or Oozie.
  • Solid understanding of distributed systems concepts, including partitioning, replication, and fault tolerance.
  • Strong scripting skills in Python or Shell.
  • Excellent troubleshooting, debugging, and documentation skills.

Preferred Qualifications

  • Experience operating Hadoop on cloud platforms such as AWS EMR, Azure HDInsight, or Databricks.
  • Familiarity with modern lakehouse formats (Delta, Iceberg, Hudi).
  • Exposure to data governance tooling such as Apache Atlas or Collibra.
  • Experience with Kubernetes-based data platforms (Spark-on-K8s, Trino).
  • Hands-on experience with CI/CD and infrastructure-as-code in data engineering workflows., Amazon Web Services (AWS), Apache, Apache HBase, Apache Hadoop, Apache Hive, Apache Spark, Apache Sqoop, Artificial Intelligence (AI), Big Data, Cloud Computing, Computer Science, Consulting, Continuous Deployment/Delivery, Continuous Integration, Data Management, Data Processing, Debugging Skills, Distributed Computing, Documentation, EAD, Ecosystems, Electronic Medical Records, HDFS (Hadoop Distributed File System), High Tech Industry, Identify Issues, Java, Machine Learning, Machine Tool, Microsoft Windows Azure, NoSQL, Production Systems, Python Programming/Scripting Language, Replication and Remote Mirroring, SQL (Structured Query Language), Scala Programming Language, Software Development, Software Engineering, Structured Data, Unstructured Data

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · WWC 2024

Videos

See all

Related articles

See all