Data Engineer

AI AI LLC
San Jose, CA, United States
15 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
1 year minimum
Working hours
Regular working hours

Tech stack

Third Normal Form Application Programming Interfaces (APIs) Agile Methodology Airflow Amazon Web Services Amazon S3 JIRA Microsoft Azure BigQuery Data Architecture Information Engineering Data Governance
+37 more
Data Integration Extract Transform Load (ETL) Data Transformation Data Systems Relational Databases Database Queries DevOps Data Flow Control Github Python (Programming Language) NoSQL Performance Tuning Scrum Methodology Systems Development Life Cycle Role-Based Access Control Cloudera SQL Databases Data Streaming Systems Integration Google Cloud Cloud Platform System Delivery Pipeline Snowflake Apache Spark Cloudformation Data Lakes Pyspark Git Flow Google Cloud Functions Data Analytics Star Schema Apache Kafka Terraform Stream Processing Stream Analytics Data Pipelines Databricks

Requirements

Are you an experienced Data Engineering professional with a passion for building scalable, reliable, and high-performance data systems? Do you have hands-on experience designing and optimizing end-to-end real-time and batch pipelines, and developing cloud-native data architectures using modern technologies such as AWS, GCP, Azure, Databricks, and Snowflake?

We are looking for a Senior Data Engineer to architect, design, and implement scalable, high-performance data solutions. The ideal candidate will be an expert in at least one major cloud data ecosystem (AWS, Azure, GCP, Snowflake, or Databricks) and possess a deep understanding of the end-to-end data lifecycle, from ingestion to business intelligence. Qualification & Skill Set Requirements Core Technical Competencies Experience: 5+ years of hands-on data engineering experience in a production environment. Languages: Strong proficiency in Python, SQL (complex queries, performance tuning), and PySpark/Apache Spark. Data Modeling: Expert knowledge of data modeling (3NF, Star, Snowflake Schema) and Lakehouse/Warehouse architectures. ETL/ELT & Orchestration: Proven experience building pipelines using tools like dbt, Airflow, Dagster, or native cloud orchestrators (Glue, Data Factory, Composer). Integrations: Experienced in integrating data from diverse sources: APIs, RDBMS/NoSQL databases, flat files, and streaming platforms (Kafka, Kinesis, Pub/Sub). Cloud Platform Expertise (Specialization-Specific) Candidates should demonstrate deep expertise in anyone of the following: Snowflake: SnowSQL, Streams, Tasks, Snowpark, and cost optimization. Databricks: Delta Lake, Unity Catalog, Delta Live Tables (DLT), and Spark optimization. GCP: BigQuery, Dataflow, Dataproc, Pub/Sub, and Cloud Functions. Azure: Synapse Analytics, Data Factory, Azure Databricks, and Stream Analytics. AWS: Redshift, S3, Lake Formation, Glue, and Lambda. Professional Practices SDLC & DevOps: Proficient in Git workflows, CI/CD pipelines (GitHub Actions, Azure DevOps, AWS CodePipeline), and IaC (Terraform/CloudFormation). Data Governance: Strong understanding of data quality, lineage, observability, security (RBAC, encryption), and compliance frameworks. Agile: Active experience in Agile/Scrum environments using Jira or Azure Boards. Mentorship: Ability to lead projects and provide technical guidance to junior/mid-level engineers. Responsibilities Architecture: Architect, design, and implement scalable, reliable data solutions and pipelines aligned with business analytics needs. Optimization: Manage and fine-tune cloud resources and workloads for maximum performance, reliability, and cost-efficiency. Data Transformation: Lead the development of ETL/ELT processes for both batch and real-time data processing. Collaboration: Partner with Product, Engineering, and Data Science teams to deliver effective, data-driven solutions. Governance & Quality: Promote and enforce best practices in data governance, security, and data quality frameworks. Mentorship: Provide technical leadership and mentorship to the team, ensuring architecture quality and best practices. Documentation: Maintain comprehensive documentation of data architectures, configurations, and workflows. Fusemachines is an Equal Opportunities Employer, committed to diversity and inclusion. All qualified applicants will receive, Are you an experienced Data Engineering professional with a passion for building scalable, reliable, and high-performance data systems? Do you have hands-on experience designing and optimizing end-to-end real-time and batch pipelines, and developing cloud-native data architectures using modern technologies such as AWS, GCP, Azure, Databricks, and Snowflake?

We are looking for a Senior Data Engineer to architect, design, and implement scalable, high-performance data solutions. The ideal candidate will be an expert in at least one major cloud data ecosystem (AWS, Azure, GCP, Snowflake, or Databricks) and possess a deep understanding of the end-to-end data lifecycle, from ingestion to business intelligence. Qualification & Skill Set Requirements Core Technical Competencies Experience: 5+ years of hands-on data engineering experience in a production environment. Languages: Strong proficiency in Python, SQL (complex queries, performance tuning), and PySpark/Apache Spark. Data Modeling: Expert knowledge of data modeling (3NF, Star, Snowflake Schema) and Lakehouse/Warehouse architectures. ETL/ELT & Orchestration: Proven experience building pipelines using tools like dbt, Airflow, Dagster, or native cloud orchestrators (Glue, Data Factory, Composer). Integrations: Experienced in integrating data from diverse sources: APIs, RDBMS/NoSQL databases, flat files, and streaming platforms (Kafka, Kinesis, Pub/Sub). Cloud Platform Expertise (Specialization-Specific) Candidates should demonstrate deep expertise in anyone of the following: Snowflake: SnowSQL, Streams, Tasks, Snowpark, and cost optimization. Databricks: Delta Lake, Unity Catalog, Delta Live Tables (DLT), and Spark optimization. GCP: BigQuery, Dataflow, Dataproc, Pub/Sub, and Cloud Functions. Azure: Synapse Analytics, Data Factory, Azure Databricks, and Stream Analytics. AWS: Redshift, S3, Lake Formation, Glue, and Lambda. Professional Practices SDLC & DevOps: Proficient in Git workflows, CI/CD pipelines (GitHub Actions, Azure DevOps, AWS CodePipeline), and IaC (Terraform/CloudFormation). Data Governance: Strong understanding of data quality, lineage, observability, security (RBAC, encryption), and compliance frameworks. Agile: Active experience in Agile/Scrum environments using Jira or Azure Boards. Mentorship: Ability to lead projects and provide technical guidance to junior/mid-level engineers. Responsibilities Architecture: Architect, design, and implement scalable, reliable data solutions and pipelines aligned with business analytics needs. Optimization: Manage and fine-tune cloud resources and workloads for maximum performance, reliability, and cost-efficiency. Data Transformation: Lead the development of ETL/ELT processes for both batch and real-time data processing. Collaboration: Partner with Product, Engineering, and Data Science teams to deliver effective, data-driven solutions. Governance & Quality: Promote and enforce best practices in data governance, security, and data quality frameworks. Mentorship: Provide technical leadership and mentorship to the team, ensuring architecture quality and best practices. Documentation: Maintain comprehensive documentation of data architectures, configurations, and workflows. Fusemachines is an Equal Opportunities Employer, committed to diversity and inclusion. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or any other characteristic protected by applicable federal, state, or local laws.

About the company

Founded in 2013, Fusemachines is a global provider of enterprise AI products and services, on a mission to democratize AI. Leveraging proprietary AI Studio and AI Engines, the company helps drive the clients’ AI Enterprise Transformation, regardless of where they are in their Digital AI journeys. With offices in North America, Asia, and Latin America, Fusemachines provides a suite of enterprise AI offerings and specialty services that allow organizations of any size to implement and scale AI. Fusemachines serves companies in industries such as retail, manufacturing, and government. Fusemachines continues to actively pursue the mission of democratizing AI for the masses by providing high-quality AI education in underserved communities and helping organizations achieve their full potential with AI. Type: Remote Full-time

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.workingnomads.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

5:47 min

Integrating user stories and test automation via Jira tools

Christoph Ruggenthaler · LIVE

Videos

See all

Related articles

See all