> Markdown version of [/jobs/ext/592955-senior-data-engineer-chinese-mandarin-speaker](https://www.wearedevelopers.com/jobs/ext/592955-senior-data-engineer-chinese-mandarin-speaker). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer (Chinese Mandarin Speaker) - **Company:** Bitus Labs - **Location:** Irvine, CA, United States - **Experience:** Expert - **Salary:** $130,000.0 - **Contract:** Permanent contract - **Skills:** Query Performance, Java (Programming Language), Airflow, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Apache HTTP Server, Code Review, Continuous Integration, Data Validation, Information Engineering, Data Governance, Data Infrastructure, Extract Transform Load (ETL), Data Transformation, Data Security, DevOps, Digital Architecture, Memory Management, Github, Gradle, Identity and Access Management, Python (Programming Language), Apache Maven, Online Analytical Processing, Operational Databases, Pair Programming, Performance Tuning, Query Optimization, Azure Machine Learning, SQL Databases, Data Streaming, AWS Cdk, Sql Optimization, Delivery Pipeline, Apache Spark, Boto3, Amazon Virtual Private Cloud (VPC), Cloudformation, Pandas, Build Management, Data Lakes, Pyspark, Data Lineage, Druid, Apache Flink, Production Code, AWS Glue, Integration Frameworks, Apache Kafka, Build Tools, Spark Streaming, Machine Learning Operations, Data Lakehouse, Vertica, Terraform, Data Pipelines, Programming Languages - **Published:** June 18, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=38690ace7c9a024c ## About the Role * 5+ years of professional data engineering experience, with at least 3 years on AWS cloud platforms. * Proven track record of delivering production data pipelines at scale (TB+ datasets, highthroughput SLAs). * Experience with data lakehouse architectures - medallion pattern, open table formats (Iceberg preferred; Delta Lake or Hudi acceptable). Programming Languages * Java: Strong command of Java (8+) for Spark jobs, custom Iceberg connectors, and performance-critical pipeline components. Familiarity with Maven/Gradle build systems. * Python: Proficient in Python 3 for AWS Glue scripts, orchestration logic, data quality checks, and automation tooling. Experience with pandas, PySpark, boto3, and packaging best practices. AWS Core Services * Storage & Compute: S3, Glue (jobs, crawlers, Data Catalog), EMR (Spark/Flink), Lambda, EC2. * Streaming: Kinesis Data Streams, Kinesis Firehose, or MSK (Managed Kafka). * Orchestration: Step Functions, MWAA (Managed Airflow), or EventBridge Scheduler. * Querying: Athena, Redshift, or Redshift Spectrum. * Security & Governance: IAM, KMS, Lake Formation, Secrets Manager, VPC. * DevOps: AWS CDK or CloudFormation; CodePipeline or equivalent CI/CD tools. Data Processing Frameworks * Apache Spark (PySpark and/or Spark Java API) - distributed transformations, performance tuning, memory management. * Apache Iceberg - table maintenance, time travel, snapshot management, partition evolution. * SQL - advanced SQL for data transformation, window functions, CTEs, query optimization. Preferred / Nice to Have * AWS Certified Data Engineer - Associate or AWS Certified Solutions Architect certification. * Experience with dbt for SQL-based transformation layers on top of the lakehouse. * Familiarity with ML platform integration: feature stores (SageMaker Feature Store), model serving data needs, or MLflow experiment tracking. * Experience with real-time OLAP engines such as Apache Druid or ClickHouse. * Contributions to open-source data tooling or internal platform libraries. * Exposure to data mesh or data product thinking - defining domain ownership and data contracts. Tech Stack at a Glance Languages Java (8+), Python 3 Cloud Platform AWS (S3, Glue, EMR, Kinesis, Athena, Lambda, Step Functions, Lake Formation, CDK) Processing Apache Spark, Apache Flink, Spark Structured Streaming Table Format Apache Iceberg (primary), Delta Lake / Hudi (familiarity) Streaming Amazon Kinesis, MSK (Kafka), Kinesis Firehose Orchestration Apache Airflow (MWAA), AWS Step Functions IaC & CI/CD AWS CDK / Terraform, GitHub Actions / CodePipeline ## Description We are looking for a Senior Data Engineer to join our Data Platform team and take ownership of building and scaling our AWS-based data lakehouse. You will architect and deliver robust, production-grade data pipelines, work closely with data scientists, analytics engineers, and product teams, and set the technical direction for how data flows across the organization. This is a hands-on engineering role - you will write production code in Java and Python every day, while also contributing to platform design decisions, mentoring junior engineers, and driving best practices around data quality, reliability, and governance., Data Lakehouse Architecture & Development * Design and build scalable medallion-architecture data lakehouses (Bronze / Silver / Gold) on AWS S3 using Apache Iceberg table format. * Develop and maintain high-throughput ETL/ELT pipelines using AWS Glue, EMR (Spark), and Lambda. * Implement schema evolution, partitioning strategies, and compaction processes for Iceberg tables to optimize storage and query performance. * Write production-quality pipeline code in both Java and Python, selecting the appropriate language per performance and maintainability requirements. Real-Time & Batch Streaming * Build and operate event-driven data pipelines using Amazon Kinesis Data Streams, Kinesis Firehose, or Apache Kafka (MSK). * Design exactly-once and at-least-once processing semantics for streaming workloads using Apache Flink or Spark Structured Streaming on EMR. AWS Platform Engineering * Manage infrastructure as code using AWS CDK or Terraform for repeatable, auditable data platform deployments. * Optimize cost and performance across AWS services including S3, Glue, Athena, Redshift Spectrum, EMR, Lambda, Step Functions, and EventBridge. * Implement data security best practices: IAM least-privilege policies, KMS encryption, VPC networking, and Lake Formation fine-grained access control. * Build and maintain CI/CD pipelines for data workloads using AWS CodePipeline, GitHub Actions, or equivalent. Data Quality & Governance * Implement data quality frameworks (e.g., Great Expectations, Deequ) and integrate validation steps into pipeline orchestration. * Define and enforce data contracts between producing and consuming systems. * Contribute to data cataloguing and lineage tracking using AWS Glue Data Catalog or Apache Atlas. Collaboration & Technical Leadership * Partner with data scientists, ML engineers, and analysts to understand data requirements and deliver performant, well-documented datasets. * Mentor mid-level and junior engineers through code reviews, design discussions, and pair programming. * Document architecture decisions (ADRs) and contribute to internal engineering knowledge base. ## Related Videos - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Getting to Know Your Legacy (System) with AI-Driven Software Archeology](https://www.wearedevelopers.com/videos/1437-getting-to-know-your-legacy-system-with-ai-driven-software-archeology) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)