> Markdown version of [/jobs/ext/2830716-big-data-engineer](https://www.wearedevelopers.com/jobs/ext/2830716-big-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Big Data Engineer - **Company:** TP-Link Systems Inc. - **Location:** Irvine, CA, United States - **Experience:** Experienced - **Salary:** $100,000.0 - $160,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Big Data, Information Systems, Extract Transform Load (ETL), Data Warehousing, Database Development, Software Debugging, Linux, Dimensional Modeling, Apache Hadoop, Apache Hive, Identity and Access Management, Python (Programming Language), Online Analytical Processing, Online Transaction Processing, Query Optimization, SQL Databases, Sqoop, Data Streaming, Parquet, Data Logging, Data Ingestion, Apache Spark, Boto3, Git, Build Management, Pyspark, Debezium, Information Technology, Data Lineage, Apache Flink, Apache Kafka, Spark Streaming, Vertica - **Published:** September 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=39ca6ce91460b142 ## About the Role * Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent practical experience. * 2+ years of hands-on data development in a production environment - scheduled pipelines with real downstream consumers. * Experience owning a pipeline or component in production: you were the person accountable for it, including when it broke. * Strong SQL: window functions, complex multi-table joins, incremental and idempotent writes; able to read an execution plan and diagnose data skew or slow queries. * Production-grade Python: maintainable, testable ETL and tooling code (PySpark / pandas / boto3) with proper error handling and logging. * Hands-on with Spark on EMR, or equivalent Hadoop / Hive experience, including partitioning, shuffle, and memory tuning. * Production experience with Airflow: DAG design, dependencies, retries, idempotent reruns, backfills, and SLA alerting. * Experience with a data ingestion or CDC tool in production (DataX, Sqoop, Debezium, or similar): extracting from OLTP sources, incremental sync, and handling upstream schema changes. * Experience with an OLAP engine - StarRocks, Doris, or ClickHouse: table models, partitioning and bucketing, materialized views, and query tuning. * Solid grasp of data warehouse modeling: dimensional modeling, slowly changing dimensions, and layering conventions. * Comfortable on AWS (S3 with Parquet / ORC and sensible partitioning, IAM basics, day-to-day EMR operations), Linux, and Git. * Comfortable picking up new tools as the platform evolves - we'd rather hire someone who learns a new engine quickly than someone who has only ever used ours. * Effective use of AI to solve data problems - using it to move faster on SQL, debugging, and unfamiliar schemas, with the judgment to catch output that looks right but isn't., * AWS cost optimization: EMR instance sizing and Spot strategy, S3 lifecycle policies, StarRocks vs. Athena trade-offs. * Lakehouse table formats: Iceberg, Hudi, or Delta. * Streaming: Kafka with Flink or Spark Structured Streaming. * dbt, Glue Data Catalog, data lineage, or a data quality framework. * QuickSight dataset, SPICE, and row-level permission design. ## Description * Own a component of the data platform end to end - its pipelines, tables, quality checks, monitoring, and recovery. * Design and build ETL on EMR (Spark) and Airflow, and data ingestion from OLTP sources via DataX / CDC, including incremental sync and upstream schema changes. * Model the data for your area: dimensional models and layering (ODS / DWD / DWS / ADS) that serve real analytical use cases. * Own data quality for your component - write the checks and freshness monitoring, and resolve issues before they reach downstream users. * Tune runtime and infrastructure cost for the jobs you own, and choose the right engine for a workload (StarRocks serving vs. batch on EMR). * Build components other engineers can reuse, and automate work that would otherwise repeat. * Work directly with analysts and business stakeholders to turn requirements into models and datasets that actually get used. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Stop Committing Your Secrets - GIt Hooks To The Rescue!](https://www.wearedevelopers.com/videos/573-stop-committing-your-secrets-git-hooks-to-the-rescue) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk)