> Markdown version of [/jobs/ext/3590574-data-engineer-data-pipelines-and-etl](https://www.wearedevelopers.com/jobs/ext/3590574-data-engineer-data-pipelines-and-etl). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer Data Pipelines and ETL - **Company:** CBS Corporation - **Location:** San Francisco, CA, United States - **Experience:** Experienced - **Salary:** $98,400.0 - $147,600.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Data Analysis, Cloud Computing, Cloud Database, Cloud Engineering, Computer Programming, Data Validation, Information Engineering, Data Integration, Extract Transform Load (ETL), Data Warehousing, Database Queries, Distributed Computing Environment, Distributed Systems, Python (Programming Language), Machine Learning, Metadata Standards, Operational Databases, Performance Tuning, Query Optimization, Cloud Services, Standard Sql, Azure Machine Learning, SQL Databases, Data Streaming, Unstructured Data, Data Processing, Apache Spark, Indexer, Usage Tracking, Data Layers, Event Driven Architecture, Data Lakes, Information Technology, Data Lineage, Apache Kafka, Data Management, Machine Learning Operations, Stream Processing, Data Pipelines, Databricks - **Published:** October 5, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/28079339/Data-Engineer-Data-Pipelines-Etl-California-San-Francisco-1403 ## About the Role Advanced Data Pipeline & ETL/ELT Expertise * 2-4+ years of experience building and scaling ETL/ELT pipelines in production environments. * Experience with workflow orchestration tools such as Airflow, Composer, or similar platforms. * Strong understanding of distributed data processing concepts. SQL & Data Modeling for Analytics & ML * Expert-level SQL skills for large-scale transformation and analytics. * Experience designing scalable warehouse schemas and ML-ready data layers. * Strong experience optimizing complex queries across multi-terabyte datasets. Programming & ML Data Integration * Proficiency in Python (or similar language) for data processing and ML pipeline integration. * Experience with distributed processing frameworks such as Spark. * Familiarity integrating data pipelines with ML platforms such as Vertex AI (preferred), Databricks ML, or equivalent. Streaming & Event-Driven Systems * Experience building real-time data pipelines using Kafka, Pub/Sub, or similar technologies. * Understanding of feature streaming, low-latency data processing, and event-driven architectures. * Ability to architect and build real-time dashboards using Superset. Cloud & Modern AI Data Platforms * Experience designing cloud-native data architectures (GCP preferred). * Experience with lakehouse architectures and cloud data warehouses. * Familiarity with vector databases, embeddings pipelines, and AI-serving infrastructure is a plus., * Bachelor's or Master's degree in Computer Science, Engineering, or related field (or equivalent experience). * 2-4+ years of experience in data engineering, data pipeline development, or related fields. * Strong foundation in modern data engineering principles, distributed systems design, and cloud-native architectures. * Demonstrated ability to design and operate large-scale production data systems. * Proven track record of technical leadership and cross-functional collaboration. * Strong problem-solving skills and ability to thrive in complex, fast-paced environments. * Detail-oriented and committed to engineering excellence and continuous improvement. ## Description Build and Maintain Scalable Data Pipelines * Design, develop, and maintain scalable batch and streaming data pipelines for large-scale structured and unstructured datasets. * Build robust ETL/ELT frameworks supporting analytics, BI, experimentation, and machine learning use cases. * Optimize pipelines for performance, reliability, scalability, and cost efficiency. * Implement advanced ingestion patterns including CDC, incremental loads, and event-driven processing. Data Modeling & Data Warehouse Architecture * Design scalable, dimensional, and hybrid data models optimized for analytics and ML use cases. * Develop reusable transformation layers (semantic layers) that serve BI, ML, and AI applications. * Write optimized, production-grade SQL for large-scale analytics workloads. * Contribute to query optimization, indexing, partitioning, and performance tuning across distributed systems and cloud warehouses. Modern Data Pipeline Development * Build and maintain modular data components following established framework patterns. * Contribute to architectural decisions across streaming systems, data lakes, and warehouses. Data Quality, Governance & Observability * Implement automated data validation, anomaly detection, and monitoring frameworks. * Establish data lineage and metadata standards to support reproducibility in ML workflows. * Enforce governance, privacy, and security best practices, particularly for sensitive AI datasets. * Ensure responsible AI data usage and compliance standards.