> Markdown version of [/jobs/ext/2120757-data-engineer-self-service-analytics-real-time-data-platforms](https://www.wearedevelopers.com/jobs/ext/2120757-data-engineer-self-service-analytics-real-time-data-platforms). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - Self-Service Analytics & Real-Time Data Platforms - **Company:** Paramount Skydance Corporation - **Location:** San Francisco, CA, United States - **Experience:** Experienced - **Salary:** $99,000.0 - $147,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Business Analytics Applications, Data Analysis, Computing Platforms, Cloud Computing, Cloud Database, Cloud Engineering, Computer Programming, Databases, Information Engineering, Data Files, Data Integration, Extract Transform Load (ETL), Data Security, Data Systems, Data Warehousing, Database Design, Database Queries, Distributed Computing Environment, Distributed Systems, Python (Programming Language), Machine Learning, Metadata, Operational Databases, Systems Development Life Cycle, Query Optimization, Power BI, Cloud Services, Standard Sql, Azure Machine Learning, Software Engineering, SQL Databases, Data Streaming, Data Processing, Scripting, Apache Spark, Data Layers, Event Driven Architecture, Information Technology, Real Time Data, Apache Kafka, Data Management, Video Streaming, Data Pipelines, Databricks - **Published:** August 19, 2026 - **Apply:** https://www.careerbuilder.com/job-details/data-engineer-self-service-analytics-and-real-time-data-platforms-san-francisco-ca--697fbd2d-ae92-42ae-9928-c638aebea79d ## About the Role Advanced Data Pipeline & ETL/ELT Expertise * 2-4+ years of experience building and scaling ETL/ELT pipelines in production environments. * Proven experience with workflow orchestration tools such as Airflow, Composer, or similar platforms. * Working knowledge of distributed data processing concepts. * SQL & Data Modeling for Analytics & ML * Expert-level SQL skills for large-scale transformation and analytics. * Experience designing scalable warehouse schemas and ML-ready data layers. * Proven experience optimizing complex queries across multi-terabyte datasets. Programming & ML Data Integration * Proficiency in Python (or similar language) for data processing and ML pipeline integration. * Experience with distributed processing frameworks such as Spark. * Experience integrating data pipelines with ML platforms such as Vertex AI (preferred), Databricks ML, or equivalent. This includes model training, batch/online inference, and pipeline orchestration. Streaming & Event-Driven Systems * Experience building real-time data pipelines using Kafka, Pub/Sub, or similar technologies. * Knowledge of feature streaming, low-latency data processing, and event-driven architectures. * Ability to work closely with the streaming team to architect and build real-time dashboards using Superset. Cloud & Modern AI Data Platforms * Experience designing cloud-native data architectures (GCP preferred). * Experience with lakehouse architectures and cloud data warehouses. * Knowledge of vector databases, embeddings pipelines, and AI-serving infrastructure is a plus., * Bachelors or Masters degree in Computer Science, Engineering, or a related field (or equivalent experience). * 2-4+ years of experience in data engineering, data pipeline development, or related fields. * Solid foundation in modern data engineering principles, distributed systems design, and cloud-native architectures. * Demonstrated ability to design and operate large-scale production data systems. * Excellent problem-solving skills with the ability to work in dynamic, high-velocity environments. * Motivated, thorough, and committed to engineering excellence and ongoing improvement., Apache Spark, Artificial Intelligence (AI), Artificial Intelligence (AI) Agents, Best Practices, Business Intelligence, Business Model, Business Operations, Cloud Architecture, Cloud Computing, Computer Science, Customer Support/Service, Data Analysis, Data Management, Data Modeling, Data Processing, Data Quality, Data Sets, Data Warehousing, Database Design, Database Extract Transform and Load (ETL), Digital Video, Distributed Computing, Equal Employment Opportunity (EEO), Film, GCP (Good Clinical Practices), Large-Scale Systems, MCP - Microsoft Certified Professional, Metadata, Metrics, Power BI, Problem Solving Skills, Production Systems, Python Programming/Scripting Language, Query Optimization, Reporting Dashboards, SQL (Structured Query Language), Scalable System Development, Software Engineering, Sports Reporting, Streaming Technology, Television Advertising, Use Cases, Video On Demand ## Description Self-Service Analytics & Real-Time Data Platforms * Design, develop, and maintain scalable batch (ETL/ELT) and near real-time streaming data pipelines. These pipelines will process large-scale structured and unstructured datasets. * Design and maintain semantic layers, metrics frameworks, and curated data products. * Enable self-service analytics through governed and reusable business data models. * Implement monitoring, observability, and operational best practices. * Develop governed data access patterns for AI, conversational analytics, and MCP-based applications. * Build AI-ready data products that support machine learning, GenAI, AI agents, and chatbot applications. * Partner with Product, Analytics, BI, and Engineering stakeholders to deliver trusted data solutions. Data Modeling & Platform Architecture * Design scalable data models optimized for analytics, real-time reporting, and AI use cases. * Develop reusable semantic and transformation layers that provide consistent business definitions. * Drive best practices for data quality, governance, metadata, and discoverability. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Beyond Dashboards: Fixing Text-to-SQL with Semantic RAG](https://www.wearedevelopers.com/videos/2036-beyond-dashboards-fixing-text-to-sql-with-semantic-rag) - [A Data Mesh needs Open Metadata](https://www.wearedevelopers.com/videos/505-a-data-mesh-needs-open-metadata) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Data Analytics with Microsoft Fabric: End-to-End Use Case with Data Agents](https://www.wearedevelopers.com/videos/1547-data-analytics-with-microsoft-fabric-end-to-end-use-case-with-data-agents) - [Data: The Deciding Factor in AI Success](https://www.wearedevelopers.com/videos/100310-data-the-deciding-factor-in-ai-success) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)