> Markdown version of [/jobs/ext/1265093-software-engineer-data-optimization-in-san-mateo](https://www.wearedevelopers.com/jobs/ext/1265093-software-engineer-data-optimization-in-san-mateo). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer-Data Optimization in San Mateo - **Company:** Energy Jobline - **Location:** San Mateo, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Big Data, Cloud Computing, Encodings, Information Engineering, Extract Transform Load (ETL), Data Structures, Data Visualization, Distributed Computing Environment, Python (Programming Language), Machine Learning, Operational Databases, Performance Tuning, Search Technologies, Software Engineering, SQL Databases, Visual Analytics, Web Services, Data Processing, ReactJS, Apache Spark, Indexer, Backend, Fastapi, Pyspark, Kubernetes, Front End Software Development, Data Pipelines, Amazon Elastic Mapreduce (EMR), Docker, Databricks - **Published:** July 14, 2026 - **Apply:** https://www.energyjobline.com/job/software-engineer-data-optimization-san-mateo-31144552 ## About the Role * 3+ years of professional software engineering experience with a focus on data processing and pipeline engineering. * Experience designing, building, and optimizing distributed data processing pipelines at scale (e.g., ETL/ELT) using technologies like Spark, Databricks, AWS EMR, AWS Batch, or Ray Core/Data. * Strong proficiency in Python with experience building production data pipelines and web services (FastAPI, Uvicorn, or similar async frameworks). * Experience building and maintaining data visualization dashboards, with proficiency in SQL, and familiarity with PySpark/Scala for large-scale data manipulation. * Experience with workflow orchestration tools (Airflow, Prefect, Dagster, or similar) for managing complex multi-step data processing pipelines. * Familiarity with vector similarity search and indexing technologies (FAISS, Annoy, ScaNN, Milvus, or similar). Experience working with cloud infrastructure (AWS S3, EC2) and container orchestration (Kubernetes, Docker). Strong understanding of data structures, algorithms, and performance optimization for data-intensive workloads. * Familiarity with React.js or a comparable modern frontend framework for building interactive data-driven applications. * Excellent communication and collaboration skills; ability to work effectively with ML researchers, data engineers, and product stakeholders. ## Description In this role, you'll develop scalable data pipelines, backend services, and user-facing tools that support dataset creation, embedding search, and ML data optimization. While primarily focused on backend and data engineering, you'll also contribute to frontend features that improve workflow management and user experience. The ideal candidate is a strong Python engineer who thrives in data-intensive environments and enjoys building scalable systems that support machine learning and large-scale data processing. As a Software Engineer - Autonomy Behavior ML Data Optimization Team, you'll: * Design, build, and maintain Airflow pipelines that orchestrate end-to-end dataset creation, refresh, and management workflows. * Develop scalable Python/FastAPI services supporting embedding search, bulk operations, and dataset management. * Build and optimize distributed data processing pipelines using Ray for embedding , clustering, FAISS index creation, cache , and large-scale dataset processing. * Improve pipeline reliability, observability, and performance through monitoring, optimization, and operational tooling. * Develop dashboards and visualization tools that provide insight into datasets, pipeline health, and system performance. * Contribute to React-based administrative tools and workflow management features that improve dataset management and search capabilities. * Partner with ML researchers and cross-functional engineering teams to integrate new embedding technologies and deliver scalable software solutions. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)