> Markdown version of [/jobs/ext/2152904-tech-lead-data-inference-engineer](https://www.wearedevelopers.com/jobs/ext/2152904-tech-lead-data-inference-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Tech Lead, Data & Inference Engineer - **Company:** Catalyst - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon Web Services, Application Integration Architecture, Microsoft Azure, Big Data, Cloud Computing, Configuration Management, Code Review, Computer Engineering, Continuous Delivery, Continuous Integration, Data Architecture, Data Systems, Data Warehousing, Distributed Data Store, Python (Programming Language), Machine Learning, Language Modeling, Systems Development Life Cycle, Query Optimization, SQL Databases, Systems Integration, Large Language Models, Apache Spark, Caching, Data Lakes, Kubernetes, Information Technology, Low Latency, Apache Flink, Apache Kafka, Software Version Control - **Published:** August 20, 2026 - **Apply:** https://www.careerbuilder.com/job-details/tech-lead-data-inference-engineer-nc--6e1e1218-0d1f-470d-805f-4c6fb2a5f00e ## About the Role * Bachelors or Masters degree in Computer Science, Computer Engineering, Electrical Engineering, or Mathematics. * Excellent written and verbal communication; proactive and collaborative mindset. * Comfortable in hybrid or distributed environments with strong ownership and accountability. * A founder-level bias for actionable to identify bottlenecks, automate workflows, and iterate rapidly based on measurable outcomes. * Demonstrated ability to teach, mentor, and document technical decisions and schemas clearly. Core Experience * 6 to 12 years of experience building and scaling production-grade data systems, with deep expertise in data architecture, modeling, and pipeline design. * Expert SQL (query optimization on large datasets) and Python skills. * Hands-on experience with distributed data technologies (Spark, Flink, Kafka) and modern orchestration tools (Airflow, Dagster, Prefect). * Familiarity with dbt, DuckDB, and the modern data stack; experience with IaC, CI/CD, and observability. * Exposure to Kubernetes and cloud infrastructure (AWS, GCP, or Azure). * Bonus: Strong Node.js skills for faster onboarding and system integration. * Previous experience at a high-growth startup (10 to 200 people) or early-stage environment with a strong product mindset. Skills: Advertising, Apache Kafka, Application Integration, Application Programming Interface (API), Artificial Intelligence (AI), Business-to-Business (B2B) Marketing, Caching, Code Reviews, Computer Engineering, Computer Science, Continuous Deployment/Delivery, Continuous Integration, Contract Creation, Cost Control, Customer Relations, Data Modeling, Data Quality, Data Science, Demand Generation, Electrical Engineering, Embedded Systems, Funding, Machine Learning, Mathematics, Mentoring, Modeling Languages, Presentation/Verbal Skills, Problem Solving Skills, Process Modeling, Quality Assurance, Root Cause Analysis, Scalable System Development, Source Code/Configuration Management (SCM), Startup, Team Player, Technical Leadership, Technical Writing, Training/Teaching, Use Cases, Writing Skills, eCommerce ## Description A fast moving and venture backed advertising technology startup based in San Francisco. They have raised twelve million dollars in funding and are transforming how business to business marketers reach their ideal customers. Their identity resolution technology blends business and consumer signals to convert static audience lists into high match and cross channel segments without the use of cookies. By transforming first party and third party data into precision targetable audiences across platforms such as Meta, Google and YouTube, they enable marketing teams to reach higher match rates, reduce wasted advertising spend and accelerate pipeline growth. With a strong understanding of how business buyers behave in channels that have traditionally been focused on business to consumer activity, they are redefining how business brands scale demand generation and account based efforts., * Lead the design, development and scaling of an end to end data platform from ingestion to insights, ensuring that data is fast, reliable and ready for business use. * Build and maintain scalable batch and streaming pipelines, transforming diverse data sources and third party application programming interfaces into trusted and low latency systems. * Take full ownership of reliability, cost and service level objectives. This includes achieving ninety nine point nine percent uptime, maintaining minutes level latency and optimizing cost per terabyte. Conduct root cause analysis and provide long lasting solutions. * Operate inference pipelines that enhance and enrich data. This includes enrichment, scoring and quality assurance using large language models and retrieval augmented generation. Manage version control, caching and evaluation loops. * Work across teams to deliver data as a product through the creation of clear data contracts, ownership models, lifecycle processes and usage based decision making. * Guide architectural decisions across the data lake and the entire pipeline stack. Document lineage, trade offs and reversibility while making practical decisions on whether to build internally or buy externally. * Scale integration with application programming interfaces and internal services while ensuring data consistency, high data quality and support for both real time and batch oriented use cases. * Mentor engineers, review code and raise the overall technical standard across teams. Promote data driven best practices throughout the organization. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [Event based cache invalidation in GraphQL](https://www.wearedevelopers.com/videos/433-event-based-cache-invalidation-in-graphql) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)