> Markdown version of [/jobs/ext/2154759-sde-iv-data-engineer](https://www.wearedevelopers.com/jobs/ext/2154759-sde-iv-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SDE IV - Data Engineer - **Company:** InMobi - **Location:** United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Training Data, Java (Programming Language), Application Programming Interfaces (APIs), Airflow, Apache HTTP Server, Big Data, Code Review, Databases, Information Engineering, Data Governance, Data Infrastructure, Data Retention, Data Stores, Github, Python (Programming Language), Machine Learning, Meta-Data Management, Pair Programming, Query Optimization, Software Engineering, SQL Databases, Data Streaming, Workflow Management Systems, Data Logging, Aerospike, Data Processing, Feature Engineering, Snowflake, Apache Spark, Backend, Data Lakes, Kubernetes, Information Technology, Apache Flink, Cassandra, Apache Kafka, Vertica, Api Design, Restful APIs, Stream Processing, Data Pipelines, Dynatrace, Microservices - **Published:** August 20, 2026 - **Apply:** https://job-boards.greenhouse.io/inmobi/jobs/6919206 ## About the Role + 8+ years of software engineering experience, with at least 4 years focused on data engineering or large-scale data systems. + Prior experience in ad-tech, programmatic advertising, or real-time bidding (RTB) environments strongly preferred. + Experience in ML engineering, feature engineering, building offline training pipelines, or collaborating closely on model productionisation - is preferred. + Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field. * Data Engineering (Core) + Deep expertise in distributed stream processing frameworks - Apache Kafka, Apache Flink, or Apache Spark Structured Streaming. + Strong command of SQL and experience optimizing complex analytical queries at petabyte scale. + Hands-on experience with real-time analytical databases, particularly StarRocks; familiarity with Snowflake or ClickHouse is a plus. + Proficiency with workflow orchestration tools: Apache Airflow, Prefect, or Dagster. + Solid understanding of data modelling - dimensional modelling, OBT patterns, and eventdriven schemas relevant to ad impression and attribution data. + Familiarity with open table formats: Apache Iceberg, Delta Lake, or Apache Hudi. * Backend & Systems: + Proficiency in Python (primary) and/or Java/Scala for building production-grade services and data processing jobs. + Experience designing and building RESTful APIs and/or gRPC-based microservices. + Working knowledge of low-latency data stores (Aerospike, Cassandra) used for real-time lookups in the bidding path. + Comfortable with containerised deployments on Kubernetes and CI/CD pipelines (GitHub Actions, ArgoCD, or similar). * Leadership & Collaboration: + Demonstrated experience leading technical squads or functioning as a technical anchor on cross-functional projects. + Strong written and verbal communication skills; ability to translate technical complexity for nonengineering stakeholders. + Track record of driving projects from ambiguous requirements to production with high engineering quality ## Description * Data Stack + Design and own end-to-end data pipelines - batch and real-time - handling bid requests, win/loss events, impression logs, click streams, and conversion signals at scale. + Build and maintain our data lakehouse architecture (ingestion * storage * serving layer) optimized for low-latency DSP analytics and ML feature generation. + Define standards for data quality, lineage, observability, and SLA adherence across all data products. + Own and evolve our StarRocks deployment for real-time analytical queries - schema design, ingestion patterns, and query optimisation for DSP reporting workloads. + Drive FinOps discipline across the data platform - track compute and storage costs, enforce resource quotas, identify optimisation opportunities, and report on unit economics. + Establish and lead Data Governance practices: metadata management, data cataloguing, access control policies, PII classification, and data retention standards compliant with GDPR and CCPA. * Backend Microservices + Contribute to and review the design of backend services that power DSP functionality - bid shading, pacing, budget management, and reporting APIs. + Architect data contracts and event schemas between microservices, ensuring consistency across the bid request lifecycle. + Champion best practices in API design, service reliability (SLOs, circuit breakers, retries), and observability (distributed tracing, structured logging, metrics). + Collaborate with the bidder team on low-latency data reads (Aerospike, Cassandra, or similar) for real-time targeting signal lookups * Technical Leadership + Lead architecture reviews, set engineering standards, and drive adoption of best practices across the data and backend engineering teams. + Mentor SDE II/III engineers through code reviews, design sessions, and pair programming; grow the next generation of technical leaders. + Partner with ML engineers on the full model lifecycle - feature engineering pipelines, offline training data generation, experiment tracking, and model serving infrastructure. + Translate complex business requirements (campaign KPIs, ROAS, viewability, brand safety) into robust technical designs. + Own the technical roadmap for data infrastructure - capacity planning, migration strategies, and build vs. buy decisions. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries)