> Markdown version of [/jobs/ext/363053-staff-sre-ads](https://www.wearedevelopers.com/jobs/ext/363053-staff-sre-ads). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff SRE, Ads - **Company:** Reddit - **Location:** Amsterdam, Netherlands (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Big Data, BigQuery, Cloud Computing, Cloud Engineering, Computer Programming, Linux, Distributed Systems, Monitoring of Systems, Python (Programming Language), Machine Learning, Performance Tuning, Recommender Systems, Reliability Engineering, Data Logging, Apache Spark, Build Management, Kubernetes, Apache Flink, Apache Kafka, Vertica - **Published:** June 21, 2026 - **Apply:** https://nl.indeed.com/viewjob?jk=671c3c1156b33ea4 ## About the Role Do you have experience in Spark?, * 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large scale distributed systems. * Strong experience supporting high traffic, user facing production environments. * Deep understanding of distributed systems, networking, Linux systems, cloud native architectures. * Experience designing highly available systems with strong operational and reliability practices. * Strong understanding of observability systems including metrics, logging, tracing, and alerting. * Good programming skills in languages such as Go, Python, or similar. * Experience improving reliability through SLOs, automation, incident management, and performance optimization. * Demonstrated ability to troubleshoot complex issues across a modern distributed system stack. * Strong collaboration and communication skills with the ability to influence technical direction across teams., * Experience supporting advertising technology platforms or other large-scale revenue-critical systems. * Deep understanding of reliability challenges associated with ad-serving, real-time auctions, budget pacing, campaign delivery, measurement, attribution, or billing systems. * Experience operating high-QPS, low-latency services where latency directly impacts business outcomes. * Experience establishing reliability programs that deliver meaningful, measurable business outcomes * Experience with Kubernetes, cloud infrastructure, and large-scale distributed systems. * Familiarity with Kafka, ClickHouse, Spark, Flink, BigQuery, or similar large-scale data platforms. * Experience partnering with Product, Data Science, and Ads Engineering organizations. * Experience supporting machine learning inference or recommendation systems at scale. ## Description * Lead reliability initiatives across multiple Ads domains including ad serving, auctions, targeting, reporting, measurement, and billing. * Partner with engineering leadership to improve reliability, scalability, operational excellence, and engineering efficiency across the Ads organization. * Drive architecture reviews and influence technical decisions impacting critical revenue-generating systems. * Design and build platforms, tooling, and automation that improve reliability and developer productivity at scale. * Participate in on-call rotations, lead complex incident investigations and coordinate cross-functional response efforts during major production events. * Identify systemic reliability risks and drive long-term solutions that improve platform resilience. * Establish reliability metrics around advertiser-critical user journeys such as campaign creation, ad delivery, auction participation, reporting, attribution, and billing. * Mentor engineers and provide technical leadership across multiple teams. * Influence roadmap planning and ensure reliability considerations are incorporated into product and infrastructure investments. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)