> Markdown version of [/jobs/ext/3309694-senior-data-ops-engineer-data-activation-products-activision](https://www.wearedevelopers.com/jobs/ext/3309694-senior-data-ops-engineer-data-activation-products-activision). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Ops Engineer, Data Activation & Products - Activision - **Company:** Activision Publishing, Inc. - **Location:** Santa Monica, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $102,800.0 - $190,204.0 - **Contract:** Temporary contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Bash Shell, Cloud Computing, Continuous Integration, Directed Acyclic Graph (Directed Graphs), Data Validation, Information Engineering, Data Infrastructure, Linux, DevOps, Programming Tools, Domain Name System (DNS), Github, Monitoring of Systems, Hypertext Transfer Protocols (HTTP), Python (Programming Language), Operational Databases, Reliability Engineering, Prometheus, Software Engineering, Data Streaming, Datadog, Scripting, Cloud Platform System, Cloud Monitoring, Delivery Pipeline, Grafana, Apache Spark, Backend, Event Driven Architecture, Gitlab-ci, Git Flow, Kubernetes, Low Latency, Apache Kafka, Spark Streaming, Data Management, Terraform, Splunk, Jenkins, Databricks - **Published:** September 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=9b45e2a0c2dfe213 ## About the Role * 5+ years of experience in SRE, DevOps, cloud infrastructure, platform engineering, software engineering, data platform operations, or related production-support roles. * Hands-on experience supporting Kubernetes-based workloads, deployment systems, cloud infrastructure, or production application environments. * Familiarity with Linux, HTTP, DNS, containers, Kubernetes, Git-based workflows, and scripting in Bash, Python, or similar languages. * Experience with monitoring, logs, metrics, dashboards, alerting, and incident management practices. * Comfort working with event systems such as Kafka, Google Pub/Sub, Kinesis, or similar technologies, including topics, subscriptions, consumers, retries, lag, and dead-letter queues. * Strong troubleshooting mindset, clear communication, and comfort operating in a production-support environment. * Interest in data platforms, data engineering systems, orchestration, streaming workloads, event-driven architecture, internal developer tools, and production data services. Nice to Have * Experience with Kubernetes deployment and release tooling such as Helm, ArgoCD, or similar GitOps workflows. * Experience with CI/CD automation using GitHub Actions, GitLab CI, Jenkins, or similar pipelines. * Familiarity with infrastructure as code using Terraform or similar provisioning tools. * Familiarity with Databricks, Spark, Spark Structured Streaming, Airflow, Astronomer, dbt, Kafka, Pub/Sub, object storage, or lakehouse architectures. * Experience with observability tools such as Grafana, Prometheus, Cloud Monitoring, Datadog, Splunk, or similar platforms. * Experience supporting internal applications, APIs, event-driven services, streaming consumers, or backend workers. * Exposure to data reliability concepts such as freshness, latency, completeness, pipeline health, data quality checks, and dependency-aware alerting. * Exposure to SLOs, SLAs, error budgets, postmortems, or formal reliability practices. ## Description We are looking for a Site Reliability Engineer to help improve the reliability, observability, and operational maturity of our data platforms, Kubernetes-based deployment systems, internal applications, and cloud environments. This role sits at the intersection of SRE, DevOps, and data. The ideal candidate is comfortable operating production systems, troubleshooting across infrastructure and applications, and helping teams deploy and support services more safely. The role does not require someone to be a data engineer, but they should be excited about the systems that support modern data engineering, including Databricks, Spark, Airflow/Astronomer, streaming pipelines, event systems, internal tools, and backend services. We are especially interested in a creative, curious engineer who enjoys learning new technology and using AI-accelerated development practices to solve problems faster and more thoughtfully. You should be excited to experiment with agentic development tools, automation frameworks, and emerging platform capabilities, while applying sound engineering judgment. This role has the option to be based in our Los Angeles (Pen Factory) office and follows an onsite work schedule of Monday through Thursday or be remote. Work arrangements may change at the company's discretion to meet business needs. Priorities can often change in a fast-paced environment like ours, so this role includes, but is not limited to, the following responsibilities: * Monitor service health, respond to alerts, and participate in incident response for cloud, Kubernetes, application, and data platform environments. * Investigate reliability issues across Kubernetes, networking, DNS, application runtime behavior, Databricks jobs, Spark workloads, orchestration systems, event systems, and dependent services. * Support the reliability of internal applications, APIs, workers, streaming consumers, event-driven services, and deployment workflows used by data engineering and business teams. * Build and maintain dashboards, alerting, runbooks, and operational documentation that improve detection and recovery speed. * Improve observability for Databricks environments, including job health, Spark streaming workloads, structured streaming metrics, cluster behavior, failures, latency, throughput, and cost signals. * Help route Spark streaming metrics, operational logs, event-system signals, and platform health signals into monitoring tools such as Grafana. * Contribute to alerting patterns for Databricks workflows, Airflow/Astronomer DAGs, dbt jobs, data freshness, pipeline failures, event lag, dead-letter queues, and production data dependencies. * Contribute scripts and automation that reduce repetitive operational work and improve environment hygiene. * Support release and deployment reliability by validating changes, improving rollback readiness, and strengthening change safety. * Partner with data engineers, analytics engineers, and software engineers to improve reliability across pipelines, services, internal tools, event systems, and data products. * Participate in post-incident follow-up and help close corrective actions that prevent recurrence. * Support platform modernization and migration efforts, including orchestration platform changes, deployment system improvements, and shared reliability standards., * Kubernetes and deployment workflows are easier to operate, monitor, and troubleshoot. * Databricks, Spark, streaming, orchestration, event-driven, and dbt environments have clearer dashboards, alerts, and runbooks. * Spark streaming metrics, event-system health signals, and platform logs are easier to access, visualize, and operationalize through tools such as Grafana. * Incidents are detected faster, resolved more efficiently, and followed up with meaningful corrective actions. * Data engineers, analytics engineers, and software engineers can deploy changes more safely and with greater confidence. * Repetitive operational work is reduced through automation and better platform standards. * Platform migrations and modernization efforts are supported with strong reliability, observability, and operational practices. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Shifting Stress to Progress— Understanding DevOps to do DevOps Better](https://www.wearedevelopers.com/videos/268-shifting-stress-to-progress-understanding-devops-to-do-devops-better) - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [React Developer Salary [2023]](https://www.wearedevelopers.com/magazine/198-react-developer-salary-2023) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)