> Markdown version of [/jobs/ext/2680453-technology-data-innovation](https://www.wearedevelopers.com/jobs/ext/2680453-technology-data-innovation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Technology Data & Innovation - **Company:** DEUTSCHE BANK A.G. - **Location:** Cary, NC, United States (Remote available) - **Salary:** $100,000.0 - $153,000.0 - **Contract:** Permanent contract - **Skills:** Microsoft Excel, Artificial Intelligence, Build Automation, Bash Shell, Computer Networks, DevOps, Distributed Systems, Python (Programming Language), PostgreSQL, Linux System Administration, Load Testing, MongoDB, Reliability Engineering, Ansible, Prometheus, Virtualization Technology, Data Logging, Istio, Grafana, Containerization, Kubernetes, Low Latency, Bare Metal, Apache Kafka, Splunk, Golang - **Published:** September 2, 2026 - **Apply:** https://careers.db.com/professionals/search-roles/#/professional/job/74436 ## About the Role * Proven experience in Site Reliability Engineering, Production Engineering, DevOps, or a closely related infrastructure role * Strong Linux system administration capability and infrastructure-level scripting experience using Python, Ansible, and Bash * Practical knowledge of observability stacks and telemetry pipelines, including Prometheus, Grafana, Splunk, metrics, logging, alerting, dashboards, and Open Telemetry-style concepts * Strong understanding of incident management, root cause analysis, operational readiness, computer networking, virtualization, containerization, and distributed systems behavior under failure * Proven ability to leverage AI tools to enhance productivity, optimize workflows to solve business problems, while applying critical judgment to ensure responsible and ethical use of data and AI outputs Skills That Will Help You Excel * Experience with Istio / Envoy, service mesh observability, traffic management, OPA Gatekeeper, admission controls, or policy-driven operational guardrails * Familiarity supporting stateful services such as PostgreSQL, Kafka, MongoDB, or comparable platform dependencies * Practical knowledge of capacity planning, load testing, chaos testing, failure-injection techniques, alert tuning, and self-healing automation * Exposure to low-latency or regulated environments with strict uptime, change control, compliance, time synchronization, deterministic performance, or SR-IOV workload constraints * Ability to read and understand Golang code when troubleshooting platform components, with strong written and verbal communication skills and a continuous learning mindset ## Description As a Site Reliability Engineer on the CaaS Private platform team, you will help operate and improve an on-prem, multi-tenant Kubernetes platform running on bare metal. You will strengthen the reliability, observability, scalability, and operational excellence of a platform that supports critical, low-latency, and regulated workloads. You will partner closely with platform, network, security, and application teams to define service level objectives, improve resilience, reduce operational toil, and build automation that allows the platform to run safely at scale. Join us here, and you will turn operational challenges into measurable engineering improvements that application teams can rely on every day., * Define, implement, and continuously improve SLI/SLOs, alerting standards, and error budgets for the CaaS Private platform and critical services * Build and maintain observability across metrics, logs, alerts, and dashboards to provide clear insight into platform health, saturation, latency, and failure modes * Lead or coordinate incident response for platform-impacting events, ensuring timely mitigation, clear communication, blameless postmortems, and durable follow-up actions * Automate repetitive operational tasks and remediation workflows to reduce toil, improve platform consistency, and accelerate recovery time * Improve reliability, upgrade safety, and operational readiness for Kubernetes clusters, ingress paths, service mesh components, node services, and critical platform dependencies * Partner with platform, network, security, and application teams on capacity planning, release readiness, troubleshooting, operational documentation, and adoption of best practices, * Hands-on Kubernetes expertise, including operating clusters on bare metal or private cloud environments and supporting platform services at scale ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [The Biggest German Tech Companies](https://www.wearedevelopers.com/magazine/424-the-biggest-german-tech-companies)