Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+5 more
Job description
We are seeking a Staff Site Reliability Engineer (Platform Engineering) to serve as the foundational Technical Lead for a premier global Financial Services enterprise. In this role, you will be the primary architect and visionary for core technology foundations that underpin high-volume, ultra-low latency financial marketplaces.
You will bridge the gap between high-level business strategy and deep technical implementation, ensuring our Google Cloud Platform-native stack provides mission-critical reliability, extreme scalability, and operational excellence. Your goal is to evolve the platform from “Infrastructure as a Service” to “Reliability as a Product.”, * Technical Vision & Strategy: Define and execute the 12-18 month technical roadmap for the Platform SRE ecosystem, building high-level Internal Development Platform (IDP) abstractions in Python.
- Architectural Leadership: Serve as the final technical authority for core infrastructure architectures spanning Google Cloud Platform, GKE, and enterprise-grade Kafka messaging clusters.
- Incident Command & Resilience: Lead response strategies for complex, cross-functional outages; foster a blameless engineering culture focused on code-driven, automated resiliency.
- Reliability Governance: Standardize and enforce SLIs, SLOs, and Error Budgets across all engineering pods to safeguard system integrity.
- GenAI & Intelligent Ops: Leverage Generative AI and Agentic workflows (e.g., Gemini) to build self-healing infrastructure and automated root-cause analysis frameworks.
- Engineering Mentorship: Elevate the global SRE organization through architectural office hours, design reviews, and engineering best practices.
Requirements
- Experience: 10+ years in SRE, Infrastructure, or Software Engineering roles in high-concurrency, high-availability environments.
- Leadership: 3+ years in a Staff, Principal, or Tech Lead capacity overseeing complex platform engineering domains.
- Cloud Native Mastery: Expertise in Google Cloud Platform (Networking, IAM, GKE) and scaling Kafka event buses for low-latency operations.
- Software Engineering: Expert-level proficiency in Python (and ideally Go) for writing production-grade distributed systems and custom Kubernetes operators.
- IaC & GitOps: Hands-on mastery of Terraform module design and GitOps patterns via ArgoCD.
- Location: Chicago-based or willing to relocate to Chicago (hybrid schedule requiring 2 days/week on-site)., * Prior experience in Financial Markets, High-Frequency Trading (HFT), or heavily regulated financial ecosystems.
- Google Cloud Platform Professional Cloud Architect or Certified Kubernetes Administrator (CKA/CKAD).
- Full-Stack exposure (Node.js or modern web frameworks).
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
7 Cloud Computing Trends Coming in 2025 for Developers
The Best X (Twitter) Accounts for Developers
Highest Paying Tech Companies for Developers