Reliability Engineer

Informatic Technologies
Chicago, IL, United States
15 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Cloud Computing Cloud Engineering Distributed Systems High-Frequency Trading Infrastructure as a Service (IaaS) Identity and Access Management Python (Programming Language) Node.Js Reliability Engineering Software Engineering Web Application Frameworks Google Cloud
+5 more
Generative AI Kubernetes Low Latency Apache Kafka Terraform

Job description

We are seeking a Staff Site Reliability Engineer (Platform Engineering) to serve as the foundational Technical Lead for a premier global Financial Services enterprise. In this role, you will be the primary architect and visionary for core technology foundations that underpin high-volume, ultra-low latency financial marketplaces.

You will bridge the gap between high-level business strategy and deep technical implementation, ensuring our Google Cloud Platform-native stack provides mission-critical reliability, extreme scalability, and operational excellence. Your goal is to evolve the platform from “Infrastructure as a Service” to “Reliability as a Product.”, * Technical Vision & Strategy: Define and execute the 12-18 month technical roadmap for the Platform SRE ecosystem, building high-level Internal Development Platform (IDP) abstractions in Python.

  • Architectural Leadership: Serve as the final technical authority for core infrastructure architectures spanning Google Cloud Platform, GKE, and enterprise-grade Kafka messaging clusters.
  • Incident Command & Resilience: Lead response strategies for complex, cross-functional outages; foster a blameless engineering culture focused on code-driven, automated resiliency.
  • Reliability Governance: Standardize and enforce SLIs, SLOs, and Error Budgets across all engineering pods to safeguard system integrity.
  • GenAI & Intelligent Ops: Leverage Generative AI and Agentic workflows (e.g., Gemini) to build self-healing infrastructure and automated root-cause analysis frameworks.
  • Engineering Mentorship: Elevate the global SRE organization through architectural office hours, design reviews, and engineering best practices.

Requirements

  • Experience: 10+ years in SRE, Infrastructure, or Software Engineering roles in high-concurrency, high-availability environments.
  • Leadership: 3+ years in a Staff, Principal, or Tech Lead capacity overseeing complex platform engineering domains.
  • Cloud Native Mastery: Expertise in Google Cloud Platform (Networking, IAM, GKE) and scaling Kafka event buses for low-latency operations.
  • Software Engineering: Expert-level proficiency in Python (and ideally Go) for writing production-grade distributed systems and custom Kubernetes operators.
  • IaC & GitOps: Hands-on mastery of Terraform module design and GitOps patterns via ArgoCD.
  • Location: Chicago-based or willing to relocate to Chicago (hybrid schedule requiring 2 days/week on-site)., * Prior experience in Financial Markets, High-Frequency Trading (HFT), or heavily regulated financial ecosystems.
  • Google Cloud Platform Professional Cloud Architect or Certified Kubernetes Administrator (CKA/CKAD).
  • Full-Stack exposure (Node.js or modern web frameworks).

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · WWC 2023

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

1:44 min

Career transition into cloud native and data management

Michael Cade · LIVE

3:55 min

Identifying underlying Node.js runtime vulnerabilities using fuzzing tools

Sonya Moisset · WWC 2023

Videos

See all

Related articles

See all