> Markdown version of [/jobs/ext/2253396-staff-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2253396-staff-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer - **Company:** WEX Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $120,600.0 - $150,900.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Databases, Continuous Integration, Information Leak Prevention, Data Security, Fault Tolerance, PostgreSQL, Load Testing, MySQL, NoSQL, Reliability Engineering, Prometheus, Systems Architecture, Data Logging, Grafana, Multi-Cloud, Reliability of Systems, Event Driven Architecture, Containerization, Kubernetes, Storage Technologies, Operational Systems, Splunk, Dynatrace, Docker, Elk Stack, Microservices - **Published:** August 26, 2026 - **Apply:** https://wexinc.wd5.myworkdayjobs.com/WEXInc/job/San-Francisco-CA/Staff-Site-Reliability-Engineer_R22610 ## About the Role If you're a senior technical leader passionate about building reliable systems, leading through influence, and making a meaningful impact with AI-enabled operations, this is a fantastic opportunity for you., * 8+ years of experience with a focus on large-scale system reliability. * Expertise in system architecture, cloud platforms, and automation frameworks. * Deep knowledge of Kubernetes, service meshes, and distributed tracing. * Experience with monitoring and logging platforms (Grafana, ELK stack, Splunk, etc.). * Knowledge of containerization and orchestration (Docker, Kubernetes). * Experience designing high-availability, fault-tolerant architectures. * Strong understanding of database reliability engineering (MySQL, PostgreSQL, NoSQL), plus networking, databases, and storage architectures. * Excellent incident command and crisis management skills. * Hands-on experience building AI agents and skills/tools that integrate with operational systems (APIs, observability, ticketing, CI/CD). * Working knowledge of AI ecosystems and agent architectures, including orchestration, tool calling, context/memory, evaluation, and human-in-the-loop patterns. * Practical understanding of AI security and governance for production use, secure permissions, data leakage prevention, secrets handling, and guarded autonomous actions. * Demonstrated ability to reduce TOIL with AI by automating repetitive operational work and delivering measurable efficiency and reliability gains. Nice to Have * Experience with multi-region and multi-cloud deployments. * Deep expertise in scalable microservices and event-driven architectures. * Strong experience with advanced observability tools (OpenTelemetry, Jaeger, Prometheus). * Leadership in driving large-scale SRE transformations. * Experience designing and developing AI agents, skills, and copilots for SRE/platform engineering, including evaluation and safe rollout practices. * Familiarity with enterprise agent platforms, skill registries, and observability for AI/agent workflows. * Ability to influence engineering culture and process improvements, including adoption of AI-assisted operations under change control, safety, and audit requirements. ## Description We are looking for a highly motivated, high-potential Staff Site Reliability Engineer (SRE) to join our team as a technical leader and drive transformative impact across WEX's platform reliability and operational excellence., * Architect and oversee the implementation of mission-critical systems with a focus on availability, scalability, and operational excellence. * Define and enforce SRE best practices and operational standards across engineering and platform teams. * Lead cross-functional initiatives to enhance system reliability, performance, and efficiency at scale. * Serve as a technical advisor for engineering leadership on reliability, architecture, and operational risk. * Develop capacity planning and load testing strategies that proactively identify and mitigate scalability risks. * Design self-healing and auto-recovery mechanisms that reduce manual intervention during failures. * Drive cloud cost optimization and budgeting initiatives without compromising reliability. * Design, build, and govern AI agents and reusable skills that automate operational workflows and reduce TOIL. * Evaluate and integrate AI ecosystems, including models, agent frameworks, orchestration, tooling interfaces, and evaluation practices, into SRE and platform workflows. * Apply AI security and governance controls, including least-privilege tool access, secure data and prompt handling, auditability, and safe automation boundaries. * Lead AI-enabled initiatives for incident response, runbook automation, anomaly detection, and capacity/performance insights, with clear measurement of TOIL reduction and reliability outcomes. * Mentor engineers on production-grade agentic solutions and help embed AI into day-to-day reliability practices. ## Related Videos - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Coding for Good: Achieving social change with an app](https://www.wearedevelopers.com/videos/1645-coding-for-good-achieving-social-change-with-an-app) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)