> Markdown version of [/jobs/ext/2039346-site-reliability-architect](https://www.wearedevelopers.com/jobs/ext/2039346-site-reliability-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Architect - **Company:** The Summit - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $136,000.0 - $175,000.0 - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), Amazon Web Services, Audit Trail, Microsoft Azure, Software as a Service, Python (Programming Language), Reliability Engineering, Cloud Services, Ansible, Prometheus, Ruby, Scripting, Grafana, Multi-Cloud, Influxdb, Graphite, Terraform, Golang, Programming Languages - **Published:** August 12, 2026 - **Apply:** https://ats.rippling.com/summithq/jobs/2d0d9398-3f69-4cdf-b379-ff5e4c2b3184 ## About the Role * Have a proven track record in designing and implementing observability and reliability platforms at scale. * Are comfortable influencing architecture and strategy across engineering, operations, and product teams. * Enjoy mentoring engineers and shaping organizational practices, not just managing tools. * Think strategically but aren't afraid to get hands-on when solving complex problems. * Thrive in environments where you can innovate, set standards, and bring structure to complex ecosystems. Bonus Points: * Expertise with observability stacks such as ELK, Grafana/Graphite/InfluxDB, LogicMonitor, Prometheus, or related platforms. * Advanced automation experience using Ansible, Terraform, or equivalent tooling. * Experience designing reliability frameworks in multi-cloud environments (Azure, AWS, or hybrid). * Knowledge of scripting and development languages (Python, Go, Ruby, or Javascript). * Familiarity with compliance-heavy industries where reliability, security, and auditability are paramount. ## Description Lead the vision, strategy, and architecture for observability and reliability platforms. Define SLOs/SLIs/SLAs, implement monitoring/alerting automation, lead incident response and post-incident analysis, evaluate tools, and mentor engineering teams to improve resilience, scalability, and reliability across multi-cloud environments. The summary above was generated by AI At Summit, we're on the lookout for talent that doesn't just think "outside the box," but brings their own unique perspective to the table. With our relentless pursuit of excellence and curiosity, we lead innovation in our industry. We humanize technology by actively listening to our clients, crafting tailored proposals, and delivering on the promise of technology with precision and purpose. Summit is a leading provider of enterprise-class Application Hosting, Managed Services, and Cloud Solutions for regulated industries, with deep experience supporting compliance, security, and performance in complex IT environments. Our mission is to simplify the complex - ensuring our clients' technology environments are secure, performant, and purpose-built for their most critical applications. As our Site Reliability Architect, you will own the vision, strategy, and architecture of Summit's observability and reliability platforms. You will ensure that the systems supporting our most critical applications are designed with resilience, scalability, and efficiency in mind. Beyond managing platforms, you will shape the standards and practices that make Summit's infrastructure visible, measurable, and reliable. You will collaborate across engineering, operations, and product teams to embed observability into the fabric of our technology, mentor engineers on best practices, and influence architecture decisions that improve performance and reliability. The right candidate will thrive in a fast-moving environment, combining deep technical expertise with the ability to see the bigger picture and guide the evolution of Summit's reliability ecosystem. What You'll Do: * Define the observability and reliability architecture strategy across Summit's platforms and services. * Implement site reliability concepts based around SLOs, SLIs, and SLAs, across teams and platforms * Partner with engineering and operations leadership to ensure system design aligns with resilience and scalability goals. * Lead the design, implementation, and governance of observability frameworks and standards across teams. * Oversee the automation of monitoring, alerting, and incident response processes to maximize efficiency and reduce human error. * Serve as the escalation point and lead for complex, cross-platform incidents, guiding resolution and post-incident analysis. * Evaluate and introduce emerging tools, frameworks, and practices to strengthen Summit's reliability and observability posture. * Mentor and coach engineering teams, fostering a culture of reliability, automation, and continuous improvement. What You'll Deliver: * A scalable, standards-driven observability framework embedded across Summit's technology stack. * Clear alignment of monitoring, alerting, and automation with business-critical definitions of health. * Reduced time-to-detect and time-to-resolve for incidents through automation and well-designed response frameworks. * Improved reliability practices across teams through coaching, knowledge-sharing, and collaboration. * Strategic roadmaps for observability and monitoring platforms that support Summit's growth and evolving client needs. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Coroutine explained yet again 60 years later](https://www.wearedevelopers.com/videos/690-coroutine-explained-yet-again-60-years-later) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)