> Markdown version of [/jobs/ext/211790-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/211790-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** FlightSafety International Inc. - **Location:** Seattle, WA, United States (Remote available) - **Experience:** Expert - **Salary:** $164,700.0 - $266,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon Web Services, Application Performance Management, Microsoft Azure, C Sharp (Programming Language), Software as a Service, Code Review, Databases, Continuous Integration, Software Debugging, Linux, DevOps, Distributed Systems, Domain Name System (DNS), Python (Programming Language), Networking Basics, Performance Tuning, Release Management, Reliability Engineering, Site Reliability Engineering Practices, Prometheus, Runbook, Software Engineering, Data Logging, Load Balancing, Grafana, Caching, Deployment Automation, Data Analytics, DocuSign, Golang - **Published:** May 23, 2026 - **Apply:** https://seeker-sp.worksourcewa.com/jobview/GetJob.aspx?JobID=293435494 ## About the Role We are looking for a self-motivated, driven and creative Senior Site Reliability Engineer to join the Site Reliability team. Metrics and analytics drive engineering at DocuSign and ensure that we are dedicating valuable engineering cycles to the right places. This role is a unique opportunity to impact the entire DocuSign team and drive adoption., * 8+ years of experience in Site Reliability Engineering, DevOps, or Software Engineering roles with ownership of production systems at scale (or equivalent experience) * Experience coding in at least one modern language (e.g., Go, Python, C#, Java), with the ability to design, implement, test, and debug productiongrade automation and services * Practical experience operating largescale services in public cloud (Azure preferred; AWS/GCP acceptable with willingness to learn Azure) * Experience with Linux, networking fundamentals, and common infrastructure components (load balancers, DNS, certificates, queues, caches, databases) * Experience with Observability stacks (e.g., Prometheus/Grafana, OpenTelemetry/Chronicle, centralized logging) * Experience with CI/CD systems and deployment strategies (blue/green, canary, rolling updates) * Experience with incident management and oncall operations for 24x7 services * Experience in building dashboards and metrics analysis Preferred * Strong analytical and problem-solving skills * Experience in highavailability, regulated, or customerfacing SaaS environments * Background in reliability practices such as chaos testing, capacity modeling, and performance tuning * Exposure to release management/unified release practices and safe rollout strategies (feature flags, staged rollouts, configurationdriven changes) * Demonstrated leadership driving crossteam initiatives: reliability programs, migrations, or major refactors * Strong written and verbal communication skills; ability to explain complex technical topics to both engineers and nontechnical stakeholders ## Description We are looking for a Senior Site Reliability Engineer (Senior SRE) to lead reliability initiatives for highimpact services. In this role, you will own the reliability, scalability, and performance of one or more critical systems, lead the design and implementation of automation to eliminate toil and reduce operational risk, drive improvements in observability, incident response, and production readiness across teams and partner closely with product engineering, platform, security, and release management to ship changes safely and quickly. Senior SREs at Docusign operate as handson technical leaders: they set the reliability bar for their domain, mentor other engineers, and lead crossfunctional projects that materially improve availability and customer experience. Ideally, you have a background in software development, incident management, service catalogs, request tracing systems, time series telemetry platforms, application performance management tools or log management tools. The role requires an on-call rotation every 4 weeks. This position is an individual contributor role reporting to the Senior Manager, SRE. Responsibility * Design, implement, and operate highly available, scalable services in cloud environments (primarily Azure, with some multicloud scenarios) * Define and evolve SLOs/SLIs, error budgets, and capacity strategies for owned services; use them to guide engineering tradeoffs and release decisions * Analyze patterns in incidents and outages; own longterm reliability improvements for your domain and contribute to reliability strategy across services * Write high quality code that is easy to maintain and test * Ensure design and architecture is extensible across projects, and participate in technical design and code reviews * Identify operational toil and lead automation efforts to eliminate it-deployment, runbook, and remediation workflows that make incidents rarer and faster to resolve * Develop robust, welltested tooling and shared libraries that are adopted across multiple teams * Improve CI/CD pipelines and guardrails to reduce change failure rate while increasing deployment velocity * Design and implement logging, metrics, tracing, and alerting for complex distributed systems; ensure signals are actionable and aligned to business impact * Build and automate tools and solutions for incident impact analysis and effective mitigation * Participate in and often lead incident response for Sev0-Sev2 events: triage, mitigation, coordination, and clear communication * Perform and contribute to blameless postincident reviews, rootcause analysis, and followthrough on corrective actions * Work with Operations and Incident Command teams during and post incidents to drive excellence in Incident Management Process * Compose and analyze dashboard to highlight areas of the business that need attention and help drive organizational KPI * Create and respond to system generated alerts to maintain system health * Work with Operations and Engineers to fill any gaps in alerting and telemetry * Act as the primary SRE partner for one or more engineering teams-shaping architecture, reviewing designs, and embedding reliability best practices * Mentor and coach other SREs and software engineers on topics such as debugging, observability, incident management, and performance optimization * Contribute to and help standardize SRE practices, runbooks, and production readiness criteria across CPE and product teams * Work with Product Management, collaborators and other developers to understand design requirements and provide estimates for development * Learn and grow in all key technologies in Docusign and be a partner to Eng and Operations teams Job Designation Remote: Employee is not required to be in or near an office frequently and works from a designated remote work location for the majority of the time. Positions at Docusign are assigned a job designation of either In Office, Hybrid or Remote and are specific to the role/job. Preferred job designations are not guaranteed when changing positions within Docusign. Docusign reserves the right to change a position's job designation depending on business needs and as permitted by local law. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Now is the time for industrialized software development](https://www.wearedevelopers.com/magazine/601-now-is-the-time-for-industrialized-software-development) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026) - [Events like RSAC Get You CISOs. Developers Decide What Actually Gets Deployed.](https://www.wearedevelopers.com/magazine/693-events-like-rsac-get-you-cisos-developers-decide-what-actually-gets-deployed) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)