Principal Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+5 more
Job description
-
Enterprise-Wide Ownership: Define the standards for end-to-end service ownership, holding the organization accountable for availability, performance, and overall operational health.
-
Architectural Influence: Lead cross-departmental initiatives to influence system design at the architectural level, driving fault tolerance, strict compliance, and operational sustainability across public and private clouds.
-
Advanced Observability Vision: Dictate the enterprise strategy for observability frameworks, ensuring the Four Golden Signals (Latency, Traffic, Errors, and Saturation) provide actionable, predictive insights across all platforms.
Strategic Leadership & Reliability Architecture
-
Enterprise-Wide Ownership: Define the standards for end-to-end service ownership, holding the organization accountable for availability, performance, and overall operational health.
-
Architectural Influence: Lead cross-departmental initiatives to influence system design at the architectural level, driving fault tolerance, strict compliance, and operational sustainability across public and private clouds.
-
Advanced Observability Vision: Dictate the enterprise strategy for observability frameworks, ensuring the Four Golden Signals (Latency, Traffic, Errors, and Saturation) provide actionable, predictive insights across all platforms.
Data-Driven Reliability Governance
-
SLO & Error Budget Authority: Establish the governance models for defining and managing SLIs and SLOs across multiple product lines.
-
Delivery Alignment: Champion Error Budgets as the ultimate technical arbiter at the executive level, balancing feature velocity with the absolute requirement for platform stability.
Incident Management & Cultural Transformation
-
Enterprise Incident Command: Lead incident response for the most critical, high-severity events.
-
Blameless Culture Champion: Foster a âWin Togetherâ environment by championing a Blameless Postmortem culture globally, ensuring root cause analyses focus strictly on systemic and process improvements rather than individual error.
Requirements
-
Experience: 12 to 15+ years of hands-on Cloud Operations, SRE, or reliability-focused engineering experience, with a proven track record of end-to-end enterprise service ownership.
-
Proven Scope: Demonstrated ability to operate at a Principal/Architect scope, driving large-scale reliability outcomes and operational excellence across global organizations.
-
Software Engineering: Expert-level software engineering skills in C#, .NET, Java, Python, or React.
-
Principles: Deep expertise in scaling core SRE principles (SLIs, SLOs, error budgets) across complex, distributed systems.
-
Technical Stack: Mastery of AWS, Kubernetes, CI/CD pipelines, Infrastructure-as-Code, and extensive knowledge of Linux and Windows environments and relational databases.
-
Education: Bachelorâs or Masterâs degree in Computer Science or a related technical field.
-
Commitment: Participation in an executive on-call rotation with flexible hours as required
Skills & Requirements Knowledge, Skills, and Abilities: A fast learner. A problem solver. Ability to document procedures. Able to meet deadlines. Good communication skills. Able to deliver the message effectively to a technical and non-technical audience. Able to comply with processes and procedures. Able to maintain professional composure in any situations. Flexible in working extended hours on occasions or as required. Exposure in the insurance industry is desired but not mandatory. Driven to improve, personally and professionally Operate best in a fast-paced, flexible work environment with ability to work in a team.
Additional Requirements and Details: High speed internet to accommodate working from home needs. Occasional lifting and/or moving up to 10 pounds. Frequent repetitive hand and arm movements required to operate a computer. Specific vision abilities required by this job include close vision (working on a computer, etc.). Frequent sitting and/or standing., The selected candidate must be legally authorized to work in the United States.
Benefits & conditions
- The opportunity to work in a space where modern technology meets a stable and vital industry
- Medical, vision & dental plans
- Life, AD&D
- Short Term and Long Term Disability
- Pension Plan & Employer Match
- Maternity, Paternity and Parental Leave
- Employee and Family Assistance Program (EFAP)
- Education Assistance
- Additional programs - Employee Referral and Internal Recognition
Why Vertafore is the place for you: *US Only
- The opportunity to work in a space where modern technology meets a stable and vital industry
- We have a Flexible First work environment! Our North America team members use our offices for collaboration, community and team-building, with members asked to sometimes come into an office and/or travel depending on job responsibilities. Other times, our teams work from home or a similar environment.
- Medical, vision & dental plans
- PPO & high-deductible options
- Health Savings Account & Flexible Spending Accounts Options:
- Health Care FSA
- Dental & Vision FSA
- Dependent Care FSA
- Commuter FSA
- Life, AD&D (Basic & Supplemental), and Disability
- 401(k) Retirement Savings Plain & Employer Match
- Supplemental Plans - Pet insurance, Hospital Indemnity, and Accident Insurance
- Parental Leave & Adoption Assistance
- Employee Assistance Program (EAP)
- Education & Legal Assistance
- Additional programs - Tuition Reimbursement, Employee Referral, Internal Recognition, and Wellness
- Commuter Benefits (Denver)
About the company
The insurance industry runs on Vertafore. We equip agencies, MGAs, and carriers with the core digital systems, specialized AI, and data-driven foundation to eliminate distribution drag across the insurance lifecycle, spanning sales, servicing, and back-office operations.
Underpinned by unmatched speed and performance power, we are the trusted backbone thatâs taking the insurance industry from friction to flow with Distribution Velocity - speed, performance, and trust - to drive growth at scale.
With over 95% of the top agencies and insurers and 50% of industry compliance transactions running through Vertafore, we lead at the intersection of innovation and trust, giving insurance professionals the confidence to transform and win in the AI era.
Our reach is global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India.
We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise-wide reliability, scalability, and performance of our critical production services. As a foundational pillar of our engineering organization, this role drives architectural standards for the full-service lifecycle-from initial design and deployment readiness to proactive production operations. At Vertafore, we view reliability as a core engineering responsibility. You will operate autonomously across AWS, hybrid data centers, and customer-hosted environments, setting the technical direction for how we treat operations as a software engineering challenge. This role is pivotal in transitioning cross-departmental teams toward a highly proactive, engineering-first culture., Over the past 50 years, Vertafore has advanced the entire insurance distribution channel with the best software solutions in the industry. Today, weâre proud to say hundreds of thousands of Vertafore users rely on our solutions to write business faster, reduce costs, and fuel growth by increasing collaboration and streamlining processes. Vertafore leads the industry with secure, cloud-based mobile products that provide superior reporting and analytics, delivering actionable insight- right when customers need it most. We partner with other leading technology companies to deliver comprehensive solutions to improve the way our customers do business and serve their customers.
The Vertafore Way
Insurance is about relationships, and technology should make those relationships stronger. Thatâs why, at Vertafore, itâs our mission to transform the way the industry operates by putting people at the heart of insurance technology. By focusing on our customers, becoming better every day, and delivering results you can see, we provide the level of trust and security that insurance is all about. Bias to Action: Weâre united by an innate drive to take action and make a difference in the technology and insurance spaces. Win Together: We work together as one team, showing empathy and respect along the way. Show Up Curious: We work to challenge one another to push boundaries and think beyond the box. Say It, Do It: We honor every one of our commitments because integrity is important to us. Customer Success is Our Success: We cultivate authentic relationships and follow up by actively listening to their needs. We Love Insurance: We appreciate the impact insurance has on the world.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Is Software Engineering Over-Saturated?
The 12 Best Jobs for Software Engineers
How Much Does a Software Engineer Make? Realistic Software Engineering Salaries