Principal Site Reliability Engineer

The Vertafore Way
United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$160,000.0 - $180,000.0
Working hours
Shift work
Job source

Tech stack

Java (Programming Language) .NET Framework Amazon Web Services C Sharp (Programming Language) Cloud Computing Continuous Integration Relational Databases Linux Distributed Systems Fault Tolerance Python (Programming Language) Software Engineering
+5 more
ReactJS Kubernetes Information Technology Low Latency Data Analytics

Job description

  • Enterprise-Wide Ownership: Define the standards for end-to-end service ownership, holding the organization accountable for availability, performance, and overall operational health.

  • Architectural Influence: Lead cross-departmental initiatives to influence system design at the architectural level, driving fault tolerance, strict compliance, and operational sustainability across public and private clouds.

  • Advanced Observability Vision: Dictate the enterprise strategy for observability frameworks, ensuring the Four Golden Signals (Latency, Traffic, Errors, and Saturation) provide actionable, predictive insights across all platforms.

Strategic Leadership & Reliability Architecture

  • Enterprise-Wide Ownership: Define the standards for end-to-end service ownership, holding the organization accountable for availability, performance, and overall operational health.

  • Architectural Influence: Lead cross-departmental initiatives to influence system design at the architectural level, driving fault tolerance, strict compliance, and operational sustainability across public and private clouds.

  • Advanced Observability Vision: Dictate the enterprise strategy for observability frameworks, ensuring the Four Golden Signals (Latency, Traffic, Errors, and Saturation) provide actionable, predictive insights across all platforms.

Data-Driven Reliability Governance

  • SLO & Error Budget Authority: Establish the governance models for defining and managing SLIs and SLOs across multiple product lines.

  • Delivery Alignment: Champion Error Budgets as the ultimate technical arbiter at the executive level, balancing feature velocity with the absolute requirement for platform stability.

Incident Management & Cultural Transformation

  • Enterprise Incident Command: Lead incident response for the most critical, high-severity events.

  • Blameless Culture Champion: Foster a ā€œWin Togetherā€ environment by championing a Blameless Postmortem culture globally, ensuring root cause analyses focus strictly on systemic and process improvements rather than individual error.

Requirements

  • Experience: 12 to 15+ years of hands-on Cloud Operations, SRE, or reliability-focused engineering experience, with a proven track record of end-to-end enterprise service ownership.

  • Proven Scope: Demonstrated ability to operate at a Principal/Architect scope, driving large-scale reliability outcomes and operational excellence across global organizations.

  • Software Engineering: Expert-level software engineering skills in C#, .NET, Java, Python, or React.

  • Principles: Deep expertise in scaling core SRE principles (SLIs, SLOs, error budgets) across complex, distributed systems.

  • Technical Stack: Mastery of AWS, Kubernetes, CI/CD pipelines, Infrastructure-as-Code, and extensive knowledge of Linux and Windows environments and relational databases.

  • Education: Bachelor’s or Master’s degree in Computer Science or a related technical field.

  • Commitment: Participation in an executive on-call rotation with flexible hours as required

Skills & Requirements Knowledge, Skills, and Abilities: A fast learner. A problem solver. Ability to document procedures. Able to meet deadlines. Good communication skills. Able to deliver the message effectively to a technical and non-technical audience. Able to comply with processes and procedures. Able to maintain professional composure in any situations. Flexible in working extended hours on occasions or as required. Exposure in the insurance industry is desired but not mandatory. Driven to improve, personally and professionally Operate best in a fast-paced, flexible work environment with ability to work in a team.

Additional Requirements and Details: High speed internet to accommodate working from home needs. Occasional lifting and/or moving up to 10 pounds. Frequent repetitive hand and arm movements required to operate a computer. Specific vision abilities required by this job include close vision (working on a computer, etc.). Frequent sitting and/or standing., The selected candidate must be legally authorized to work in the United States.

Benefits & conditions

We do not accept resumes from agencies, headhunters, or other suppliers who have not signed a formal agreement with us. Qualifications Vertafore is a Flexible First working environment which allows team members to work from home as often as you’d like, while using our offices as a place for collaboration, community, and teambuilding. There are times you may be asked to come into an office and/or travel for specific meetings for a specific business purpose and this varies by job responsibilities.

Why Vertafore is the place for you: *Canada Only

  • The opportunity to work in a space where modern technology meets a stable and vital industry
  • Medical, vision & dental plans
  • Life, AD&D
  • Short Term and Long Term Disability
  • Pension Plan & Employer Match
  • Maternity, Paternity and Parental Leave
  • Employee and Family Assistance Program (EFAP)
  • Education Assistance
  • Additional programs - Employee Referral and Internal Recognition

Why Vertafore is the place for you: *US Only

  • The opportunity to work in a space where modern technology meets a stable and vital industry
  • We have a Flexible First work environment! Our North America team members use our offices for collaboration, community and team-building, with members asked to sometimes come into an office and/or travel depending on job responsibilities. Other times, our teams work from home or a similar environment.
  • Medical, vision & dental plans
  • PPO & high-deductible options
  • Health Savings Account & Flexible Spending Accounts Options:
  • Health Care FSA
  • Dental & Vision FSA
  • Dependent Care FSA
  • Commuter FSA
  • Life, AD&D (Basic & Supplemental), and Disability
  • 401(k) Retirement Savings Plain & Employer Match
  • Supplemental Plans - Pet insurance, Hospital Indemnity, and Accident Insurance
  • Parental Leave & Adoption Assistance
  • Employee Assistance Program (EAP)
  • Education & Legal Assistance
  • Additional programs - Tuition Reimbursement, Employee Referral, Internal Recognition, and Wellness
  • Commuter Benefits (Denver)

About the company

Vertafore is a leading technology company whose innovative software solutions are advancing the insurance industry. Our suite of products provides solutions to our customers that help them better manage their business, boost their productivity and efficiencies, and lower costs while strengthening relationships.

Our mission is to move InsurTech forward by putting people at the heart of the industry. We are leading the way with product innovation, technology partnerships, and focusing on customer success.

Our fast-paced and collaborative environment inspires us to create, think, and challenge each other in ways that make our solutions and our teams better.

We are headquartered in Denver, Colorado, with offices across the U.S., Canada, and India.

We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise-wide reliability, scalability, and performance of our critical production services. As a foundational pillar of our engineering organization, this role drives architectural standards for the full-service lifecycle-from initial design and deployment readiness to proactive production operations. At Vertafore, we view reliability as a core engineering responsibility. You will operate autonomously across AWS, hybrid data centers, and customer-hosted environments, setting the technical direction for how we treat operations as a software engineering challenge. This role is pivotal in transitioning cross-departmental teams toward a highly proactive, engineering-first culture., Over the past 50 years, Vertafore has advanced the entire insurance distribution channel with the best software solutions in the industry. Today, we’re proud to say hundreds of thousands of Vertafore users rely on our solutions to write business faster, reduce costs, and fuel growth by increasing collaboration and streamlining processes. Vertafore leads the industry with secure, cloud-based mobile products that provide superior reporting and analytics, delivering actionable insight- right when customers need it most. We partner with other leading technology companies to deliver comprehensive solutions to improve the way our customers do business and serve their customers.

The Vertafore Way

Insurance is about relationships, and technology should make those relationships stronger. That’s why, at Vertafore, it’s our mission to transform the way the industry operates by putting people at the heart of insurance technology. By focusing on our customers, becoming better every day, and delivering results you can see, we provide the level of trust and security that insurance is all about. Bias to Action: We’re united by an innate drive to take action and make a difference in the technology and insurance spaces. Win Together: We work together as one team, showing empathy and respect along the way. Show Up Curious: We work to challenge one another to push boundaries and think beyond the box. Say It, Do It: We honor every one of our commitments because integrity is important to us. Customer Success is Our Success: We cultivate authentic relationships and follow up by actively listening to their needs. We Love Insurance: We appreciate the impact insurance has on the world.

Is this role not an exact fit for you? Keep an eye on our for other positions!

Vertafore is a drug free workplace and conducts preemployment drug and background screenings.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo Ā· LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard Ā· WWC 2025

1:21 min

Exploring the target application for front end tests

Anna Mcdougall Ā· JS Congress

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter Ā· WWC 2022

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 Ā· WWC 2021

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani Ā· Europe 2026 Virtual

Videos

See all

Related articles

See all