Site Reliability Engineer

AdaptiveVets Solutions, Inc.
United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$113,050.0 - $133,000.0
Working hours
Shift work
Languages
English
Job source

Tech stack

Agile Methodology Amazon Web Services Cloud Computing CompTIA Security+ Monitoring of Systems Reliability Engineering Data Logging Mttr Kubernetes Information Technology Splunk Dynatrace

Job description

  • Implement and manage platform observability using Dynatrace and Splunk, including monitoring, alerting, centralized logging, and operational dashboards.
  • Support incident response operations, coordinating with the Monitoring and Incident Management Manager to detect, escalate, and resolve incidents within required SLA timeframes.
  • Implement and maintain automated recovery procedures including node failure detection, cordon/drain/replace workflows, and failover automation.
  • Participate in 24x7 on-call rotation to maintain platform availability and respond rapidly to Critical and High severity incidents.
  • Proactively engage in PI planning alongside supported application teams to address infrastructure dependencies and capacity planning.
  • Develop and maintain operational runbooks, SRE automation scripts, and post-incident root cause analysis documentation.
  • Monitor resource saturation metrics and autoscaling policy effectiveness, alerting before defined thresholds are exceeded.
  • Track and report SRE metrics including availability, incident frequency, MTTR, and reliability trends.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, IT, or related field (4 additional years of relevant experience may substitute)
  • 5+ years of site reliability engineering or cloud operations experience in Kubernetes-based environments
  • Hands-on experience with Dynatrace, Splunk, and enterprise monitoring and alerting platforms
  • Experience with AWS GovCloud operations, Kubernetes/EKS node management, and automated recovery procedures
  • Ability to support 24x7 on-call operations and respond rapidly to Critical incidents
  • Proficiency with SRE automation scripting and root cause analysis methodologies

Preferred Qualifications

  • Experience with SRE frameworks including SLI/SLO/SLA definition and error budget management
  • AWS Certified SysOps Administrator or equivalent
  • Certified Kubernetes Administrator (CKA)

Education:

  • Bachelor’s degree in Computer Science, Engineering, or IT (Required)

Experience:

  • Site reliability engineering in cloud Kubernetes environments: 5 years (Required)

License/Certification:

  • AWS Certified SysOps Administrator (Preferred)
  • Certified Kubernetes Administrator (CKA) (Preferred)
  • CompTIA Security+ (Preferred)

Location and Ability to Commute:

  • Remote

Security clearance:

  • Public Trust (Preferred), * Bachelor’s (Required)

Experience:

  • Agile: 4 years (Preferred)
  • Site reliability engineering in cloud Kubernetes environment: 5 years (Required)

Language:

  • English (Required)

License/Certification:

  • AWS Certified SysOps Administrator (Preferred)
  • Certified Kubernetes Administrator (CKA) (Preferred)
  • CompTIA Security+ (Preferred)

Benefits & conditions

$113,050 - $133,000 a year - Full-time, Pulled from the full job description

  • Referral program
  • Tuition reimbursement
  • Parental leave
  • 401(k)
  • Health insurance
  • 401(k) matching
  • Paid time off, * 401(k)
  • 401(k) matching
  • Dental insurance
  • Health insurance
  • Health savings account
  • Life insurance
  • Paid time off
  • Parental leave
  • Referral program
  • Tuition reimbursement
  • Vision insurance

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · WWC 2025

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

12:08 min

Comparing Keptn orchestration capabilities against alternative software operators

Thomas Schütz · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all