Site Reliability Engineer

Cadre
United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) JavaScript (Programming Language) Agile Methodology Amazon Web Services Application Performance Management Automation of Tests Cloud Computing DevOps Distributed Systems Python (Programming Language) Performance Tuning Reliability Engineering
+15 more
TypeScript Data Logging Scripting Cloud Platform System Spring Cloud Grafana AWS Lambda Deployment Automation Functional Programming Cloudwatch Api Gateway Splunk Dynatrace Serverless Computing Programming Languages

Job description

The Site Reliability Engineer will support a large federal technology modernization effort focused on improving the reliability, visibility, and performance of cloud-native applications and services across a national benefits platform. This role focuses heavily on observability, telemetry, monitoring, and performance engineering within a modern serverless environment.

You’ll work closely with development, operations, and platform teams to help build the standards, tooling, and engineering patterns used across serverless services. This role goes beyond writing Lambda code. You’ll help define how services are instrumented, deployed, monitored, and optimized across the platform.

What You’ll Do:

  • Build and maintain observability, telemetry, logging, and monitoring solutions for serverless applications and services
  • Support performance analysis, troubleshooting, and optimization efforts across distributed cloud environments
  • Develop and maintain engineering patterns and standards for AWS Lambda services
  • Implement instrumentation using AWS Distro for OpenTelemetry (ADOT)
  • Support monitoring and alerting capabilities using Dynatrace, Splunk, and related observability tools
  • Work with development and DevOps teams to integrate monitoring and telemetry into CI/CD pipelines
  • Assist with diagnosing system bottlenecks, latency issues, and application performance concerns
  • Support automated testing, deployment, and operational readiness activities
  • Help improve operational visibility, tracing, and logging consistency across environments
  • Participate in Agile delivery activities, release coordination, and operational support efforts

Requirements

Do you have experience in Technical troubleshooting support?, * 3+ years of experience supporting cloud-native applications, performance engineering, or observability platforms

  • Experience with AWS serverless technologies including Lambda, CloudWatch, API Gateway, and related services
  • Experience with observability and monitoring platforms such as Dynatrace, Splunk, Grafana, or similar tools
  • Familiarity with OpenTelemetry or AWS ADOT instrumentation practices
  • Experience supporting CI/CD pipelines and automated deployment workflows
  • Understanding of distributed systems, logging, tracing, and performance analysis concepts
  • Experience troubleshooting performance issues across cloud environments
  • Familiarity with scripting or development languages such as Python, JavaScript, TypeScript, or Java
  • Understanding of DevOps and Agile software delivery practices
  • Strong communication and collaboration skills
  • Federal government experience is a plus
  • Ability to obtain and maintain a Public Trust or other required government clearance

Benefits & conditions

Pulled from the full job description

  • Health insurance
  • 401(k) matching
  • Paid time off
  • Vision insurance
  • Dental insurance
  • Profit sharing
  • Paid holidays, Keeping people satisfied is no small undertaking. At CADRE, our commitment to supporting employees extends further than conventional benefits. While we offer competitive industry-standard packages, our focus goes beyond transactional care.

This is why we developed the CADRE Cares Program which encompasses a range of benefits and initiatives aimed at enhancing the overall well-being and job satisfaction of employees.

We want to support our employees at the times they need it the most. We envision unwavering commitment to our employees that sets a new standard for compassionate and impactful business practices.

CADRE Convoy Program:

A dedicated support system designed to cater to your individual needs and enhance your overall well-being, no matter where you are.

CADRE Connect Program:

A series of intentional touchpoints, events, and initiatives that promote open communication and encourage continuous growth and development among team members.

CADRE Compensation Program:

  • 401(k) Safe Harbor Plans with Matching & Immediate Vesting
  • Medical, Dental, & Vision Plans
  • Paid Time Off: Holidays, Vacation, Wellness, & Personal Leave Plans
  • Continuing Education & Training Budget
  • Office & Technology Budget
  • Cell Phone Budget
  • Wellness & Healthy Living Budget
  • Awards & Bonuses
  • Profit Sharing Plans

About the company

We’re more than a government contracting company. We’re a cadre, a specialized team built to support government agencies with practical, reliable solutions to complex operational and technology challenges. We strategically mobilize, manage, and maintain specialized cadres using our RO(M³) model.

CADRE GOVERNMENT SOLUTIONS is an Equal Opportunity and Affirmative Action Employer. We welcome and encourage diversity in our workforce.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all