Senior Site Reliability Engineer (SRE)

CenCore LLC
United States
29 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon Cloudfront Amazon S3 Application Performance Management Systems Engineering Cloud Computing Cloud Computing Security DevOps Disaster Recovery Identity and Access Management Key Management PostgreSQL
+20 more
Routing Performance Tuning Reliability Engineering Cloud Services Prometheus Software Engineering Software Vulnerability Management Datadog Data Logging Load Balancing Spring Cloud System Availability Grafana Amazon Virtual Private Cloud (VPC) Amazon Relational Database Service Kubernetes Infrastructure Automation Frameworks Route53 Cloudwatch Terraform

Job description

The Senior Site Reliability Engineer (SRE) will implement, secure, and operate the cloud infrastructure that supports CenCore Group’s proprietary enterprise SaaS platform. This role is responsible for maintaining a scalable, highly available, secure, and reliable cloud environment as the platform grows and supports enterprise customers., * Manage, maintain, and improve AWS-based cloud infrastructure supporting enterprise SaaS operations.

  • Operate and support Kubernetes environments, including Amazon EKS.
  • Own platform reliability, scalability, availability, disaster recovery readiness, and operational resilience.
  • Design and support cloud networking, load balancing, routing, traffic management, and related infrastructure components.
  • Implement and maintain monitoring, alerting, logging, and observability solutions to support proactive issue detection and response.
  • Establish and document operational standards, Service Level Objectives (SLOs), incident response processes, and reliability best practices.
  • Partner with software engineering and product teams to improve application performance, platform stability, and deployment reliability.
  • Apply security best practices across IAM, secrets management, encryption, vulnerability remediation, access controls, and production operations.
  • Support production operations, troubleshoot critical issues, and participate in incident resolution as needed.

Requirements

  • Professional experience supporting cloud infrastructure, site reliability, DevOps, platform engineering, or systems engineering functions.
  • Hands-on experience with AWS cloud services and production cloud operations.
  • Experience administering or operating Kubernetes environments.
  • Working knowledge of infrastructure reliability, availability, scalability, incident response, and operational support practices.
  • Experience implementing monitoring, logging, alerting, or observability tools.
  • Ability to troubleshoot complex production issues and coordinate resolution across technical teams.
  • Strong understanding of cloud security fundamentals, including identity and access management, encryption, secrets management, and vulnerability remediation.
  • Ability to document technical processes, standards, and operational procedures.

Preferred Qualifications

  • Experience with AWS services such as EKS, ALB, VPC, CloudFront, Route 53, RDS/Aurora, S3, and IAM.
  • Experience with Terraform or other Infrastructure as Code tools.
  • Experience with monitoring platforms such as Datadog, CloudWatch, Grafana, Prometheus, or similar tools.
  • PostgreSQL administration, performance tuning, or database operations experience.
  • Experience supporting enterprise SaaS, cloud-native applications, or customer-facing production platforms.
  • Experience developing disaster recovery, operational readiness, or production support documentation.

Skills / Competencies

  • Cloud infrastructure operations and automation
  • Platform reliability, scalability, and performance optimization
  • Kubernetes administration and containerized application support
  • Monitoring, observability, and incident response
  • Cloud security and operational risk awareness
  • Technical troubleshooting and root cause analysis
  • Cross-functional collaboration with engineering, product, and operations teams
  • Clear technical documentation and process improvement

Benefits & conditions

Health insurance, 401(k) matching, Paid time off, Vision insurance, Dental insurance, Paid holidays Full-time Remote, Eligible employees may enroll in company-sponsored medical, dental, and vision benefits in accordance with plan terms and enrollment requirements. The company contributes toward employee healthcare premiums, subject to plan provisions.

Employees are also eligible to participate in the company’s 401(k) retirement savings plan, including employer matching contributions in accordance with plan guidelines and eligibility requirements.

Paid time off, sick leave, and holiday benefits are provided in accordance with company policy, applicable contract requirements, and federal, state, and local laws. Benefit eligibility and accrual rates may vary based on position, location, and length of service.

The company is committed to equitable pay practices and does not seek or rely on an applicant’s wage history when making compensation decisions.

Benefits may be modified from time to time in accordance with company policy and applicable law.

About the company

At CenCore Group, we deliver security solutions that support critical national security missions. We specialize in designing, building, securing, and maintaining advanced technology environments at the intersection of emerging technology and national security. We are seeking a dependable, cleared professional to join our team in support of secure physical security operations.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all