Site Reliability Engineer (SRE) / Platform Engineer

Wintrio Llc
United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Cloud Engineering Configuration Management Cyber Security Information Systems Continuous Delivery Continuous Integration DevOps Distributed Systems
+29 more
Monitoring of Systems Python (Programming Language) Performance Tuning Windows PowerShell Reliability Engineering Prometheus Software Engineering Data Logging Scripting Cloud Platform System Spring Cloud System Availability Grafana HybridCloud Infrastructure as Code (IaC) Containerization Gitlab-ci Kubernetes Infrastructure Automation Frameworks Information Technology Hashicorp Terraform Splunk Dynatrace Devsecops Docker Elk Stack Jenkins Golang

Job description

WINTrio LLC is seeking an experienced Site Reliability Engineer (SRE) / Platform Engineer to support mission-critical Federal cloud environments by improving system reliability, scalability, performance, and operational resilience.

This role is responsible for designing and implementing highly available cloud platforms, developing observability solutions, automating infrastructure and operational processes, and optimizing the reliability of distributed systems. The successful candidate will work closely with Cloud Engineers, DevSecOps Engineers, Software Developers, and Cybersecurity teams to build resilient platforms that support continuous delivery, high availability, and enterprise-scale operations.

The ideal candidate will possess strong experience with cloud platforms, monitoring, automation, incident response, performance optimization, and modern platform engineering practices.

Job Responsibilities

  • Design and implement reliability, scalability, and resiliency strategies for cloud-native and distributed systems.

  • Build, configure, and maintain enterprise monitoring, logging, alerting, and observability platforms.

  • Automate infrastructure provisioning, deployments, operational workflows, and platform management tasks.

  • Monitor application, infrastructure, networking, and cloud platform performance to proactively identify operational issues.

  • Troubleshoot system outages, application failures, performance bottlenecks, and infrastructure incidents.

  • Lead incident response activities, root cause analysis (RCA), post-incident reviews, and corrective action planning.

  • Develop capacity planning strategies and optimize system utilization, performance, and scalability.

  • Support CI/CD pipelines, release engineering, platform engineering, and automation initiatives.

  • Collaborate with software development, DevSecOps, cloud infrastructure, cybersecurity, and operations teams to improve platform reliability.

  • Develop operational dashboards, service-level objectives (SLOs), service-level indicators (SLIs), and reliability metrics.

  • Create and maintain technical documentation, operational runbooks, and standard operating procedures (SOPs).

  • Support continuous improvement initiatives that enhance system availability, operational efficiency, and customer experience.

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, Information Systems, or a related field, or equivalent professional experience.

  • Minimum five (5) years of experience in Site Reliability Engineering (SRE), Platform Engineering, DevOps, Cloud Engineering, or Infrastructure Engineering.

  • Strong understanding of distributed systems, cloud computing, and enterprise infrastructure.

  • Experience implementing monitoring, logging, observability, and automation solutions.

  • Experience supporting AWS, Microsoft Azure, or hybrid cloud environments.

  • Experience with scripting languages such as Python, Go, Bash, or PowerShell.

  • Strong analytical, troubleshooting, and problem-solving skills.

  • Strong written and verbal communication skills.

Technical Areas

Site Reliability Engineering

  • High Availability

  • Reliability Engineering

  • Platform Engineering

  • Operational Excellence

  • Capacity Planning

  • Performance Optimization

  • Resiliency Engineering

  • Incident Management

Observability & Monitoring

  • Monitoring

  • Logging

  • Alerting

  • Distributed Tracing

  • Metrics Collection

  • Operational Dashboards

  • Service-Level Objectives (SLOs)

  • Service-Level Indicators (SLIs)

Cloud Infrastructure

  • Amazon Web Services (AWS)

  • Microsoft Azure

  • Hybrid Cloud Environments

  • Cloud Infrastructure

  • Platform Services

Automation & Platform Operations

  • Infrastructure Automation

  • Platform Automation

  • Configuration Management

  • Operational Automation

  • CI/CD Support

  • Release Engineering

Container Platforms

  • Docker

  • Kubernetes

  • Container Orchestration

  • Cloud-Native Platforms

Tools & Platforms

Monitoring & Observability

  • Prometheus

  • Grafana

  • ELK Stack

  • Splunk

Cloud Platforms

  • Amazon Web Services (AWS)

  • Microsoft Azure

Container Platforms

  • Docker

  • Kubernetes

Infrastructure Automation

  • Terraform

  • Infrastructure as Code (IaC)

CI/CD Platforms

  • Jenkins

  • GitLab CI/CD

  • Azure DevOps

Scripting & Programming

  • Python

  • Go

  • Bash

  • PowerShell

Preferred Certifications

  • Certified Kubernetes Administrator (CKA)

  • AWS Certified DevOps Engineer

  • Microsoft Azure DevOps Engineer Expert

  • Google Professional Cloud DevOps Engineer (or equivalent SRE certification), * Experience supporting highly available, mission-critical enterprise systems.

  • Experience supporting Federal Government or other regulated environments.

  • Experience implementing observability platforms and reliability engineering best practices.

  • Experience supporting Kubernetes, containerized platforms, or cloud-native applications.

  • Familiarity with Federal cybersecurity and compliance frameworks, including NIST RMF, FedRAMP, or FISMA.

  • Experience supporting enterprise modernization, platform engineering, or digital transformation initiatives.

Benefits & conditions

  • Full-time position.

  • Remote within the United States.

  • Standard business hours Monday through Friday.

  • Occasional travel may be required in support of customer meetings, technical workshops, and program activities.

WINTrio Benefits

  • Healthcare (Medical, Dental, and Vision)

  • Flexible Spending Account (FSA) and Health Savings Account (HSA)

  • 401(k) and Retirement Savings Plan

  • Annual Bonus and Profit Sharing Opportunities

  • Paid Time Off (PTO) and Vacation

  • Employee Assistance Program (EAP)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.wintrio.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all