> Markdown version of [/jobs/ext/1202746-site-reliability-engineer-sre-platform-engineer](https://www.wearedevelopers.com/jobs/ext/1202746-site-reliability-engineer-sre-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (SRE) / Platform Engineer - **Company:** Wintrio Llc - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Engineering, Configuration Management, Cyber Security, Information Systems, Continuous Delivery, Continuous Integration, DevOps, Distributed Systems, Monitoring of Systems, Python (Programming Language), Performance Tuning, Windows PowerShell, Reliability Engineering, Prometheus, Software Engineering, Data Logging, Scripting, Cloud Platform System, Spring Cloud, System Availability, Grafana, HybridCloud, Infrastructure as Code (IaC), Containerization, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Hashicorp, Terraform, Splunk, Dynatrace, Devsecops, Docker, Elk Stack, Jenkins, Golang - **Published:** July 8, 2026 - **Apply:** https://www.wintrio.com/careers/site-reliability-engineer-sre-platform-engineer/ ## About the Role * Bachelor's degree in Computer Science, Information Technology, Engineering, Information Systems, or a related field, or equivalent professional experience. * Minimum five (5) years of experience in Site Reliability Engineering (SRE), Platform Engineering, DevOps, Cloud Engineering, or Infrastructure Engineering. * Strong understanding of distributed systems, cloud computing, and enterprise infrastructure. * Experience implementing monitoring, logging, observability, and automation solutions. * Experience supporting AWS, Microsoft Azure, or hybrid cloud environments. * Experience with scripting languages such as Python, Go, Bash, or PowerShell. * Strong analytical, troubleshooting, and problem-solving skills. * Strong written and verbal communication skills. Technical Areas Site Reliability Engineering * High Availability * Reliability Engineering * Platform Engineering * Operational Excellence * Capacity Planning * Performance Optimization * Resiliency Engineering * Incident Management Observability & Monitoring * Monitoring * Logging * Alerting * Distributed Tracing * Metrics Collection * Operational Dashboards * Service-Level Objectives (SLOs) * Service-Level Indicators (SLIs) Cloud Infrastructure * Amazon Web Services (AWS) * Microsoft Azure * Hybrid Cloud Environments * Cloud Infrastructure * Platform Services Automation & Platform Operations * Infrastructure Automation * Platform Automation * Configuration Management * Operational Automation * CI/CD Support * Release Engineering Container Platforms * Docker * Kubernetes * Container Orchestration * Cloud-Native Platforms Tools & Platforms Monitoring & Observability * Prometheus * Grafana * ELK Stack * Splunk Cloud Platforms * Amazon Web Services (AWS) * Microsoft Azure Container Platforms * Docker * Kubernetes Infrastructure Automation * Terraform * Infrastructure as Code (IaC) CI/CD Platforms * Jenkins * GitLab CI/CD * Azure DevOps Scripting & Programming * Python * Go * Bash * PowerShell Preferred Certifications * Certified Kubernetes Administrator (CKA) * AWS Certified DevOps Engineer * Microsoft Azure DevOps Engineer Expert * Google Professional Cloud DevOps Engineer (or equivalent SRE certification), * Experience supporting highly available, mission-critical enterprise systems. * Experience supporting Federal Government or other regulated environments. * Experience implementing observability platforms and reliability engineering best practices. * Experience supporting Kubernetes, containerized platforms, or cloud-native applications. * Familiarity with Federal cybersecurity and compliance frameworks, including NIST RMF, FedRAMP, or FISMA. * Experience supporting enterprise modernization, platform engineering, or digital transformation initiatives. ## Description WINTrio LLC is seeking an experienced Site Reliability Engineer (SRE) / Platform Engineer to support mission-critical Federal cloud environments by improving system reliability, scalability, performance, and operational resilience. This role is responsible for designing and implementing highly available cloud platforms, developing observability solutions, automating infrastructure and operational processes, and optimizing the reliability of distributed systems. The successful candidate will work closely with Cloud Engineers, DevSecOps Engineers, Software Developers, and Cybersecurity teams to build resilient platforms that support continuous delivery, high availability, and enterprise-scale operations. The ideal candidate will possess strong experience with cloud platforms, monitoring, automation, incident response, performance optimization, and modern platform engineering practices. Job Responsibilities * Design and implement reliability, scalability, and resiliency strategies for cloud-native and distributed systems. * Build, configure, and maintain enterprise monitoring, logging, alerting, and observability platforms. * Automate infrastructure provisioning, deployments, operational workflows, and platform management tasks. * Monitor application, infrastructure, networking, and cloud platform performance to proactively identify operational issues. * Troubleshoot system outages, application failures, performance bottlenecks, and infrastructure incidents. * Lead incident response activities, root cause analysis (RCA), post-incident reviews, and corrective action planning. * Develop capacity planning strategies and optimize system utilization, performance, and scalability. * Support CI/CD pipelines, release engineering, platform engineering, and automation initiatives. * Collaborate with software development, DevSecOps, cloud infrastructure, cybersecurity, and operations teams to improve platform reliability. * Develop operational dashboards, service-level objectives (SLOs), service-level indicators (SLIs), and reliability metrics. * Create and maintain technical documentation, operational runbooks, and standard operating procedures (SOPs). * Support continuous improvement initiatives that enhance system availability, operational efficiency, and customer experience. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)