> Markdown version of [/jobs/ext/2028216-site-reliability-engineer-expert-specialist-devops](https://www.wearedevelopers.com/jobs/ext/2028216-site-reliability-engineer-expert-specialist-devops). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer/ Expert/ Specialist (DevOps) - **Company:** Stovall Custom Woodworking, Inc. - **Location:** Stanford, IL, United States (Remote available) - **Experience:** Expert - **Salary:** $92,300.0 - $166,850.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Computing, CompTIA Security+, DevOps, Information Technology Operations, Reliability Engineering, Service Design, System Availability, Containerization, Kubernetes, Information Technology, Deployment Automation, Performance Monitor, Programming Languages - **Published:** August 11, 2026 - **Apply:** https://www.careerjet.com/job/usb25aa7116acd18639d5bd2dea5369bb1/eaa ## About the Role * Bachelor's degree in computer science, Information Technology, Engineering, or a related field. * Several years of experience in IT operations, service management, or infrastructure management, including roles such as Site Reliability Engineer, Problem Manager, or DevOps Manager. * Proven experience in managing high-availability systems and ensuring operational reliability. * Extensive experience in root cause analysis (RCA), incident management, and developing permanent solutions for recurring service disruptions. * Hands-on experience with CI/CD pipelines, automation, system performance monitoring, and the implementation of infrastructure as code. * Strong background in collaborating with cross-functional teams (development, operations, engineering, etc.) to improve operational processes and service delivery. * Experience in managing deployments, risk assessments, and optimizing event and problem management processes. * Familiarity with cloud technologies, containerization, and scalable architecture, including experience with zero-downtime deployment strategies. * CompTIA Security+ or Certified Kubernetes Administrator (CKA). Functional Skills: * Collaboration * Stakeholder Management * Service Design * Communication * Problem Solving * Incident Management * Change Management * Innovation ; Technical Skills: * Cloud Infrastructure * Automation & AI * Operations Monitoring & Diagnostics * Deployment * Programming & Scripting Languages ## Description At SITA, we keep airports moving, airlines flying smoothly, and borders open. Our technology and communication innovations power the success of the global air travel industry. You'll find us in 95% of international airports, working closely with over 2,500 transportation and government clients. Each partnership brings unique challenges, and we thrive on delivering fresh solutions and cutting-edge tech to keep operations running like clockwork. We don't just move the world forward-we're proud to be recognized as a Great Place to Work® by 79% of our employees and certified in most of our growing locations. Here, we feel empowered, supported, and inspired to grow. Are you ready to love your job? The adventure begins right here, with you, at SITA. ABOUT THE ROLE & TEAM The Site Reliability Engineer is responsible for the proactive support of products to ensure high product performance, with a continuous focus on improvement. The role involves identifying and resolving the root causes of operational incidents, implementing solutions to enhance stability, and preventing recurrence. The Site Reliability Engineer manages the creation and maintenance of the event catalogue to trigger events and develops both manual remediation approaches and automated workflows to address alerts. Additionally, they oversee the deployment of IT services and solutions, ensuring seamless integration with minimal disruption. WHAT YOU'LL DO * Design, build, and maintain support systems to ensure high availability, scalability, and performance of critical infrastructure. * Lead incident response and root cause analysis for system failures, including problem investigations and coordination with relevant teams. * Implement and manage automation for system provisioning, deployment, self-healing, and performance monitoring to increase operational efficiency. * Establish and monitor SLIs/SLOs, proactively identify performance issues, and drive continuous improvements in service reliability. * Collaborate with development and operations teams to embed reliability best practices and evolve toward zero-downtime architecture. * Manage and optimize an event catalog, including event definitions, thresholds, remediation actions, and relevance across products. * Develop event response protocols, provide training, and ensure efficient handling of incidents across teams. * Drive post-incident reviews and feedback loops to enhance event definitions and service reliability. * Oversee quality and readiness of deployments, ensuring clear processes, assigned responsibilities, and minimal disruption. * Maintain deployment schedules and conduct risk assessments to ensure operational stability and deployment readiness. * Coordinate and execute deployment plans, manage resources, and incorporate feedback for continuous process improvement. * Manage CI/CD pipelines and infrastructure as code, ensuring seamless integration between development and operations. * Support and evolve DevOps practices, automating operational tasks and maintaining tools to drive ongoing efficiency. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [What If We Designed Employee Experience Like a Customer Journey?](https://www.wearedevelopers.com/videos/1824-what-if-we-designed-employee-experience-like-a-customer-journey) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)