> Markdown version of [/jobs/ext/2628310-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2628310-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Tgw Systems Inc. - **Location:** Grand Rapids, MI, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Computer Engineering, Software Debugging, Linux, Monitoring of Systems, Hardware Platform Interface, PostgreSQL, OpenShift, Oracle (Applications), Reliability Engineering, Prometheus, YAML, Data Logging, Grafana, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Non-relational Database, Terraform - **Published:** August 9, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=79dad8a2c3f4f6aa ## About the Role Education: Bachelor's degree in Computer Science, Information Technology, Computer Engineering, or a related field (or equivalent practical experience) Experience: Minimum of five (5) years' experience of delivering and supporting productive IT systems following Infrastructure as Code practices. Travel: Up to 20% domestic and international travel. Skills & Abilities * Hands-on experience with Linux, Kubernetes (Preferred Red Hat OpenShift) and GitOps delivery workflows. * Experience with debugging YAML, Helm & Terraform. * Strong troubleshooting skills across application, infrastructure, and networking layers. * Ability to communicate effectively with a variety of audiences, internal and external. Preferred Skills * Proficiency in monitoring and observability tools (Prometheus, Grafana). * Familiarity with hypervisor layer and container orchestration platforms. * Familiarity with relational (Oracle, PostgreSQL) and non-relational databases. * Experience supporting customer owned and remote on-premises environments. Physical Requirements * Ability to remain stationary at a desk for prolonged periods of time. * Ability to go to site frequently and move safely around industrial and/or warehouse environments. * Ability to lift and carry supplies up to 25 pounds at a time. * Ability to tolerate exposure to job site temperature fluctuations due to seasonal weather in geographic regions. ## Description The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, performance, and scalability of both internal (in-house) and customer environments. This role acts as a critical bridge between the product platform engineering, development and customer operation team, ensuring seamless deployment, operation, and support of our solutions across many environments., * Maintains and optimizes availability, performance, and resilience of in-house and on-premises productive environments. * Monitors system health, troubleshoots incidents, and implements proactive reliability solutions. * Manages deployments, upgrades, and patching platform level changes of multiple customers. * Demonstrates a customer-focused approach and ownership over production systems. * Serves as technical liaison to customers (internal and external), explaining TGW's network, software and hardware platform architecture along with delivery and deployment approach. * Performs hands-on standing up of customer environments, and ongoing operations. * Develops and tests high-availability and disaster recovery plans for customers. * Engages with project management and customers on infrastructure and platform topics. * Collaborates with Platform Engineering and Global Infrastructure teams to ensure TGW's standard deliverable is continuously improved with customer real world feedback. * Partners with solutions architects to ensure successful delivery of customer projects. * Documents issues and best practices for incident prevention and resolution. * Improves observability (logging, metrics, alerting) across customer environments. * Drives root cause analysis and collaborates with respective teams to implement preventative measures. ## Related Videos - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) - [CI/CD with Github Actions](https://www.wearedevelopers.com/videos/856-ci-cd-with-github-actions) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [GitOps keeps focus on apps, not on infrastructure](https://www.wearedevelopers.com/videos/182-gitops-keeps-focus-on-apps-not-on-infrastructure) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)