Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+6 more
Job description
The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, performance, and scalability of both internal (in-house) and customer environments. This role acts as a critical bridge between the product platform engineering, development and customer operation team, ensuring seamless deployment, operation, and support of our solutions across many environments., * Maintains and optimizes availability, performance, and resilience of in-house and on-premises productive environments.
-
Monitors system health, troubleshoots incidents, and implements proactive reliability solutions.
-
Manages deployments, upgrades, and patching platform level changes of multiple customers.
-
Demonstrates a customer-focused approach and ownership over production systems.
-
Serves as technical liaison to customers (internal and external), explaining TGW’s network, software and hardware platform architecture along with delivery and deployment approach.
-
Performs hands-on standing up of customer environments, and ongoing operations.
-
Develops and tests high-availability and disaster recovery plans for customers.
-
Engages with project management and customers on infrastructure and platform topics.
-
Collaborates with Platform Engineering and Global Infrastructure teams to ensure TGW’s standard deliverable is continuously improved with customer real world feedback.
-
Partners with solutions architects to ensure successful delivery of customer projects.
-
Documents issues and best practices for incident prevention and resolution.
-
Improves observability (logging, metrics, alerting) across customer environments.
-
Drives root cause analysis and collaborates with respective teams to implement preventative measures.
Requirements
Education: Bachelor’s degree in Computer Science, Information Technology, Computer Engineering, or a related field (or equivalent practical experience)
Experience: Minimum of five (5) years’ experience of delivering and supporting productive IT systems following Infrastructure as Code practices.
Travel: Up to 20% domestic and international travel.
Skills & Abilities
-
Hands-on experience with Linux, Kubernetes (Preferred Red Hat OpenShift) and GitOps delivery workflows.
-
Experience with debugging YAML, Helm & Terraform.
-
Strong troubleshooting skills across application, infrastructure, and networking layers.
-
Ability to communicate effectively with a variety of audiences, internal and external.
Preferred Skills
-
Proficiency in monitoring and observability tools (Prometheus, Grafana).
-
Familiarity with hypervisor layer and container orchestration platforms.
-
Familiarity with relational (Oracle, PostgreSQL) and non-relational databases.
-
Experience supporting customer owned and remote on-premises environments.
Physical Requirements
-
Ability to remain stationary at a desk for prolonged periods of time.
-
Ability to go to site frequently and move safely around industrial and/or warehouse environments.
-
Ability to lift and carry supplies up to 25 pounds at a time.
-
Ability to tolerate exposure to job site temperature fluctuations due to seasonal weather in geographic regions.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Where To Find Software Engineering Jobs
Is Software Engineering Over-Saturated?
Find a Developer Job: 12 Best Job Sites For Developers