> Markdown version of [/jobs/ext/2727526-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2727526-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** The Trainline - **Location:** London, UK - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Agile Methodology, Amazon Web Services, Applications Architecture, Configuration Management, Continuous Integration, Linux, Github, Python (Programming Language), Operational Data Store, Reverse Proxy, Data Streaming, Datadog, Scripting, Grafana, Terraform, New Relic (SaaS), Docker, Elk Stack - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/site-reliability-engineer-trainlinegroup-com-8578275 ## About the Role We're looking for a mid-level Site Reliability Engineer to help drive this forward. You'll bring solid production experience, a growth mindset, and a willingness to challenge and be challenged - contributing to platform reliability while developing broader technical ownership with support from senior engineers., * Experience of SRE concepts such as SLI, SLO and error budgets. * Hands-on experience with observability tooling such as New Relic, Elastic (ELK Stack), Influx, Grafana or similar * Experience working with cloud providers (preferably AWS). * Experience troubleshooting Linux operating systems. * Experience of scripting in at least one language (preferably Python) * Understanding of load balancing and reverse proxy concepts, upstream config concepts, upstream health checks, worker & data flow concepts. * Application architecture concepts (threading, queuing, readiness checks, health checks, circuit breakers, timeouts, exponential backoff, throttling). * Experience building, maintaining and evolving time series data, retention, cardinality, deviation, moving averages and other functions. * Experience with build, deployment & configuration management tooling such as GitHub Actions and Terraform. ## Description * Developing an understanding of system architecture, dependencies, and failure modes across the Trainline platform * Participating in production incident response, supporting investigation, mitigation, communication, and coordinated service restoration * Contributing to post-incident reviews and follow-up actions to improve reliability, scalability, and resilience * Taking part in the SRE on-call rotation * Designing, building, and maintaining observability using metrics, logs, events, and traces to support effective detection and diagnosis * Improving monitoring and alerting by aligning signals to business and customer impact, reducing noise and improving mean time to detection (MTTD) * Ensuring relevant operational data is surfaced quickly and clearly during live incidents * Making informed tooling and technology choices using SRE principles, balancing team and business needs * Supporting AWS-hosted infrastructure and shared platform services using infrastructure-as-code and CI/CD tooling * Collaborating with product engineering teams to ensure services are operationally ready and deployed safely * Advising on reliability and resilience practices * Writing and maintaining reliable, well-structured code and scripts to support reliability and observability goals * Prioritising work effectively and collaborating using agile processes to deliver against team and business goals Our Tech Stack * AWS * New Relic * ELK stack * Grafana * Incident.io * Docker, ECS * Terraform * Github Actions ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Fullstack Developer Salary UK](https://www.wearedevelopers.com/magazine/251-fullstack-developer-salary-uk) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk)