Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+19 more
Job description
They are now looking for an Site Reliability Engineer to become the first technical hire in the US. This is a high-impact role for someone who enjoys solving complex production issues, improving platform reliability and working closely with engineering, infrastructure and business stakeholders. As Site Reliability Engineer, you will play a key role in ensuring the reliability, scalability and smooth operation of the company’s platform and supporting infrastructure. You will contribute through incident response, proactive monitoring, automation, documentation and continuous improvement of support processes. You will also act as the on-the-ground engineering presence in the New York office, extending technical coverage into US working hours and helping bridge collaboration between US-based stakeholders and the UK engineering team. Duties and responsibilities
- Investigate, own and resolve incidents across the platform and infrastructure.
- Work closely with cross-functional teams to ensure a rapid and effective response to production issues.
- Build and maintain runbooks, post-mortems and internal documentation to support knowledge sharing.
- Improve observability through monitoring, alerting and dashboards.
- Reduce operational overhead through automation and reliability-focused engineering.
- Contribute to the ongoing improvement of support processes, tooling and applications.
- Train, coach and support members of the UK on-call support rota.
Requirements
- A strong background in application support, site reliability engineering or production support.
- Experience supporting production environments and resolving live incidents.
- Understanding of SRE principles, including SLIs, SLOs, error budgets and blameless post-mortems.
- Experience with observability and monitoring tools such as Datadog.
- Experience with cloud platforms such as AWS or Azure.
- Proficiency in at least one programming or scripting language, such as TypeScript, .NET, Bash or PowerShell.
- Knowledge of SQL and experience working with databases.
- Understanding of networking fundamentals, Linux and Windows systems.
- Strong problem-solving skills and the ability to remain calm under pressure.
- Excellent communication skills, with the ability to explain technical concepts to technical and non-technical stakeholders.
Nice to have
- Experience with infrastructure-as-code tools such as Terraform, CloudFormation or CDK.
- Knowledge of DevOps deployment practices and tooling such as TeamCity, Octopus Deploy or Docker.
- Experience in the financial services or fintech sector.
- Exposure to AI/ML platforms, prompt engineering or integrations with LLM APIs.
About the company
Site Reliability Engineer - AWS, Azure, IaC, Typescript, .NET - NY (office 3 days a week) - $120,000 - $140,000
Do you want to be the first technical hire in the US for a fast-growing leading FX hedging platform?
Do you want to be the cornerstone for the team as it grows?
This is a fantastic opportunity to a well-established UK business that has recently received some significant investment and is now breaking into the US market. The UK has been comfortably managing the UK coverage; however, they are growing quickly and they need a full-time US presence.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Find a Developer Job: 12 Best Job Sites For Developers
Dev Digest 120 - Apple and peers
Where To Find Software Engineering Jobs