Software Engineer - SRE
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
When the Tarr Steps, a footbridge assembled of heavy stones in Exmoor National Park in England, washed away in a flood in 1942, the Royal Engineers rebuilt it. Then it washed away again in 1952. So the Royal Engineers heaved in more stones: on and on, like a Sisyphean lesson in absurdity.
Of course, they do this because the Tarr Steps are a historical monument, believed to be built during the Bronze Age. Modern bridge building looks a lot different and generally requires less upkeep. While we appreciate the charm of ancient things, Mercury is engineering the future of banking*. āQuaintā and āarchaicā are not values we seek in our systems. Rebooting a tumbling server and hoping a flood of requests doesnāt wash it away can be an emergency tactic, but we actively seek out durable solutions over monotonous ops work. Youāll bring this ethos and teach others how to live it too.
Up to this point, the Stability team at Mercury has primarily built platform-level constructs that product teams adopt. However, those teams are asking for us to work more closely with them to mature their implementations and practices. We are creating an SRE team that will rotate through product teams. You will understand the teamās domain, identify opportunities for improvements in reliability/observability/performance/preparedness and create self-reinforcing, virtuous cycles.
As part of this role, you will:
- Embed with product teams, helping them improve their operational maturity by setting up and refining practices around on-call, monitoring, alerting, and run books
- Run regular game day exercises with product teams, helping them feel more prepared to investigate and quickly remediate incidents
- Jump into application code written in Haskell & TypeScript and implement reliability techniques such as retries, better error handling, better logging, circuit breaking, etc
- Steer SLOs towards meaningful customer outcomes that product teams are accountable for. Those SLOs become a strong signal for whether the product is working as intended
- Champion reliability practices through design document and code reviews
- Identify observability gaps that hinder debugging, incident response, and business intelligence and help close those
- Advocate for longer-term improvements that non-product engineering teams can drive
- Participate in the product teamās on-call rotation while embedding or as part of a more general engineering rotation, helping us improve processes and how we learn from incidents
Requirements
- Has past Site Reliability Engineering or DevOps experience
- Has measurable examples of influencing an organization towards greater reliability
- Has significant experience with PostgreSQL
- Has authored and operated Temporal workflows
- Has experience with observability platforms like Grafana or Honeycomb
- Has familiarity with OpenTelemetry
Benefits & conditions
The total rewards package at Mercury includes base salary, equity (stock options/RSUs), and benefits.
Our salary and equity ranges are highly competitive within the SaaS and fintech industry and are updated regularly using the most reliable compensation survey data for our industry. New hire offers are made based on a candidateās experience, expertise, geographic location, and internal pay equity relative to peers.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this roleā¦