Senior Data Reliability Engineer

Elliptic Enterprises Limited
Greater London, UK
18 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Big Data Software as a Service Data Integrity Distributed Systems Reliability Engineering JavaScript Pagination Plugin Enterprise Software Applications Kubernetes Data Lineage

Job description

Responsible for a diverse suite of products, you will oversee SR of enterprise grade applications that sit on the critical path running 1000s of QPS. Elliptic is known for its extensive and reliable datasets and you will play a critical role in defining and building out a market-leading foundation for data quality and control. This means building the processes, culture, and frameworks that will power observability, quality, data lineage, and remediation to form an essential pillar of our data & intelligence platform., This is a cross team role, and you will have the full support of leadership and engineering in carrying out your responsibilities - it’s not all down to you, but you will show the rest of us what good looks like.

  • Evangelise SRE & DRE across engineering
  • Lead the charge on building out a framework for data quality that will provide our customers with strong guarantees about the fidelity of our data as well support our marketing and revenue functions
  • SRE as a function define and own the on-call process:
  • Quickly establishing a strong working knowledge of our systems
  • Commanding incidents
  • Running mop-ups
  • Ensuring follow-up actions are completed to your schedule
  • Evaluating and improving our existing E2E on-call process
  • Take part in the on-call rotation, one week every 4-5 weeks (24x7x365 coverage)
  • Evaluate, manage and maintain our existing solutions for monitoring, alerting, paging, response, documentation
  • Report on uptime, availability, performance, etc across our product suite
  • Write post-mortems for both internal and external consumption
  • Represent our SRE & DRE function on sales calls with tier one enterprise financial institutions
  • Work with product, sales and customer service to define SLAs for different products and use cases
  • Work with internal product teams to define SLOs for internal consumption and measurement
  • Work with our engineering teams directly to embed DRE practices, * £500 Remote working budget to set up your home office space

Requirements

  • Thrive under high pressure situations, and are able to make tough decisions quickly
  • Fail fast, own the failure; encourage a blame free engineering culture
  • Are an inspiring thought leader, and are able to take others with you on a journey
  • Aren’t afraid to get your hands dirty and dig into code across myriad technologies
  • Understand the importance of reliability in enterprise finance systems
  • Have strong opinions based on your experience that you evolve over time as you learn from others, * Proven experience at leveling up the quality and reliability of large datasets not just services and APIs
  • Experience leading site reliability for a high volume SaaS product
  • Supported distributed systems in AWS
  • The presence and empathy required to hold teams to account
  • Defined SLAs / SLOs both internal and client facing
  • Offered post mortems to enterprise clients (verbal and written)

Bonus Points for:

  • Having a genuine interest in the crypto ecosystem and being behind the mission of the company
  • Working knowledge of Kubernetes and the challenges presented

Benefits & conditions

  • Private Health Insurance - we use Vitality!
  • Full access to Spill Mental Health Support
  • Life Assurance: we hope you will never need this - but our cover is for 4 times your salary to your beneficiaries
  • £100 Crypto for you!
  • Cycle to Work Scheme

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

9:51 min

Live demonstration of dynamic data redaction plugins

Tom Kaltofen Tom Kaltofen +1 · Europe 2026 Virtual

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all