Site Reliability Engineer, CCIP

Chainlink Labs
United States
27 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Distributed Systems Interoperability Performance Tuning Reliability Engineering Software Engineering Kubernetes Web3.js Dynatrace

Job description

As a Senior Site Reliability Engineer on the CCIP Platform team, you will ensure the reliability, scalability, and operational excellence of the systems powering Chainlink’s Cross-Chain Interoperability Protocol (CCIP). This role exists to strengthen production resilience, reduce operational toil, and enable engineering teams to ship safely while maintaining high service availability. You will influence reliability practices across the platform and help establish operational standards that scale with the business.

Your Impact

  • Improve deployment safety and increase delivery velocity by advancing production engineering practices.
  • Establish distributed tracing across the platform to improve observability and accelerate incident investigation.
  • Eliminate operational toil through automation that increases engineering efficiency and platform reliability.
  • Drive adoption of meaningful SLOs, SLIs, and error budgets that guide engineering decisions and improve service health.
  • Increase platform scalability and operational readiness as CCIP continues to grow.
  • Strengthen Chainlink’s reputation through highly available production systems while reducing operational overhead.

Requirements

  • Demonstrated experience in Site Reliability Engineering, Production Engineering, or a similar role operating large-scale distributed systems.
  • Deep expertise defining, implementing, and driving adoption of SLOs, SLIs, and error budgets across engineering organizations.
  • Built and operated production Kubernetes environments supporting critical services.
  • Applied OpenTelemetry to improve observability across distributed systems.
  • Experience improving the reliability, scalability, and operability of production infrastructure.

Preferred Requirements

  • Demonstrated technical leadership influencing reliability practices across engineering teams.
  • Experience performing capacity planning and performance tuning for high-throughput distributed services.
  • Previous experience working on Web3 infrastructure or within a crypto-native engineering organization.
  • Applied chaos engineering or fault-injection techniques to improve production resilience.
  • Partnered with software engineering teams to conduct production-readiness reviews before service launches.
  • Experience leading on-call operations, including defining rotations, escalation policies, and improving alert quality.

About the company

Chainlink is the industry-standard oracle platform bringing the capital markets onchain and powering the majority of decentralized finance (DeFi). The Chainlink stack provides the essential data, interoperability, compliance, and privacy standards needed to power advanced blockchain use cases for institutional tokenized assets, lending, payments, stablecoins, and more. Since inventing decentralized oracle networks, Chainlink has enabled tens of trillions in transaction value and now secures the vast majority of DeFi.

Many of the world’s largest financial services institutions have also adopted Chainlink’s standards and infrastructure, including Swift, Euroclear, Mastercard, Fidelity International, UBS, S&P Dow Jones Indices, FTSE Russell, WisdomTree, ANZ, and top protocols such as Aave, Lido, GMX and many others. Chainlink leverages a novel fee model where offchain and onchain revenue from enterprise adoption is converted to LINK tokens and stored in a strategic Chainlink Reserve. Learn more at chain.link.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

3:26 min

Automating programmable logic using smart contracts and JavaScript tools

Ryan Arndt · WWC 2023

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · WWC 2025

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · WWC 2022

12:08 min

Comparing Keptn orchestration capabilities against alternative software operators

Thomas Schütz · LIVE

Videos

See all

Related articles

See all