Site Reliability Engineer

Bridgesource Solutions
Seattle, WA, United States
27 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Systems Engineering Microsoft Azure Cloud Computing Continuous Integration Linux Distributed Systems InfiniBand Python (Programming Language) Remote Direct Memory Access Reliability Engineering
+8 more
Ansible Software Engineering AI Infrastructure Graphics Processing Unit (GPU) Grafana Kubernetes Terraform Programming Languages

Requirements

  • Experience in Site Reliability Engineering, Platform Engineering, Systems Engineering, Software Engineering, or Infrastructure Engineering.
  • Strong knowledge of Linux, Kubernetes, cloud platforms (AWS, Azure, or GCP), networking, and distributed systems.
  • Proficiency in Python, Go, or a similar programming language for automation and tooling.
  • Experience with Infrastructure as Code (Terraform, Ansible, etc.), CI/CD, monitoring, and observability tools.
  • Hands-on experience supporting production environments, troubleshooting complex issues, and performing root cause analysis (RCA).
  • Understanding of SLIs, SLOs, incident management, and on-call best practices.
  • Experience with AI infrastructure, GPU environments, HPC, InfiniBand, or RDMA is a plus.
  • Strong communication skills with the ability to collaborate across engineering teams. Senior and Principal-level candidates should also demonstrate technical leadership, architecture/design experience, and the ability to mentor and guide other engineers.
  • Nice to have: experience building HPC clusters for 1,000+ GPUs to power AI/ML workloads with Hyperscale clients

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all