Site Reliability Engineer

PathAI, Inc.
Boston, MA, United States
27 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$165,750.0 - $224,450.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon S3 Intelligent Platform Management Interface Cloud Computing Data Centers Python (Programming Language) Kernel-Based Virtual Machine Network Layer Machine Learning Network Planning and Design Reliability Engineering Ansible
+13 more
Prometheus Systems Integration Datadog Grafana HybridCloud Juniper Cloudformation Containerization Kubernetes Infrastructure Automation Frameworks Information Technology Terraform Golang

Job description

We’re looking for a skilled senior/staff level Site Reliability Engineer focused on designing, building, and operating our hybrid cloud/on-prem environment. What You’ll Do

If you’re the right candidate, you’ll be exercising all the skills you have and building new ones along the way:

  • Advancing the state of our operations by implementing SRE best practices - focusing on users, monitoring, and automation.
  • Engineering infrastructure patterns for cloud environments in Amazon Web Services - building in security, reliability and scalability.
  • Designing, building, and operating our data center to support our rapidly growing Machine Learning team.
  • Integrating on-premises datacenter environments with existing cloud infrastructure to create a seamless hybrid cloud environment.
  • Improving the reliability and resilience of our infrastructure through root-cause analysis and reviewing gaps in designs, and implementations of our infrastructure.
  • Participating in platform on-call rotations and assisting with urgent incident response.

Requirements

Our employees’ skills come in all shapes and sizes, but to be successful in this role with us, you’ll at least need:

  • 5+ years of relevant experience.
  • Automation: You work hard to eliminate toil by automating everything through scripting, configuration management tools (Ansible), and code (Python/GoLang).
  • You’ve built monitoring infrastructure with modern observability tools (Datadog/Grafana/Prometheus).
  • You’ve worked with infrastructure as code (Terraform/Cloudformation).
  • You’ve administered physical hardware stacks in production settings (iDRAC/IPMI/Nvidia UFM/Juniper Systems).
  • You’re opinionated on storage solutions and how they can be optimized for high performance workloads (Quobyte/S3/FSx/EFS).
  • Familiarity with modern network designs and comfort operating across network layers.
  • Some experience and opinions on virtualization, containerization, or container orchestration platforms. (EKS/ClusterAPI/KVM).
  • Operations experience: You’ve managed critical production infrastructure and are familiar with incident response, scaling, and rapid growth related challenges.
  • A bachelor’s degree in Computer Science or equivalent experience.
  • An insatiable intellectual curiosity and the ability to learn quickly in a complex space.
  • Travel: Willingness to travel up to 25% of the time.

About the company

PathAI is on a mission to improve patient outcomes with AI-powered pathology. We are transforming traditional pathology methods into powerful, new technologies. These innovations in pathology can help accelerate drug development, improve confidence in the accuracy of diagnosis, and get life-saving therapies to patients more quickly. At PathAI, you’ll work with a diverse and talented team of people, who are dedicated to solving complex problems and making a huge impact.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:19 min

Applying code assistant capabilities to infrastructure and cloud operations

Ryan J Salva · Coffee With Developers

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

Videos

See all

Related articles

See all