Senior/Staff Site Reliability Engineer - Data Center

PathAI, Inc.
Boston, MA, United States
8 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$146,250.0 - $225,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon S3 Intelligent Platform Management Interface Cloud Computing Computer Engineering Data Centers Kernel-Based Virtual Machine Network Layer Machine Learning Reliability Engineering Ansible Prometheus Software Engineering
+8 more
Datadog Grafana HybridCloud Juniper Containerization Kubernetes Infrastructure Automation Frameworks Information Technology

Job description

  • You will advance the state of our operations by implementing SRE best practices - focusing on users, monitoring, and automation.
  • You will design, build and operate our data center to support our rapidly growing Machine Learning team.
  • You will build highly-secure on-premises environments handling NIST/ISO standards.
  • You will integrate on-premises datacenter environments with existing cloud infrastructure to create a seamless hybrid cloud environment.
  • You will improve the reliability and resilience of our infrastructure through root-cause analysis and reviewing gaps in designs, and implementations of our infrastructure.
  • You will participate in platform on-call rotations and assist with urgent incident response.

Requirements

  • You have a BS in Computer Science, Computer Engineering, Electrical Engineering, Software Engineering or closely related technical field.
  • You have 8 years experience working in physical hardware/facilities, networking, automation or other relevant areas.
  • You have demonstrated experience with modern datacenter network designs and comfort operating across network layers.
  • You’ve administered physical hardware stacks in production settings (iDRAC/IPMI/Nvidia UFM/Juniper Systems).
  • You have demonstrated experience and opinions on virtualization, containerization, or container orchestration platforms. (EKS-Anywhere/ClusterAPI/KVM).
  • You have strong expertise in storage solutions and optimizing them for high-performance workloads (e.g., Quobyte, S3, FSx, EFS).
  • You are highly proficient with automation tools; you eliminate toil by automating everything through scripting, configuration management tools (Ansible/RedFish).
  • You’ve built monitoring infrastructure with modern observability tools (Datadog/Grafana/Prometheus).
  • You have a proven operational background managing critical production systems, with extensive experience in incident response, infrastructure scaling, and navigating high-growth challenges.
  • You have the ability to travel to onsite Datacenter location(s) as needed.

Preferred:

  • You have outstanding interpersonal, verbal, and written communication and influencing skills: have built and cultivated important relationships both inside and outside of the organization and externally; have proven abilities to influence internal partners and stakeholders, thought leaders, national advocacy organizations, national standard-setting bodies, and other relevant external parties.

  • You have strong analytical and critical thinking skills with attention to detail; you have the ability to manage multiple projects and drive results in a fast-paced environment; you have a collaborative mindset with demonstrated leadership capabilities..

About the company

PathAI’s mission is to improve patient outcomes with AI-powered pathology.

PathAI is transforming traditional pathology methods into powerful, new technologies. These innovations in pathology can help accelerate drug development, improve confidence in the accuracy of diagnosis, and get life-saving therapies to patients more quickly. At PathAI, you’ll work with a diverse and talented team of people, who are dedicated to solving complex problems and making a huge impact.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · World Congress 2025

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all