> Markdown version of [/jobs/ext/2873874-senior-staff-site-reliability-engineer-data-center](https://www.wearedevelopers.com/jobs/ext/2873874-senior-staff-site-reliability-engineer-data-center). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior/Staff Site Reliability Engineer - Data Center - **Company:** PathAI, Inc. - **Location:** Boston, MA, United States (Remote available) - **Experience:** Expert - **Salary:** $146,250.0 - $225,000.0 - **Contract:** Permanent contract - **Skills:** Amazon S3, Intelligent Platform Management Interface, Cloud Computing, Computer Engineering, Data Centers, Kernel-Based Virtual Machine, Network Layer, Machine Learning, Reliability Engineering, Ansible, Prometheus, Software Engineering, Datadog, Grafana, HybridCloud, Juniper, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology - **Published:** September 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=df8b10bd75082f49 ## About the Role * You have a BS in Computer Science, Computer Engineering, Electrical Engineering, Software Engineering or closely related technical field. * You have 8 years experience working in physical hardware/facilities, networking, automation or other relevant areas. * You have demonstrated experience with modern datacenter network designs and comfort operating across network layers. * You've administered physical hardware stacks in production settings (iDRAC/IPMI/Nvidia UFM/Juniper Systems). * You have demonstrated experience and opinions on virtualization, containerization, or container orchestration platforms. (EKS-Anywhere/ClusterAPI/KVM). * You have strong expertise in storage solutions and optimizing them for high-performance workloads (e.g., Quobyte, S3, FSx, EFS). * You are highly proficient with automation tools; you eliminate toil by automating everything through scripting, configuration management tools (Ansible/RedFish). * You've built monitoring infrastructure with modern observability tools (Datadog/Grafana/Prometheus). * You have a proven operational background managing critical production systems, with extensive experience in incident response, infrastructure scaling, and navigating high-growth challenges. * You have the ability to travel to onsite Datacenter location(s) as needed. Preferred: * You have outstanding interpersonal, verbal, and written communication and influencing skills: have built and cultivated important relationships both inside and outside of the organization and externally; have proven abilities to influence internal partners and stakeholders, thought leaders, national advocacy organizations, national standard-setting bodies, and other relevant external parties. * You have strong analytical and critical thinking skills with attention to detail; you have the ability to manage multiple projects and drive results in a fast-paced environment; you have a collaborative mindset with demonstrated leadership capabilities.. ## Description * You will advance the state of our operations by implementing SRE best practices - focusing on users, monitoring, and automation. * You will design, build and operate our data center to support our rapidly growing Machine Learning team. * You will build highly-secure on-premises environments handling NIST/ISO standards. * You will integrate on-premises datacenter environments with existing cloud infrastructure to create a seamless hybrid cloud environment. * You will improve the reliability and resilience of our infrastructure through root-cause analysis and reviewing gaps in designs, and implementations of our infrastructure. * You will participate in platform on-call rotations and assist with urgent incident response. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)