> Markdown version of [/jobs/ext/2714758-site-reliability-engineer-ai-accelerator-infrastructure](https://www.wearedevelopers.com/jobs/ext/2714758-site-reliability-engineer-ai-accelerator-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - AI Accelerator Infrastructure - **Company:** d Matrix - **Location:** Santa Clara, United States - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Amazon Web Services, Application Layers, Build Automation, Microsoft Azure, Bash Shell, Cloud Computing, Configuration Management, Continuous Integration, Linux, DevOps, Ethernet, InfiniBand, Job Scheduling, Python (Programming Language), Package Management Systems, Ansible, Prometheus, Runbook, Software Deployment, Virtual Local Area Networks, Datadog, Data Storage Management, Google Cloud, Cloud Platform System, Grafana, Reliability of Systems, Kubernetes, Information Technology, Bare Metal, Slurm, Terraform - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/site-reliability-engineer-ai-accelerator-infrastructure-contract-d-matrix-8752446 ## About the Role * Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience); 5+ years in SRE, infrastructure engineering, or systems administration. * Strong Linux systems knowledge: networking, storage, systemd, package management, kernel parameters, and performance diagnostics. * Hands-on experience with colocation or on-premises server infrastructure - physical hardware, rack networking, and bare-metal provisioning. * IaC experience with Terraform and/or Ansible - writing and maintaining production configurations, not just running existing playbooks. * Kubernetes operational experience: cluster troubleshooting, workload management, storage, and networking. * Prometheus + Grafana or DataDog: building dashboards, writing alert rules, and understanding signal quality. * Python and/or Bash scripting: production-quality automation, not just one-off scripts. * Incident response experience: structured triage, RCA production, and follow-through on action items. * Comfort operating in fast-moving startups: you own your systems, document what you build, and iterate without waiting for perfect requirements. Strongly Preferred * Experience operating customer-facing infrastructure or platform services with external reliability expectations. * Cloud infrastructure operations across AWS, Azure, or GCP-including hybrid environments spanning cloud and on-prem. * HPC job scheduler experience: Slurm, LSF, or equivalent - operations and troubleshooting. * Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink - configuration and troubleshooting. * Go programming for SRE tooling - health-check services, exporters, or auto-remediation agents. * Experience with large-scale infrastructure automation: host lifecycle management, fleet auto-healing, or AIOps-driven operations. ## Description The Role: 6-month contract with potential conversion to full-time You will be a core member of the d-Matrix SRE team, responsible for the reliability, automation, and observability of the infrastructure that the company runs on. You will work across colocation, on-premises lab environments, and cloud platforms-and you will own your systems end-to-end, from initial provisioning through live incident response. You will work closely with the senior DevOps lead team, whose CI/CD pipelines and automation layer depend on the infrastructure you operate. You will also support customer-facing environments where d-Matrix partners collaborate on hardware and software deployments. What You Will Do Infrastructure Operations * Own reliability and availability of assigned infrastructure domains: colo server fleets, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services. * Perform hands-on infrastructure work: server provisioning, OS configuration, network setup, storage management, and hardware troubleshooting from bare metal up. * Support and operate high-speed interconnect environments - InfiniBand, RoCE, or high-speed Ethernet - in lab and colo settings. * Conduct capacity planning and hardware lifecycle management for assigned infrastructure domains. Automation & Infrastructure as Code * Own IaC and configuration management (Terraform, Ansible) for your infrastructure domains - all provisioning and changes through code, not manual steps. * Build automation to eliminate toil: host lifecycle management, fleet health checks, auto-remediation workflows, and self-service tooling for engineering teams. * Develop networking automations for cluster interconnects, VLAN management, and lab network configurations. * Contribute to shared IaC modules and automation libraries used across the SRE team. Observability & Incident Response * Design and maintain monitoring dashboards, alerting, and SLIs (Prometheus/Grafana, DataDog) for your infrastructure domains-ensuring signal quality and actionable alerts. * Participate in on-call rotation; triage and resolve incidents from bare metal to application layer, distinguishing infrastructure faults from software or hardware product issues. * Produce high-quality RCA reports for P0/P1 incidents with root cause analysis and tracked action items. * Detect performance issues, recommend solutions, and implement fixes that permanently improve system reliability. Customer & Platform Services * Support and operate platform services used by external customers for hardware and software deployment collaboration with d-Matrix. * Ensure QoS and uptime commitments for customer-facing environments; escalate reliability risks proactively. * Document platform configurations, access procedures, and operational runbooks for customer environments. Documentation & Collaboration * Maintain high-quality runbooks, architecture diagrams, and troubleshooting guides-documentation is part of the job, not an afterthought. * Partner with the DevOps team to ensure infrastructure reliability supports CI/CD pipeline performance and developer experience. * Serve as a technical resource for engineering teams-sharing operational knowledge and raising infrastructure risks early. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)