> Markdown version of [/jobs/ext/2716885-data-center-operations-lead-partner-site-operations](https://www.wearedevelopers.com/jobs/ext/2716885-data-center-operations-lead-partner-site-operations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Center Operations Lead - Partner Site Operations - **Company:** Anthropic's Mission - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $320,000.0 - **Contract:** Permanent contract - **Skills:** Data Centers, Information Technology Operations, Break Fix, Information Technology - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/data-center-operations-lead-partner-site-operations-anthropic-3-9617346 ## About the Role * Have 8+ years of experience in data center operations (hardware, IT infrastructure, or critical facilities) as a manager, technical lead or related role, including accountability for production availability. * Have managed vendors, MSPs, or contract workforces to measurable outcomes: SOWs, SLAs, operational reviews, and corrective action. * Carry hands-on technical depth in server, network, and rack-level infrastructure, enough to independently verify vendor claims and audit quality. * Have built or substantially improved operational processes, not just run them. * Have served in an incident command or lead-responder role and communicate clearly under ambiguity. * Can support non-standard hours, including an on-call rotation and availability during deployment surges and maintenance windows. * Bachelor's degree in relevant domain or equivalent practical experience. Strong candidates may also have * Experience with third-party colocation providers or partner-operated sites, delivering IT operations outcomes inside a facility someone else runs. * Experience standing up operations at a new site or data hall, from commissioning handoff through first deployment. * Experience with GPU/accelerator or high-density liquid-cooled infrastructure. * Familiarity with multi-vendor sites where facilities and IT operations are performed by different partners. * Experience leading projects from initiation to completion across teams you didn't own. * Background in incident management frameworks, contract/SLA design, or EHS programs., Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. ## Description Anthropic's Data Center Operations (DCO) team ensures compute fleet availability through hardware and IT operations. At our partner-operated sites, this role manages the interface between Anthropic and the strategic site operations partner performing day-to-day data hall work. As the site lead, you own site outcomes for your assigned sites including: deployment velocity, availability, and incident response. Rather than managing operations staff directly, you provide tactical direction, set priorities, and define the standards for the vendor's on-site teams, paired with performance oversight and ongoing operational assessment to ensure all operational commitments are met. You will define the operational processes, quality gates, and governance rhythms for partner-operated sites. Expect to build the playbook as much as you run it, not just at a site level, but defining and developing program improvements fleet-wide. What you'll own * Operational outcomes. Own site availability, deployment milestones, and repair turnaround, verified with independent data rather than vendor self-reporting. * Vendor direction. Set daily and weekly priorities and lead the operating cadence, including standups and business reviews. * Process definition. Author and improve procedures for deployment, break-fix, change management, security, and EHS compliance. Analyze operational trends and standardize lessons across the program. * Performance management. Track vendor performance against SLAs and staffing commitments, driving corrective actions when necessary. * Incident response and on-call. Participate in the incident escalation on-call rotation. When designated Anthropic Incident Commander for a site-specific incident, direct vendor response, own communications, and close out post-incident actions. * Internal interface. Translate engineering requirements into vendor direction and communicate site constraints and risks to leadership. Representative work * Leading weekly operations reviews and scorecards with vendor site leads. * Directing deployment surges to meet first-compute-online milestones. * Analyzing failure patterns to identify root causes and driving fixes with owners. * Creating break-fix ownership matrices and training vendor teams. * Serving as Incident Commander for facility events and producing post-mortems. * Establishing operational readiness for new data halls, including spares and security. * Identifying process gaps and codifying improvements as program standards. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Data Mining Accessibility](https://www.wearedevelopers.com/videos/802-data-mining-accessibility) - [The Sustainability Race: AI's Promises, Pitfalls and Potential](https://www.wearedevelopers.com/videos/100155-the-sustainability-race-ai-s-promises-pitfalls-and-potential) - [Enabling intelligent logistics automation: home-grown Industrial IoT platform at Austrian Post](https://www.wearedevelopers.com/videos/2018-enabling-intelligent-logistics-automation-home-grown-industrial-iot-platform-at-austrian-post) - [5 Years in Cloud Native: The Good, the Bad, and the Bill](https://www.wearedevelopers.com/videos/100111-5-years-in-cloud-native-the-good-the-bad-and-the-bill) - [Beyond the 9–5: Designing Work Around Humans](https://www.wearedevelopers.com/videos/1321-beyond-the-9-5-designing-work-around-humans) ## Related Articles - [A Guide to Green Tech and Green IT Careers](https://www.wearedevelopers.com/magazine/374-a-guide-to-green-tech-and-green-it-careers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Should Tech Managers Be Developers First? Pros and Cons](https://www.wearedevelopers.com/magazine/327-should-tech-managers-be-developers-first-pros-and-cons) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025)