Data Center Operations Coordinator

Together Ai
San Francisco, CA, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$150,000.0 - $200,000.0
Working hours
Regular working hours
Job source

Tech stack

Microsoft Excel Computer-Aided Design Artificial Intelligence JIRA Data Centers Issue Tracking Systems Networking Hardware Inventory Management Software AI Infrastructure Comptia Server+ Information Technology Servicenow

Job description

We’re looking for a detail-oriented Data Center Operations professional to manage and track all break/fix activities across multiple data center locations. This role acts as the central point of coordination for hardware incidents, vendor dispatches, ticket management, asset tracking, and operational reporting to ensure maximum uptime and fast issue resolution., * Track and manage all break/fix incidents across multiple data centers

  • Monitor ticket queues and ensure SLA compliance for incident response and resolution
  • Coordinate with on-site technicians, remote hands teams, vendors, and engineering groups
  • Maintain accurate records of failed hardware, replacements, RMAs, and repair status
  • Escalate critical outages and recurring infrastructure issues to leadership and engineering teams
  • Schedule and oversee maintenance windows and emergency repair activities
  • Provide daily/weekly operational status reports and incident summaries
  • Ensure all work follows data center operational procedures and change management policies
  • Identify trends in hardware failures and recommend process improvements

Requirements

  • Experience working in data center operations, IT infrastructure, or hardware support
  • Strong understanding of server, storage, and networking hardware
  • Experience with ticketing systems such as ServiceNow, Jira, or Remedy
  • Ability to manage multiple priorities across several sites simultaneously
  • Excellent communication and organizational skills
  • Familiarity with SLA management and incident escalation processes
  • Proficiency with Excel, reporting dashboards, and inventory tracking tools, * Experience supporting enterprise or hyperscale data centers
  • Knowledge of remote hands operations and vendor management
  • Understanding of ITIL processes and change management
  • CompTIA Server+, Network+, or similar certifications

Benefits & conditions

We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $150,000-200,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

About the company

Together AI is a research-driven AI infrastructure company on a mission to dramatically lower the cost of modern AI by co-designing software, hardware, algorithms, and models. We believe open and transparent AI systems create the best outcomes for society - and we’re building the physical and computational foundation to make that real. Our team has been behind landmark advances including FlashAttention, Hyena, FlexGen, and RedPajama., Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:08 min

Intersecting neurodivergent support and technical operations consulting

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

56 sec

Integrating automated approval workflows into the portal

Markus Eisele Markus Eisele · WWC 2025

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

4:03 min

Managing massive power consumption scaling in AI data centers

Stephan Gillich Stephan Gillich +3 · WWC 2024

2:27 min

Establishing a simulated technical environment for the workflow demo

Tobias Dunn-Krahn · LIVE

Videos

See all

Related articles

See all