System Administrator, Data Centers

Utilidata, Inc.
Ann Arbor, MI, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$120,000.0 - $150,000.0
Working hours
Regular working hours
Job source

Tech stack

Access Network Ubuntu (Operating System) CentOS Configuration Management Data Centers Linux Monitoring of Systems Inventory Management Software Linux System Administration Linux Servers Log Analysis Performance Tuning
+20 more
Red Hat Enterprise Linux Ansible Prometheus Security Information and Event Management AI Infrastructure High Performance Computing Grafana Software Troubleshooting Containerization Kubernetes Performance Monitor Hardware Infrastructure CIS Benchmarks Puppet Splunk Docker Elk Stack User Administration Vulnerability Analysis User Accounts

Job description

The Systems Administrator is responsible for the day-to-day operational support, maintenance, and reliability of Karman systems deployed in high-density data center environments. This role focuses on Linux systems administration, system health, uptime, and operational excellence. This is not an architecture or design role, but a hands-on position ensuring stable, secure, and optimized production systems. This position is based onsite at our company headquarters in Ann Arbor, Michigan, with flexibility for occasional remote work. Candidates will be expected to collaborate cross-functionally with remote teams based across the country., * Perform day-to-day Linux systems administration across production environments (RHEL, Ubuntu, CentOS or similar)

  • Install, configure, patch, upgrade, and maintain Linux servers and associated infrastructure
  • Monitor system performance, availability, and resource utilization; proactively address issues before impact
  • Troubleshoot and resolve operating system, application, and infrastructure-related incidents
  • Execute established deployment procedures and configuration standards for Karman systems
  • Support rack-based systems, including hardware checks and coordination around PDUs and PSUs as needed
  • Maintain system security through patching, access control management, and adherence to best practices
  • Perform log analysis, root cause analysis, and incident documentation
  • Contribute to documentation of operational runbooks, troubleshooting guides, and system procedures
  • Support Tailscale network access, user accounts, and permission policies
  • Support observability stack (Prometheus/Grafana) including dashboards, user access management, metrics collection, and alerting
  • Provide internal technical support to engineering teams, including general escalations and package install requests
  • Work with IT and security teams on vulnerability scanning, remediation timelines, and patch prioritization
  • Support change management processes, coordinating with IT on network, firewall, and access control changes affecting production environments
  • Participate in security reviews, audits, and tabletop exercises as needed
  • Coordinate with IT on asset lifecycle management, including provisioning, decommissioning, and inventory tracking for data center hardware
  • Coordinate with the SOC on security incident detection, triage, and response for Karman production systems
  • Participate in on-call rotation to support system uptime and rapid incident response

Requirements

Do you have experience in System performance monitoring?, * 5+ years of hands-on Linux systems administration experience in production environments

  • Strong experience with Linux installation, configuration management, patching, and performance tuning
  • Experience supporting systems in data center or high-availability environments
  • Solid understanding of system monitoring, log management, and troubleshooting methodologies
  • Experience working within established infrastructure standards and operational processes
  • Comfortable working in a fast-paced, operationally focused environment

Enhanced Qualifications (Nice to Have)

  • Experience with configuration management or automation tools (Ansible, Puppet, Chef, or similar)
  • Experience with monitoring/observability tools (Prometheus, Grafana, ELK stack, or similar)
  • Familiarity with containerized environments (Docker, Kubernetes)
  • Exposure to high-performance computing, AI infrastructure, or GPU-based systems
  • Familiarity with security frameworks such as NIST, CIS benchmarks, or SOC 2 controls.
  • Experience with SIEM platforms (Splunk, Sentinel, or similar)

Benefits & conditions

Pulled from the full job description

  • Health insurance
  • 401(k) matching
  • Paid time off
  • Vision insurance
  • Dental insurance
  • Stock options, Salary Range: $120,000 to $150,000 base compensation depending on experience and stock options. Salary will be commensurate with an individual’s skills, training, years of experience, and in line with internal compensation bands., * Creating a diverse and inclusive workplace that is welcoming, supportive, affirming and respectful
  • Empowering employees to solve problems and work together to make a difference
  • Providing mentorship and growth opportunities as part of a collaborative team
  • A flexible work environment with flexible paid time off
  • Competitive compensation and benefits, including health, dental, vision, and employer-match 401k

ZyLCcgPqOW

About the company

Utilidata is a fast-growing NVIDIA-backed edge AI company enabling greater visibility and control of power utilization in energy-intensive infrastructure, like the electric grid and data centers. Karman, the company’s distributed AI platform powered by a custom NVIDIA module, is transforming the way utility companies operate the grid edge and will enable data centers to unlock more compute for the same provisioned power.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all