Data Center Site Manager

Lambda Inc.
Chicago, IL, United States
2 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Secure Shell (SSH) Airflow Computer Engineering Data Centers Dynamic Host Configuration Protocol Data Center Infrastructure Management (CIM) Distributed Data Store Document-Oriented Databases Domain Name System (DNS) Ethernet Hypertext Transfer Protocols (HTTP) Network Topologies
+17 more
InfiniBand IP Routing Network Layer Linux System Administration Network Architecture Open Shortest Path First (OSPF) Simple Network Management Protocols Software Engineering Syslog TCP/IP Virtual Local Area Networks Network Switches Remote Desktop Protocol (RDP) Transport Layer Security File Transfer Protocol (FTP) Deployment Automation NetBIOS

Job description

  • Maintain high availability, reliability, and security in the data center environment
  • Ensure new server, storage and network infrastructure is properly racked, labeled, cabled, and configured
  • Troubleshoot hardware and software issues in some of the world’s most advanced systems
  • Document data center layout and network topology in DCIM software
  • Work with supply chain & manufacturing teams to ensure timely deployment of systems and project plans for large-scale deployments
  • Assess current and future state data center requirements based on growth plans and technology trends
  • Manage a parts depot inventory and track equipment through the delivery-store-stage-deploy-handoff process in each of our data centers
  • Create installation standards and documentation for placement, labeling, and cabling to drive consistency and discoverability across all data centers
  • Oversee deployments and day-to-day operations of the data center
  • Maintain uptime for assets and infrastructure, and ensure customer SLAs are met
  • Participate in technical discussions and provide expertise on data center integration and deployment strategies
  • Understand power/cooling requirements as well as cabling needs required within data center space to support high performance infrastructures.
  • Work closely with cross-functional teams, including Hardware Engineering, Software Engineering, Supply Chain, Customer Experience and Sales, to align data center solutions with business goals
  • Ensure the data center complies with Lambda’s standards and policies

Requirements

  • Have 5+ years experience with critical infrastructure systems supporting data centers, such as power distribution, air flow management, environmental monitoring, capacity planning, DCIM software, structured cabling, and cable management
  • Have basic understanding of Linux administration
  • Have experience in setting up networking appliances (Ethernet and InfiniBand) across multiple data center locations
  • Are someone who pays attention to detail and has the ability to follow instructions
  • Are action-oriented and have a strong willingness to learn
  • Have a desire to mentor other team members and help them reach their full potential

Nice to Have

  • Experience with troubleshooting and theoretical knowledge the following network layers, technologies, and system protocols: TCP/IP, OSPF, SNMP, SSL, HTTP, FTP, SSH, Syslog, DHCP, DNS, RDP, NETBIOS, IP routing, Ethernet, switched Ethernet, 802.11x, NFS, and VLANs
  • Experience with working in large-scale distributed data center environments
  • Experience working with auditors to meet all compliance requirements (ISO/SOC)
  • Experience Supermicro & Nvidia hardware
  • Previous data center team management experience

Benefits & conditions

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast
  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
  • Our values are publicly available: https://lambda.ai/careers
  • We offer generous cash & equity compensation
  • Health, dental, and vision coverage for you and your dependents
  • Wellness and commuter stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible paid time off plan that we all actually use

About the company

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda’s mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you’d like to build the world’s best AI cloud, join us.

*Note: This position requires presence in our Elk Grove Data Center 5 days per week.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

7:31 min

Essential foundational skills and concepts for infrastructure roles

Megha Kadur · LIVE

46 sec

Automating telemetry collection through robust Telegraf deployment

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

1:44 min

Career transition into cloud native and data management

Michael Cade · LIVE

2:30 min

Discovering and instrumenting services using systemd process enumeration

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · World Congress 2022

Videos

See all

Related articles

See all