Data Center Lab Operations Engineer

CareerCircle
Reno, NV, United States
4 days ago
Apply on www.careercircle.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$72,800.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Systems Engineering Software Quality Data Centers Data Center Infrastructure Management (CIM) Software Debugging Linux Domain Name System (DNS) Firmware Networking Hardware IP Addressing Python (Programming Language)
+19 more
Linux System Administration Machine Learning Networking Basics Network File Systems Network Administration Ansible Software Engineering System Testing TCP/IP Network Switches Network Routing Graphics Processing Unit (GPU) Transport Layer Security High Performance Computing Infrastructure Automation Frameworks Operational Systems Slurm Hardware Infrastructure Network Server

Job description

We are seeking a Data Center Lab Operations Engineer to support a high-performance compute environment used for developing and validating next-generation technologies. This role will partner closely with engineering, QA, and operations teams to maintain critical infrastructure, troubleshoot complex hardware and software issues, and drive continuous improvements across compute and testing operations., * Support and maintain a large-scale compute farm consisting of builders, testers, packagers, and supporting infrastructure.

  • Collaborate with hardware, software, QA, and systems engineering teams to deploy, troubleshoot, and optimize engineering environments.
  • Monitor system health, availability, and performance while ensuring SLA targets are achieved.
  • Lead troubleshooting and recovery efforts for infrastructure, hardware, and software incidents.
  • Assist with server deployments, hardware installations, upgrades, and maintenance activities.
  • Develop and maintain SOPs, documentation, and operational processes.
  • Collect and analyze infrastructure metrics to drive improvements in reliability and efficiency.
  • Implement automation and process improvements using scripting and infrastructure management tools., Use of Artificial Intelligence (AI): We may use Artificial Intelligence (AI) to support parts of our hiring process, including sourcing, screening, and evaluating candidates. AI helps assess applications and qualifications, but final decisions are made by our hiring team. By applying, you acknowledge and agree that your application may be reviewed using AI tools. Related Jobs Data Center Technician TEKsystems

Sparks, NV*On-Site

Operations IP Addressing Network Routing Firmware Updates Business Valuation Category 6 Cabling Category 5 Cabling Full Stack Development Network Administration Artificial Intelligence Business Transformation Submittals (Construction) Troubleshooting (Problem Solving) +0 Data Center Engineer TEKsystems

Sparks, NV*On-Site

Linux Debugging Operations System Recovery Safety Assurance Hardware Support Quality Assurance Business Valuation Mechanical Aptitude Software Engineering Linux Administration Full Stack Development Artificial Intelligence Business Transformation Hardware Troubleshooting High Performance Computing Standard Operating Procedure Slurm (Batch Scheduling Software) +0

Google IT Automation with Python Data Center Technician TEKsystems

Sparks, NV*On-Site

Firmware Operations Leadership IP Addressing Problem Solving Desktop Support Network Switches Help Desk Support Business Valuation Category 6 Cabling Category 5 Cabling Networking Hardware Full Stack Development Data Center Operations Artificial Intelligence Business Transformation Hardware Troubleshooting Troubleshooting (Problem Solving) +0

Requirements

Operations Leadership Management Automation Data Centers Communication Problem Solving Asset Management Safety Assurance Machine Learning Influencing Skills Business Valuation Process Improvement Systems Engineering Mechanical Aptitude Network File Systems Full Stack Development Hardware Installations Artificial Intelligence Business Transformation Performance Engineering Hardware Troubleshooting Infrastructure Management Software Quality (SQA/SQC) High Performance Computing Python (Programming Language) Influencing Without Authority Troubleshooting (Problem Solving) Slurm (Batch Scheduling Software), * 5+ years of experience supporting data centers, engineering labs, compute environments, or similar infrastructure operations.

  • Strong understanding of Linux and Windows administration.
  • Experience troubleshooting hardware, servers, networking, and operating systems.
  • Knowledge of networking fundamentals including TCP/IP, DNS, NFS, SSL, and related protocols.
  • Experience with scripting and automation tools such as Python, Shell, or Ansible.
  • Familiarity with DCIM tools and asset management platforms.
  • Strong documentation, communication, and problem-solving skills.
  • Ability to work cross-functionally with both technical and non-technical stakeholders., * Experience supporting HPC or AI/ML compute clusters.
  • Familiarity with Slurm, BCM, or other cluster management platforms.
  • Experience working with GPUs, server hardware, PCBs, and system validation environments.
  • CCNA or similar infrastructure certification.
  • Knowledge of storage, networking, and dense data center architectures.
  • Experience with liquid cooling technologies.
  • Strong mechanical aptitude and comfort working with tools and hardware installations., Individual compensation offered for this position within this range will depend on many factors, including qualifications, skills, relevant experience, job knowledge, geographic location, internal equity, and other pertinent job-related factors.

About the company

We’re partners in transformation. We help clients activate ideas and solutions to take advantage of a new world of opportunity. We are a team of 80,000 strong, working with over 6,000 clients, including 80% of the Fortune 500, across North America, Europe and Asia. As an industry leader in Full-Stack Technology Services, Talent Services, and real-world application, we work with progressive leaders to drive change. That’s the power of true partnership. TEKsystems is an Allegis Group company., We’re a leading provider of business and technology services. We accelerate business transformation for our customers. Our expertise in strategy, design, execution and operations unlocks business value through a range of solutions. We’re a team of 80,000 strong, working with over 6,000 customers, including 80% of the Fortune 500 across North America, Europe and Asia, who partner with us for our scale, full-stack capabilities and speed. We’re strategic thinkers, hands-on collaborators, helping customers capitalize on change and master the momentum of technology. We’re building tomorrow by delivering business outcomes and making positive impacts in our global communities. TEKsystems and TEKsystems Global Services are Allegis Group companies. Learn more at TEKsystems.com.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careercircle.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · World Congress 2022

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all