AI DevOps Systems Administrator

Column Technical Services
Scottsdale, AZ, United States
2 months ago
Apply on indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$130,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Continuous Integration Data Security Linux DevOps Linux System Administration Machine Learning Network Interface Ansible Azure Machine Learning Server Virtualization AI Infrastructure
+11 more
Data Logging Infrastructure as Code (IaC) Containerization Storage Technologies Information Technology Deployment Automation Machine Learning Operations Virtual Agents Terraform Docker Server Operating Systems & Platforms

Job description

Column Technical Services is seeking a highly skilled AI DevOps Systems Administrator to architect, support, and evolve the infrastructure powering our cutting-edge Artificial Intelligence and Machine Learning initiatives in a secure, classified environment based in Scottsdale, AZ. In this role, you’ll be at the forefront of innovation, driving reliable model development and deployment by optimizing pipelines, maximizing compute performance, and ensuring robust scalability and security across platforms. This is a unique opportunity to work with advanced technologies while making a direct impact on mission-critical systems. If you’re passionate about AI infrastructure, thrive in high-performance environments, and are ready to take on meaningful, complex challenges, we encourage you to apply.

Sponsorship is not available for this role. Candidates must currently reside in or near Scottsdale, Arizona. Applicants must hold an active TS/SCI with Polygraph clearance.

In this position, you will work closely with data scientists and machine learning engineers to enable seamless transitions from experimentation to production.

Core Responsibilities

  • Architect, deploy, and support scalable environments for AI/ML training and inference workloads
  • Build and maintain automated CI/CD workflows for machine learning models and AI-driven applications
  • Administer and fine-tune Linux-based systems across physical and virtual infrastructures
  • Implement and manage containerized environments using tools such as Docker and Kubernetes to support scalable ML services
  • Utilize Infrastructure as Code (IaC) solutions (e.g., Terraform, Ansible) to automate provisioning, configuration, and system management
  • Optimize allocation and usage of GPU resources for compute-intensive workloads
  • Establish monitoring, logging, and alerting frameworks to ensure system health, availability, and performance
  • Partner with engineering teams to troubleshoot issues, improve workflows, and meet infrastructure requirements

Additional Responsibilities

As a senior-level contributor, you will serve as a key technical point of contact, supporting users and participating in system design and evolution efforts to align with emerging technologies. You will:

  • Install, configure, and maintain software and system components
  • Diagnose and resolve technical issues, including access control and permissions
  • Provide guidance and training to users on system functionality
  • Manage daily operations of server environments across both physical and virtual platforms
  • Configure, maintain, and troubleshoot hardware, operating systems, and network interfaces
  • Investigate and resolve system alerts, ensuring continuity of services
  • Develop scripts to streamline and automate repetitive operational tasks
  • Collaborate directly with stakeholders to identify, isolate, and resolve system-related issues impacting broader services

Requirements

Do you have experience in Terraform?, Do you have a Bachelor’s degree?, * A collaborative mindset with a strong commitment to team success and shared outcomes

  • Solid understanding of how systems, servers, and services interconnect within a broader IT ecosystem
  • Advanced expertise in supporting both physical and virtual server environments
  • Deep knowledge of access controls, permissions, and security practices to ensure appropriate and secure data access
  • A proactive approach to identifying opportunities to leverage AI for operational efficiency, continuous improvement, and innovation

What You’ll Experience

  • Work with advanced and often highly classified technologies
  • Be part of a forward-thinking team focused on innovation and exploration
  • Continuous learning opportunities aligned with emerging advancements, * Minimum of 8 years of relevant experience OR a Master’s degree with 6+ years of experience
  • Bachelor’s degree in Computer Science, a related discipline, or equivalent experience
  • Deep expertise in server-based operating systems
  • Strong proficiency in Linux environments, containerization, and AI/ML infrastructure
  • Proven ability to serve as a subject matter expert and mentor team members
  • Advanced troubleshooting skills across operating systems, networking, and storage technologies
  • Hands-on experience building, deploying, and maintaining enterprise-scale server environments
  • Exposure to or experience working with AI/ML workloads is highly desirable
  • Willingness to travel occasionally, * How have you tuned Linux systems to support high-performance AI/ML workloads?
  • What’s your experience scaling containerized ML workloads with Docker and Kubernetes?
  • How have you automated ML model deployment using CI/CD pipelines?
  • Do you have an active TS/SCI with Polygraph clearance?
  • Will you now or in the future require sponsorship to work in the USA?

Benefits & conditions

Pulled from the full job description

  • Health insurance
  • 401(k) matching
  • Paid time off
  • Vision insurance
  • Dental insurance
  • Flexible spending account
  • Life insurance, * 401(k) matching
  • Dental insurance
  • Flexible spending account
  • Health insurance
  • Life insurance
  • Paid time off
  • Vision insurance

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

4:15 min

Introduction to artificial intelligence driven development

Natalie Pistunovich · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:34 min

Transitioning from traditional software development to artificial intelligence consulting

Patrick Schnell Patrick Schnell · Coffee With Developers

Videos

See all

Related articles

See all