Machine Learning & AI Infrastructure Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+9 more
Job description
Kforce has a client in Austin, TX that is seeking a Machine Learning & AI Infrastructure Engineer. This is not a traditional AI Engineer or Data Scientist role. The hiring team is specifically seeking a unique blend of: HPC Administrator + Kubernetes Administrator + AI Infrastructure Operations Engineer. Candidates who have owned, operated, supported, and troubleshot production AI or HPC environments will be the strongest fit. Experience administering and maintaining systems is significantly more important than architecture-only experience., * Administer and support AI and HPC cluster environments
- Manage day-to-day operations of large-scale compute infrastructure
- Deploy, maintain, and troubleshoot Kubernetes-based platforms
- Ensure reliability, performance, and scalability across compute, storage, and networking environments
- Support AI model training and inference infrastructure
- Automate operational processes through scripting and tooling
- Partner with engineering teams and customers to optimize platform performance
- Troubleshoot complex infrastructure, networking, storage, and containerization issues
- Support both internal platforms and customer-facing environments, * AI Infrastructure
- Machine Learning Platforms
- HPC Operations
- Research Computing
- Biotechnology
- Academic Medical Centers
- Digital Biology
- Financial Services AI Platforms
- Automotive AI Initiatives
- Large-Scale Data Science Environments
Requirements
- Strong experience administering High Performance Computing (HPC) environments
- Experience with AI cluster administration and infrastructure operations
- Hands-on Kubernetes administration experience in on-premises environments
- Experience provisioning and managing PV/PVC storage through Kubernetes CSI drivers
- Strong Linux administration skills, specifically Ubuntu
- Scripting experience with Bash and/or Python
- Proven troubleshooting and operational support experience
- Ability to manage and maintain production infrastructure environments
Experience with one or more of the following:
- Dell PowerScale/Isilon
- VAST Storage
- NetApp ONTAP
- DDN IntelliFlash
- DDN Exascaler
- Lustre Parallel File Systems, * NVIDIA ecosystem experience
- NVIDIA Base Command Manager (BCM)
- Bright Cluster Manager
- MLOps platform exposure
- Containerization technologies and orchestration platforms
- High-performance networking experience
- RDMA technologies
- InfiniBand networking
- NVIDIA UFM
- Parallel file system administration
- Storage Technologies (highly desired)
Benefits & conditions
Why Consider This Opportunity?
- 100% Remote Environment
- Exposure to cutting-edge AI, GenAI, and HPC technologies
- Flat organizational structure with minimal bureaucracy
- Direct impact on strategic technology initiatives
- Opportunity to work on platforms that support healthcare, research, drug discovery, and other meaningful AI-driven innovations
- High visibility and collaboration with industry-leading technical teams
Compensation & Benefits:
- Base Salary: $175,000 - $200,000+
- Annual Bonus: Typically 10%-15%
- Medical, Dental, and Vision Coverage
- 401(k)
- Additional performance-based incentives, The pay range is the lowest to highest compensation we reasonably in good faith believe we would pay at posting for this role. We may ultimately pay more or less than this range. Employee pay is based on factors like relevant education, qualifications, certifications, experience, skills, seniority, location, performance, union contract and business needs. This range may be modified in the future.
We offer comprehensive benefits including medical/dental/vision insurance, HSA, FSA, 401(k), and life, disability & ADD insurance to eligible employees. Salaried personnel receive paid time off. Hourly employees are not eligible for paid time off unless required by law. Hourly employees on a Service Contract Act project are eligible for paid sick leave.
Note: Pay is not considered compensation until it is earned, vested and determinable. The amount and availability of any compensation remains in Kforce’s sole discretion unless and until paid and may be modified in its discretion consistent with the law.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
Stephan Gillich - Bringing AI Everywhere
What Industries Outside of AI Are Hiring The Most AI Experts?