Associate ML Infrastructure Engineer

OpenKyber LLC
United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Microsoft Windows Artificial Intelligence C++ (Programming Language) Ubuntu (Operating System) CentOS Software Quality Computer Programming Linux Python (Programming Language) Tensorflow Ruby
+10 more
Software Engineering Network Switches Data Processing Pytorch Large Language Models Kubernetes Slurm Machine Learning Operations Fedora Docker

Job description

About the Role: This role sits at the center of our team’s expansion and will have a direct impact on documentation, processes, and customer relationships. Work will vary from tool evaluation and integration to day-to-day hands-on lab and infrastructure support. In addition to core engineering responsibilities, this role requires a strong grasp of AI technologies and hands-on proficiency with AI-assisted development tools such as Claude Code. As AI reshapes how engineering work gets done, this engineer will be expected to leverage AI tooling in their own workflows and help drive adoption across the team. This role will operate as a point of contact for assigned projects, coordinate internal resources, and be supported by supervisors and technical resources. This person must be willing to take direction from leadership while also providing guidance and driving task ownership across team members. Our goal is to build and maintain strong, long-lasting customer relationships.

What You’ll Do: Serve as the point of contact for assigned projects, coordinating internal resources to deliver results. Leverage AI-assisted development tools including Claude Code to accelerate engineering workflows, automate repetitive tasks, and improve code quality. Evaluate, adapt, and integrate new tools and workflows into RL-R’s environment, with a focus on AI-powered capabilities. Maintain current knowledge of AI/LLM developments both internal and industry-wide and apply them to day-to-day engineering work. Lead proof-of-concept implementations for new tools and technologies. Evaluate potential problems and technical issues; develop and implement solutions. Create and maintain documentation, training materials, and best practices guides. Communicate with clients to identify and define project requirements, scope, and objectives. Ensure compliance with security, privacy, and data handling requirements.

Requirements

  • Strong technical background in software engineering or a related technical field.
  • Strong programming skills with Python (other languages C++, Java, Ruby, etc. are a bonus)
  • Hands-on experience with AI-assisted development tools (e.g., Claude Code or similar).
  • Demonstrated experience with LLMs, AI/ML tools, and integration patterns.
  • Experience with Linux OS (CentOS, Fedora, Ubuntu).
  • Experience with Windows OS.
  • Basic network troubleshooting skills.
  • Strong oral and written communication skills.
  • Experience understanding and communicating technical information to non-experts.
  • Experience planning and designing customer solutions.
  • Experience running projects from start to finish.
  • Strong, proactive work ethic.
  • Ability to learn and adapt to new situations and environments.
  • Typing speed of at least 50 words per minute
  • Ability to regularly lift 30 lbs (computers and associated equipment)
  • Ability to occasionally complete repetitive motion by lifting shoulder above head (testing cables in racks)
  • Ability to provide own vehicle transportation for travel between support sites
  • Ability to perform repetitive manual tasks
  • Ability to utilize a step stool to reach cable trays and racks, * 4+ years experience as a Systems Administrator (Windows and/or Linux).
  • 2+ years of networking experience and troubleshooting.
  • Proficiency with Claude Code including skills, hooks, MCP servers, and multi-agent workflows.
  • Experience with ML frameworks (TensorFlow, PyTorch).
  • Experience building MCP (Model Context Protocol) servers or custom tool integrations.
  • Experience installing and configuring Kubernetes, Docker, OpenCue, and Slurm.
  • Knowledge of network switch CLI and commands.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · WWC 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

50 sec

Why developer happiness matters in web frameworks

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

Videos

See all

Related articles

See all