AI Platform Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+8 more
Job description
The AI Platform Engineer is responsible for deploying and managing the computing and storage platform backbone that powers AI models. Working on-site alongside AI Data Engineering, IT Infrastructure, Security, and business stakeholders, this role is central to the reliability, performance, and scalability of our AI/ML infrastructure., * Provision and maintain CPU and GPU compute clusters, high-bandwidth network equipment, and scalable, high-throughput storage systems for AI/ML training and inference, working on-site and virtually to coordinate hands-on configuration and troubleshooting with infrastructure teams.
- Develop and maintain scripts and infrastructure-as-code (IaC) modules for automated configuration of AI/ML infrastructure using tools such as Terraform and Ansible; participate in in-person and virtual code and design reviews with the platform engineering team.
- Develop and maintain CI/CD pipelines and container orchestration platforms (Kubernetes, Flux); troubleshoot pipeline failures and escalate to more senior engineers or other teams as needed, collaborating in person and virtually to resolve issues quickly and effectively.
- Set up and monitor data stores (file systems, block storage, traditional and vector databases) to support LLMs and RAG pipelines, with regular on-site coordination with AI Data Engineering teams to ensure data infrastructure meets evolving model requirements.
- Monitor system performance and log metrics to ensure reliability and uptime; implement system improvements under guidance of more senior engineers, including attending in-person and virtual planning and incident review sessions.
- Collaborate in person with AI Data Engineering, IT Infrastructure, Security, and Business Leaders to optimize infrastructure for performance and cost, including participating in cross-functional working sessions, architecture reviews, and stakeholder briefings.
- Deploy and maintain Model Context Protocol (MCP) connectors and ensure secure, scalable infrastructure for MCP integrations, coordinating directly with security and engineering teams to validate configurations and address emerging requirements.
Requirements
- At least two (2) years of experience working in an infrastructure, DevOps, or platform engineering role, with at least some exposure to AI/ML workloads.
- Proficiency in cloud computing, container orchestration, and infrastructure-as-code.
- Knowledge of networking and distributed storage;
- Experience with monitoring and observability tools to track performance and system health.
- Ability to troubleshoot complex systems and collaborate in person with Data Engineering, Infrastructure, AI Engineering, Security, and Business Leaders.
- Strong problem-solving, communication, and teamwork skills; comfortable working in a collaborative, on-site team environment.
- Foundational understanding of AI/ML concepts.
Preferred Qualifications:
- Knowledge of GPU acceleration.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on recruiting.ultipro.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Stephan Gillich - Bringing AI Everywhere
Navigating the AI Shift
What Industries Outside of AI Are Hiring The Most AI Experts?
MLOps And AI Driven Development