AI Infrastructure Engineer (Storage)

CommonAI CIC
Cambridge, UK
9 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Microsoft Azure Bash Shell Cloud Computing Cloud Storage System Configuration Data Integrity Extract Transform Load (ETL) Data Security
+27 more
Linux Distributed Data Store Monitoring of Systems Python (Programming Language) Linux System Administration Open Source Technology Performance Tuning Cloud Services Ansible Prometheus Data Streaming AI Infrastructure Ceph (Software) Scripting Data Storage Management Google Cloud Cloud Platform System Grafana Containerization AI Platforms Kubernetes Storage Technologies Data Analytics Data Management Terraform Data Pipelines Docker

Job description

CommonAI CIC is a non-profit membership organisation, founded on a belief in collaborative engineering for the safe and responsible development of foundational AI technologies. A place where AI startups, enterprises large and small, public sector bodies and academia can share resources and knowledge, to codevelop and grow businesses, fast.

We support technology-focused start ups, each with unique data management challenges, and are seeking an experienced Infrastructure Engineer to help them design, deploy and maintain high-performance storage systems for their AI and data-driven workloads. The successful candidate will combine deep experience architecting and managing distributed, cloud, and tiered storage solutions with strong Linux and automation skills.

In this role you will:

  • Design, implement, and maintain storage platforms that support large-scale AI and data pipelines
  • Manage distributed storage systems such as Ceph, Lustre, or BeeGFS
  • Oversee tiered storage architectures, optimising data movement across high-performance, object, and archival tiers
  • Ensure data integrity, availability, and security across on-premises and cloud environments
  • Develop automation and monitoring tools using Bash, Python, or similar scripting languages
  • Manage and secure container images and related storage used for AI and ML workloads
  • Integrate storage systems with public cloud services (AWS, Azure, GCP) and hybrid environments
  • Troubleshoot complex storage and data flow issues, collaborating closely with AI platform and infrastructure teams
  • Contribute to ongoing architecture improvements, performance tuning, and capacity planning

Requirements

  • Strong Linux system administration background
  • Proven experience installing, configuring, and maintaining Ceph clusters or similar technologies in a production environment
  • Familiarity with distributed filesystems (e.g., Lustre, BeeGFS) and cloud-based storage services (e.g. EC2)
  • Experience with tiered storage management and lifecycle data policies
  • Scripting and automation proficiency (e.g. Bash, Python, Terraform/OpenTofu, Ansible)
  • Understanding of data security best practices and compliance considerations
  • Experience working with container technologies (e.g. Docker, Kubernetes) and image storage registries
  • Strong analytical, troubleshooting, communication and documentation skills

We also value:

  • Knowledge of GPU compute environments or AI training infrastructure
  • Experience with monitoring and observability tools (Prometheus, Grafana, etc.)
  • Contributions to open-source storage, data management, or infrastructure projects
  • Familiarity with object storage systems (S3, RADOS Gateway, MinIO, etc.)

Benefits

  • A collaborative and supportive work environment
  • The opportunity to have a high impact in a growing organisation
  • Competitive salary package and pension
  • Professional development opportunities
  • Networking opportunities with influential people from across the tech sector and academia
  • A vibrant office environment located a few minutes walk away from Cambridge train station

CommonAI CIC is an equal opportunity employer and is committed to creating an inclusive and diverse workplace.

Requirements

  • Strong Linux system administration background
  • Proven experience installing, configuring, and maintaining Ceph clusters or similar technologies in a production environment
  • Familiarity with distributed filesystems (e.g., Lustre, BeeGFS) and cloud-based storage services (e.g. EC2)
  • Experience with tiered storage management and lifecycle data policies
  • Scripting and automation proficiency (e.g. Bash, Python, Terraform/OpenTofu, Ansible)
  • Understanding of data security best practices and compliance considerations
  • Experience working with container technologies (e.g. Docker, Kubernetes) and image storage registries
  • Strong analytical, troubleshooting, communication and documentation skills

We also value:

  • Knowledge of GPU compute environments or AI training infrastructure
  • Experience with monitoring and observability tools (Prometheus, Grafana, etc.)
  • Contributions to open-source storage, data management, or infrastructure projects
  • Familiarity with object storage systems (S3, RADOS Gateway, MinIO, etc.)

Benefits & conditions

  • A collaborative and supportive work environment
  • The opportunity to have a high impact in a growing organisation
  • Competitive salary package and pension
  • Professional development opportunities
  • Networking opportunities with influential people from across the tech sector and academia
  • A vibrant office environment located a few minutes walk away from Cambridge train station

About the company

CommonAI CIC is a non-profit membership organisation, founded on a belief in collaborative engineering for the safe and responsible development of foundational AI technologies. A place where AI startups, enterprises large and small, public sector bodies and academia can share resources and knowledge, to codevelop and grow businesses, fast.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

2:27 min

Core infrastructure components required for an AI factory

Thomas Schmidt Thomas Schmidt · World Congress 2024

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all