Senior DevOps Engineer

Brahma
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
£ 95K

Job location

Remote

Tech stack

Artificial Intelligence
Airflow
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Azure
Cloud Computing
Configuration Management
Computer Networks
Continuous Integration
Linux
DevOps
Github
Identity and Access Management
Python
Key Management
Network Segmentation
Role-Based Access Control
Ansible
Prometheus
Shell Script
Visual Effects
Ceph
Autoscaling
Grafana
Caching
Gitlab
Kubernetes
Kafka
Slurm
Machine Learning Operations
Terraform
Software Version Control

Job description

Own build, deploy, and runtime reliability across BRAHMA AI's hybrid estate. Deliver secure, scalable infrastructure for Gen AI based workflows and products across hybrid environments. Partner with infrastructure and multidisciplinary product and research teams to help them innovate and ship fast.We are hiring remotely across the EMEA region., * Design, implement, and operate Slurm and Kubernetes-based platforms across cloud and on-prem GPU nodes, including autoscaling, rollout strategies, and multi-cluster operations.

  • Build CI/CD pipelines for services, model training, and model serving; standardise artifact/version management and environment promotion.
  • Implement Infrastructure as Code with Terraform/Terragrunt and configuration management; enforce drift detection and repeatable environments.
  • Design and implement observability stacks (metrics, logs, tracing); drive incident response and postmortems.
  • Secure the stack with least privilege, secrets management, network policy, and hardened baselines; support ISO/MPA controls with the security team.
  • Operate model-serving infrastructure for real-time and batch workloads; optimise GPU utilisation, concurrency, and latency.
  • Drive cost visibility and efficiency across compute, storage, and egress; forecast capacity
  • and plan lifecycle of hardware and licenses., * Model serving stacks and GPU telemetry/optimization.
  • On-prem operations for GPU/CPU fleets.
  • HPC/VFX pipeline exposure; render farms; real-time engines.
  • Storage systems (S3/MinIO, Ceph/Lustre/NFS), CDN, and caching strategies.
  • Messaging/streaming (Kafka) and workflow/orchestration (Argo, Airflow).

Requirements

  • 6+ years in DevOps/SRE/Platform roles running production systems.
  • Expert with Kubernetes and containers (runtime, scheduling, networking, autoscaling).
  • Strong with Terraform and at least one configuration management tool (Ansible
  • preferred).
  • CI/CD (GitHub Actions [preferred] / GitLab), release strategies, and artifact registries.
  • Observability in production (Prometheus/Grafana preferred).
  • Linux mastery, shell scripting, and a high-level language (Python preferred).
  • Cloud proficiency (AWS/GCP/Azure) and Security fundamentals: IAM, secrets management, network segmentation, image provenance.
  • Experience with data-/media-heavy workloads or ML pipelines in production.
  • Location: EU/UK time zones (±2h)., * Pragmatic and systems-thinking oriented.
  • Bias to automate and simplify.
  • Clear communication during incidents and reviews.
  • Ownership across design, operations, and quality.

About the company

BRAHMA AI is the next generation of enterprise media technology formed through the integration of Prime Focus Technologies and Metaphysic. By combining CLEAR®, CLEAR® AI, ATMAN, and VAANI into one ecosystem, BRAHMA AI enables enterprises to manage, create, and distribute content with intelligence, security, and efficiency. Proven, scalable, and enterprise-tested, BRAHMA AI is helping global organizations accelerate growth, efficiency, and creative impact in the AI-powered era.

Apply for this position