> Markdown version of [/jobs/ext/2510620-ai-pipeline-engineer-security-automation-platform-imunify360](https://www.wearedevelopers.com/jobs/ext/2510620-ai-pipeline-engineer-security-automation-platform-imunify360). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Pipeline Engineer - Security Automation Platform (Imunify360) - **Company:** Cloudlinux - **Location:** Madrid, Spain (Remote available) - **Contract:** Permanent contract - **Skills:** PHP (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon S3, Configuration Management, Databases, Continuous Integration, Extract Transform Load (ETL), Linux, Fault Tolerance, Job Scheduling, Python (Programming Language), Release Management, Reliability Engineering, Ansible, Prometheus, WordPress, Workflow Management Systems, Ceph (Software), Delivery Pipeline, Large Language Models, Grafana, State Machines, Backend, Data Layers, Gitlab-ci, Data Management, Vertica, Puppet, Docker, Security Orchestration, Automation & Response, Web Api - **Published:** August 10, 2026 - **Apply:** https://es.indeed.com/viewjob?jk=0cfbe850b7c9e9d2 ## About the Role To thrive in this role, you should have: * 5+ years of professional backend / platform / infrastructure engineering experience; * Demonstrable experience building and operating multi-stage data or automation pipelines - CI/CD systems, ETL/ELT, build and release automation, job orchestration, ML/data platforms, or similar. This is the single most important requirement. We will ask you to walk us through one in detail; * Real depth in at least one of Python, Go or Rust. We use all three, and we are not hiring a language specialist - which one you bring genuinely does not matter. Depth is simply how we verify the experience behind it is real, so expect specific questions about systems you have designed and shipped: why they are shaped the way they are, how they behave under failure, and what you would build differently today; * Systems design judgement, more than raw coding throughput. The difficult part of this role is deciding what to build, working out where it will break, and making it prove its own correctness - not volume of code produced; * Practical experience with workflow orchestration and job scheduling (Airflow, Temporal, Prefect, Dagster, Argo, custom schedulers - whatever you have actually run in production); * A working instinct for reliability engineering : idempotency, retries with backoff, exactly-once vs at-least-once, checkpointing and resumability, graceful degradation, backpressure, and safe handling of partial failure; * Hands-on observability experience - Prometheus/Grafana, LGTM stack, or equivalent - including designing metrics rather than only consuming dashboards someone else built; * Deep CI/CD experience, ideally GitLab CI including dynamic/child pipelines and self-hosted runners; comfort with Docker and container-based test environments; * Experience with object storage (S3/Ceph or equivalent) and with large-scale analytical stores - ClickHouse or another columnar database; * Comfort designing state machines and long-running processes that survive restarts, and reasoning about concurrency across multiple in-flight rollouts; * Excellent debugging skills across system, network and data layers; * Strong communication skills and comfort working in a distributed team; * At least upper-intermediate proficiency in spoken and written English. Nice to have: * Experience with progressive delivery - canary and percentage-based rollouts, feature flags, automated rollback, blast-radius control; * Experience running AI/LLM systems in production , particularly cost control, token accounting, evaluation harnesses, and dealing with non-deterministic components inside a deterministic pipeline; * Experience with fleet-scale telemetry and with building quality gates on top of noisy production signal; * Familiarity with WordPress, PHP, or WAF/ModSecurity concepts; * Experience with configuration management (Ansible, Puppet, Salt) and with Linux service operations. You do not need a cybersecurity background. Most of the hard problems here are orchestration, reliability, correctness under concurrency and observability. Domain knowledge is learnable and we have specialists to learn it from - pipeline engineering judgement is what we cannot substitute. We value engineers who are: * Curious and fearless problem solvers - not afraid to dig into existing systems, investigate root causes, and propose improvements; * Sceptical by default - who ask what would falsify a conclusion before acting on it, and who trust measurements over plausible reasoning; * Pragmatic and detail-oriented - focused on building reliable, maintainable systems, and allergic to solutions that require a human to remember something; * Owners - comfortable being the person accountable for whether a pipeline ran correctly last night; * Effective communicators - able to articulate ideas clearly, exchange feedback constructively, and foster collaboration across teams; * Engaging and proactive - contributing energy, initiative, and a positive presence that strengthens team culture. ## Description * Designing, building and operating the automated pipelines described above, end to end; * Turning fragile multi-stage batch jobs into resumable, idempotent, observable systems with explicit state machines and recovery paths; * Defining and enforcing latency budgets and SLOs per stage, and making violations visible and actionable rather than silent; * Building the observability layer - metrics, dashboards, alerting and health gates - so the pipeline reports its own condition instead of needing someone to go and look; * Designing and implementing guardrails: automatic hold and rollback, blast-radius limits, kill switches, and safe-by-default behavior when an upstream dependency is unavailable; * Making the systems low-maintenance: eliminating manual steps, removing standing human babysitting, and reducing the operational surface rather than adding to it; * Writing and maintaining unit and integration tests for logic that is genuinely hard to test - concurrency, partial failure, external API flakiness, multi-stage state; * Investigating and resolving complex issues across ClickHouse, GitLab CI, S3/object storage, Prometheus/Grafana and third-party APIs; * Collaborating with the security analysts and the Server team on architecture, and pushing back when a proposed design will not survive contact with production. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)