> Markdown version of [/jobs/ext/1304329-staff-senior-software-devops-engineer](https://www.wearedevelopers.com/jobs/ext/1304329-staff-senior-software-devops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff/Senior Software DevOps Engineer - **Company:** OLIX - **Location:** London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Test Suite, Amazon Web Services, Cloud Computing, Continuous Integration, Linux, DevOps, Field-Programmable Gate Array (FPGA), Github, Hardware-In-The-Loop Simulation, Python (Programming Language), Remote Access Technology, Prometheus, System Programming, Test Execution Engine, Datadog, Scripting, Grafana, Hardware Testing, Information Technology - **Published:** July 17, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=ebd19e88cdc48aed ## About the Role * Experience build/test infrastructure, CI/CD, developer-productivity, or large-scale systems / release engineering, with demonstrated end-to-end ownership of a big test or CI system. * Scaling a large test suite: staged lanes, parallelism with real isolation, content-addressed caching, and keeping feedback fast and affordable as the suite grows. * Fleet and scarce-resource management: heterogeneous CI runner fleets across cloud and on-prem (VMs, containers, and bare-metal hardware), and giving a team shared, monitored access to scarce or expensive resources - custom accelerators, FPGA/prototype rigs, lab hardware, or contended compute - including reservations, remote access, hardware-in-the-loop testing, and artifact portability across heterogeneous hosts. * Performance-analysis support: trustworthy regression baselines, determinism and noise handling, sound metric aggregation (geometric vs arithmetic mean vs median), and fast attribution and bisection. * Observability and metrics platforms: CI-health and readiness dashboards, a metrics store that scales to long-lived, high-cardinality time series, and self-serve access for engineers. * Strong scripting and systems programming (e.g. Python plus a systems language), and fluency with containers, Linux, and cloud infrastructure (AWS or similar). * A reproducible, safe-by-default mindset: hermetic builds, least privilege, fail-closed defaults, and blast-radius control. * Excellent communication and the ability to align and influence cross-functional teams (compiler, runtime, modelling) without relying on formal authority. * Bachelor's degree or higher in computer science, electrical engineering, mathematics, or a related field. Nice to Have * GitHub Actions or comparable CI at scale; scaling CI runner fleets on cloud infrastructure (e.g. AWS); hardware-in-the-loop or lab automation for custom silicon or FPGA bring-up; time-series and observability stacks (Prometheus/Grafana, Datadog, columnar warehouses). * Adjacent depth is welcome: HPC / cluster batch scheduling, release engineering, or developer-productivity platforms. ## Description We're searching for a Staff/Senior Software DevOps Engineer to own the build, test, and CI flows that the entire DX-1 software stack - compiler, runtime, simulator, and framework integration - depends on. Ours is a large test suite that asserts token-exact correctness against golden references, and it has to run across scarce, expensive resources that span both simulation compute and hardware-in-the-loop testing - simulator and emulator boxes alongside DX-1 and prototype-platform boards. Your mission is to keep that system fast, trustworthy, observable, and affordable as the test suite, the team, and the resource pool all grow. This is a build-and-test- role, not product-serving SRE. You'll work where CI, the runner fleet, and the test hardware meet, partnering closely with the infrastructure, compiler, runtime, simulator, and modelling teams. At Senior/Staff level your impact is the velocity of every engineer who depends on this system - how fast they get a trustworthy signal, how rarely they wait on a machine or a flaky run, and how much they can self-serve without coming to you. That leverage - the standards, platforms, and shared-resource model others build on - is what we're hiring for, far more than any single system you ship. Responsibilities Own the Build & Test Pipelines: Design, build, and own CI pipelines and test-execution - PR, merge, and nightly lanes - that gates the entire software stack, balancing fast feedback against coverage and cost. Scale Test Execution: Split a large, slow suite into staged lanes, parallelize it with real test isolation, and cache aggressively with content-addressed keys, so feedback stays fast and cheap as the suite and the team grow rather than getting solved by throwing machines at it. Manage the Fleet & Scarce Resources: Run CI across a heterogeneous fleet of cloud and self-hosted machines, and give the team fair, monitored, fail-fast shared access to scarce and expensive hardware - keeping it reliable, well-utilised, and never a silent bottleneck. Build the Performance & Readiness Signal: Stand up perf-regression baselines the team trusts (pinned hardware, rolling baselines, sound metric aggregation, determinism), and turn CI and test signal into CI-health and product-readiness dashboards that drive real decisions. Own Software Observability: Choose the metrics store that scales to many series on daily runs with long-lived history, making dashboards for observable software. Set Standards: Define the flows that keep builds and test runs hermetic and reproducible; and make the system fail closed and contain the blast radius when something is misconfigured or a job is untrusted. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [The journey from developer to devops - what i've learnt along the way](https://www.wearedevelopers.com/videos/238-the-journey-from-developer-to-devops-what-i-ve-learnt-along-the-way) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [The 8 Best Code Testing Tools](https://www.wearedevelopers.com/magazine/402-the-8-best-code-testing-tools) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read)