> Markdown version of [/jobs/ext/2727189-software-devops-engineer](https://www.wearedevelopers.com/jobs/ext/2727189-software-devops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software DevOps Engineer - **Company:** Olix - **Location:** UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Test Suite, Amazon Web Services, Cloud Computing, Continuous Integration, Linux, Field-Programmable Gate Array (FPGA), Github, Hardware-In-The-Loop Simulation, Python (Programming Language), Remote Access Technology, Prometheus, System Programming, Test Execution Engine, Datadog, Scripting, Grafana, Hardware Testing, Parallel Computation, Information Technology, Bare Metal - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/staff-senior-software-devops-engineer-olix-8754647 ## About the Role * Experience in build/test infrastructure, CI/CD, developer productivity, or large-scale systems and release engineering, with demonstrated end-to-end ownership of a large test or CI system * Experience scaling a large test suite through staged lanes, parallelism with real isolation, content-addressed caching, and maintaining fast, cost-effective feedback as the suite grows * Experience managing heterogeneous CI runner fleets across cloud and on-prem environments, including VMs, containers, and bare-metal hardware. Ability to provide teams with shared, monitored access to scarce or expensive resources such as custom accelerators, FPGA/prototype rigs, lab hardware, or contended compute, including reservations, remote access, hardware-in-the-loop testing, and artifact portability across heterogeneous hosts * Experience supporting performance analysis with trustworthy regression baselines, deterministic testing, noise handling, sound metric aggregation (geometric mean, arithmetic mean, and median), and fast attribution and bisection * Experience building observability and metrics platforms, including CI health and product readiness dashboards, metrics stores that scale to long-lived, high-cardinality time series, and self-service access for engineers * Strong scripting and systems programming skills (for example, Python plus a systems language), along with proficiency in containers, Linux, and cloud infrastructure such as AWS * A reproducible, safe-by-default mindset, including hermetic builds, least privilege, fail-closed defaults, and blast radius control * Excellent communication skills and the ability to align and influence cross-functional teams, including compiler, runtime, and modelling teams, without relying on formal authority * Bachelor's degree or higher in Computer Science, Electrical Engineering, Mathematics, or a related field Nice to Have * GitHub Actions or comparable CI at scale; scaling CI runner fleets on cloud infrastructure (e.g. AWS); hardware-in-the-loop or lab automation for custom silicon or FPGA bring-up; time-series and observability stacks (Prometheus/Grafana, Datadog, columnar warehouses) * Adjacent depth is welcome: HPC / cluster batch scheduling, release engineering, or developer-productivity platforms ## Description We're searching for a Staff/Senior Software DevOps Engineer to own the build, test, and CI flows that the entire DX-1 software stack, including the compiler, runtime, simulator, and framework integration, depends on. Ours is a large test suite that asserts token-exact correctness against golden references, and it has to run across scarce, expensive resources that span both simulation compute and hardware-in-the-loop testing, including simulator and emulator boxes alongside DX-1 and prototype-platform boards. Your mission is to keep that system fast, trustworthy, observable, and affordable as the test suite, the team, and the resource pool all grow. This is a build-and-test role, not product-serving SRE. You'll work where CI, the runner fleet, and the test hardware meet, partnering closely with the infrastructure, compiler, runtime, simulator, and modelling teams. At the Senior/Staff level, your impact is the velocity of every engineer who depends on this system: how fast they get a trustworthy signal, how rarely they wait on a machine or a flaky run, and how much they can self-serve without coming to you. That leverage, through the standards, platforms, and shared resource model others build on, is what we're hiring for far more than any single system you ship. Responsibilities Own the Build & Test Pipelines: Design, build, and own CI pipelines and test execution across PR, merge, and nightly lanes that gate the entire software stack, balancing fast feedback with coverage and cost. Scale Test Execution: Split a large, slow suite into staged lanes, parallelize it with real test isolation, and cache aggressively using content-addressed keys so feedback stays fast and cost-effective as the suite and the team grow, rather than relying on simply adding more machines. Manage the Fleet & Scarce Resources: Run CI across a heterogeneous fleet of cloud and self-hosted machines, and give the team fair, monitored, fail-fast shared access to scarce and expensive hardware, keeping it reliable, well utilized, and never a silent bottleneck. Build the Performance & Readiness Signal: Stand up performance regression baselines the team trusts using pinned hardware, rolling baselines, sound metric aggregation, and deterministic testing. Turn CI and test signals into CI health and product readiness dashboards that drive real decisions. Own Software Observability: Choose the metrics store that scales to many time series across daily runs with long-lived history, making dashboards for observable software. Set Standards: Define the flows that keep builds and test runs hermetic and reproducible, and make the system fail closed while containing the blast radius when something is misconfigured or a job is untrusted. ## Related Videos - [GitOps for the people](https://www.wearedevelopers.com/videos/815-gitops-for-the-people) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Monitoring as Code - Managing your dashboards at scale](https://www.wearedevelopers.com/videos/753-monitoring-as-code-managing-your-dashboards-at-scale) - [Lights, Camera, GitHub Actions!](https://www.wearedevelopers.com/videos/734-lights-camera-github-actions) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The 8 Best Code Testing Tools](https://www.wearedevelopers.com/magazine/402-the-8-best-code-testing-tools) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers)