> Markdown version of [/jobs/ext/2304161-systems-engineer](https://www.wearedevelopers.com/jobs/ext/2304161-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Systems Engineer - **Company:** Refractal - **Location:** Greater London, UK - **Contract:** Permanent contract - **Skills:** Systems Engineering, Automation of Tests, Software as a Service, Cloud Computing, Continuous Delivery, Data Integrity, Relational Databases, Identity and Access Management, Python (Programming Language), PostgreSQL, Open Source Technology, Role-Based Access Control, Service Design, Data Logging, Data Storage Technologies, Multi-Agent Systems, Backend, Low Latency, Machine Learning Operations, Dynatrace - **Published:** August 30, 2026 - **Apply:** https://www.collegerecruiter.com/job/2830557912-systems-engineer ## About the Role We assess demonstrated ability, rather than a specific credential or number of years in industry. You should have: * Strong experience with backend, distributed or infrastructure systems. Experience with Python is preferred. * Production experience with observability and distributed tracing. You can connect actions, tool calls, requests, model calls and enforcement decisions across agents and services. * Experience building and testing responsive services under load. You understand latency, timeouts, retries, backpressure and failure handling. * Experience with relational databases and multi-tenant data. This should include Postgres, Row-Level Security, migrations and access-control tests. * Hands-on experience with containers, cloud infrastructure and continuous delivery. Experience with GCP, Cloud Run and Cloud SQL is preferred. * A good understanding of operational security. This includes least privilege, secret management, hardened containers and trust boundaries. * Experience with UX design and Product Management is preferred. Particularly strong signals include: * A production system that you owned where latency, correctness or integrity under load was important. * Production observability or agent-tracing work, including the correlation of multi-step or multi-agent activity with OpenTelemetry or similar tools. * Experience with multi-tenant SaaS isolation, Postgres Row-Level Security or compliance-driven infrastructure. * Model-serving or MLOps experience, including GPU inference, detector artefacts, provider routing or containerised model deployment. * Work with sandboxing, subprocess isolation, tamper-evident logging, supply-chain integrity or other security controls. * Open-source infrastructure contributions or systems write-ups that show clear engineering decisions. ## Description As a Systems Engineer, you will own the services and infrastructure that keep Refractal secure and reliable in production. These systems help customers secure the agents they deploy and defend against malicious agents operated outside their organisation. You will work across the enforcement runtime, data integrity, tenant isolation, agent tracing, model services and cloud infrastructure. Your work will cover individual actions and wider attack sequences from internal and external agents. You will build systems that receive and normalise events, connect related activity across agents, services and tools, evaluate risk, apply enforcement decisions and record the evidence. You will work directly with the founders and early customers. You will turn security and operational requirements into systems that are measurable, testable and safe under failure. You will also help set the technical direction, engineering standards and product roadmap. What You'll Work On * Build and operate the runtime enforcement pipeline. It must evaluate activity from customer agents and external malicious agents, then return allow, change or block decisions with low latency and predictable failure behaviour. * Own observability and agent tracing across the platform. Trace individual events and connect activity across internal agents, external agents, services, tools and targets. Provide useful traces, logs, metrics and alerts. * Build the tamper-evident evidence store and audit ledger. Customers must be able to verify that decision records have not been changed. * Enforce multi-tenant isolation in every service and data path. Use Postgres Row-Level Security, separate access paths and automated tests to prevent cross-tenant access. * Build the infrastructure for red-team and evaluation jobs. Manage subprocesses, timeouts, cancellation and output safely. A failed job must not hang or escape its limits. * Build reliable event and streaming services. Handle backpressure, disconnections, retries and load without losing the meaning of an enforcement decision. * Build and operate model-serving services. Manage detector artefacts, container builds, provider routing, credentials, fallbacks and GPU workloads. * Own cloud infrastructure and continuous delivery. This includes Cloud Run, Cloud SQL, IAM, Secret Manager and the checks that gate each deployment. * Threat-model the infrastructure. Protect secrets, harden containers, verify artefacts and define clear trust boundaries between customer traffic, the control plane and model services. * Oversee Product Management of software dashboard product, including UX design and design decisions. What We Are Looking For We need a systems engineer who takes end-to-end ownership of security-critical services. You can work from service design and data storage to deployment and production operation. You reason about latency, failure, isolation, integrity and cost. You use measurements instead of assumptions. You add tracing, metrics and logs before you call a service production-ready. You deliver small, correct changes and test failure cases. You understand that a slow decision, a missing trace, a leaked row or a changed audit record can be a security fault. Most importantly, you want to build the systems that let organisations deploy their own autonomous systems safely and defend against malicious agents operated by others. ## Related Videos - [Optimizing Discovery: PostgreSQL's Role in Transforming GetYourGuide's Search](https://www.wearedevelopers.com/videos/1647-optimizing-discovery-postgresql-s-role-in-transforming-getyourguide-s-search) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Meet Your New BFF: Backend to Frontend without the Duct Tape](https://www.wearedevelopers.com/videos/682-meet-your-new-bff-backend-to-frontend-without-the-duct-tape) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)