Software Engineer - SRE Core Infra
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+11 more
Job description
PhysicsX is growing rapidly and so is the infrastructure that underpins our platform. We are building core infrastructure that is reliable, secure, scalable and reproducible across multiple cloud providers and on premises environments. As the platform evolves to serve increasingly complex engineering workloads, the foundational infrastructure layer becomes ever more critical., We are looking for a Senior Software Engineer to join our Platform SRE Core Infrastructure team. This role is responsible for the design, provisioning and operation of the shared infrastructure that the entire PhysicsX platform depends on. You will work across infrastructure as code, Kubernetes cluster management, secrets management, GPU drivers, networking and multi tenancy architecture to ensure the platform is dependable, scalable and secure.
This is a role for an engineer who combines deep infrastructure expertise with a reliability engineering mindset and an appreciation for the developer experience of the teams that consume the platform.
What You Will Do
- Own the design and delivery of core infrastructure across multi cloud providers (AWS, GCP, Azure) and on premises environments using Terraform and Crossplane.
- Architect and operate Kubernetes clusters supporting both single tenant and multi tenant workloads, with a strong emphasis on isolation, performance and reliability.
- Define and implement infrastructure provisioning patterns using Crossplane compositions and Terraform modules, ensuring reproducibility and auditability across environments.
- Design and operate secrets management solutions, including dynamic secret provisioning, rotation and fine grained access control integrated with cluster identity.
- Manage and maintain GPU driver configurations and accelerated compute node pools, ensuring compatibility and performance for AI and simulation workloads.
- Own cluster networking design including CNI selection, Istio service mesh integration, ingress strategy and cross cluster connectivity.
- Implement and maintain vCluster based multi tenancy to provide strong workload isolation within shared infrastructure.
- Develop lightweight Kubernetes Operators or controllers where automation of infrastructure lifecycle tasks requires it.
- Establish SLOs and reliability targets for core infrastructure components and lead the response to production incidents.
- Partner with security and platform teams to enforce infrastructure governance, network policies and compliance controls.
- Contribute to and uphold engineering standards across the platform organisation., We collect diversity and inclusion data solely for the purpose of monitoring the effectiveness of our equal opportunities policies and ensuring compliance with employment and equality legislation. This information is confidential, used only in aggregate form, and will not influence the outcome of your application.
Requirements
- Kubernetes depth - 5 or more years of professional experience operating Kubernetes in production. You have a thorough understanding of cluster architecture, the scheduler, networking, storage and the API lifecycle. Kubernetes certifications such as CKAD, CKA or CKS are highly desirable.
- Crossplane expertise - significant hands on experience designing and operating Crossplane compositions, providers and managed resources in production environments.
- Terraform proficiency - strong experience authoring, structuring and operating Terraform at scale, including state management, module design and CI integration.
- Multi cloud and on premises - practical experience operating infrastructure across more than one cloud provider and on premises environments, with an understanding of the differences in identity, networking and storage.
- Multi tenancy architecture - experience designing and implementing both single tenant and multi tenant Kubernetes architectures, with strong views on isolation, resource governance and operational overhead.
- Secrets management - experience with tools such as Vault, External Secrets Operator or cloud native secret stores, including dynamic provisioning and rotation.
- Networking - solid knowledge of Kubernetes networking, CNI plugins, Istio service mesh and ingress patterns. Experience with cross cluster or hybrid connectivity is valuable.
- vCluster and virtual clusters - experience using vCluster or similar tooling to provide lightweight, isolated Kubernetes environments within shared clusters.
- GPU and accelerated compute - familiarity with GPU driver management, device plugins and the operational considerations of running accelerated workloads in Kubernetes.
- Kubernetes Operators - at least lightweight experience writing or extending Kubernetes Operators or controllers, ideally in Golang or Python.
- Software engineering capability - you are comfortable writing code to automate and extend infrastructure. Python and Golang are the primary languages used across the platform. Exposure to functional programming languages such as Erlang, Elixir or OCaml is appreciated.
- Platform engineering mindset - you think about infrastructure as a product consumed by engineering teams and prioritise usability, documentation and long term maintainability.
- Distributed systems experience - a solid grounding in distributed systems concepts, including failure modes, consistency and the operational challenges of running systems at scale.
Ideally
- Experience with GitOps workflows using tools such as Argo CD or Flux.
- Contributions to open source infrastructure or Kubernetes ecosystem projects.
About the company
Help shape an AI-native engineering company at a formative stage, tackling problems that genuinely matter for industry and society. This is work with real-world impact - and something you can be proud to stand behind.
Learn alongside exceptional people
Work with a high-caliber, collaborative team of engineers, scientists, and operators who care deeply about doing great work, and about helping each other get better. We come from diverse backgrounds, but we share a commitment to operating at the highest level and addressing some of the most complex challenges out there. If you’re ambitious, thoughtful, and driven by impact, you’ll feel at home.
Influence over hierarchy
We operate with a flat structure: good ideas win - wherever they come from. Questioning assumptions and challenging the status quo isn’t just welcomed, it’s expected.
Sustainable pace, long-term ambition
Building meaningful technology is a marathon, not a sprint. We believe in balancing focused, ambitious work with a life beyond it. Our hybrid model blends time together in our Shoreditch office with work-from-home days, giving you the flexibility to work sustainably while staying connected in person.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Is Software Engineering Over-Saturated?
Dev Digest 121 - AI goes offline
Highest Paying Tech Companies for Developers