> Markdown version of [/jobs/ext/3533324-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3533324-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** SANDTECH HOLDINGS LLC - **Location:** United States - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** OpenTofu, Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Collaborative Software, Data Security, DevOps, Domain Name System (DNS), Identity and Access Management, Subnetting, Python (Programming Language), Machine Learning, Operational Databases, Performance Tuning, Reliability Engineering, Runbook, Pulumi, Load Balancing, Large Language Models, Git Flow, Kubernetes, Terraform - **Published:** October 2, 2026 - **Apply:** https://startup.jobs/site-reliability-engineer-sand-tech-holdings-limited-c-10253699 ## About the Role * 5 to 8 years in platform engineering, cloud infrastructure, DevOps or site reliability, operating production systems that mattered to someone. * Cloud infrastructure, hands on. Azure and AWS. You have provisioned and run production data infrastructure rather than only managing it, with a strong understanding of networking, IAM and secure data handling, and how to design effectively within a client's existing landscape and governance constraints. * Networking you can reason about under pressure. VPCs and VNets, subnets, security groups, DNS, load balancing, private connectivity, and the patterns for getting services to talk to each other across a customer's internal and external boundaries. * Kubernetes and containers. Pods, deployments, services, namespaces. Comfortable in kubectl or any kube-api interface for inspection and troubleshooting. You do not need to have written an operator. * Infrastructure as code. You are comfortable with using IaC by default, from the start, with more than one toolchain. Demonstrable experience across one or more of Pulumi, Terragrunt, Terraform, OpenTofu, CDK or equivalents is required. Pulumi is used for Symmetri and experience with it is a strong plus. * Python for automation. As an organisation with strong data science roots we are strongly biased to python for many coding use cases, so familiarity is important and experience is a plus. * Change management and branching discipline. Standard git workflows, and the willingness to follow and improve a team's existing strategy rather than your own. * Resource planning. Sizing compute, memory and storage against what a customer actually needs and what their environment can actually give you, including when that environment is on premises and finite. * Incident response. You have been on call. You know the difference between fixing an outage and fixing the cause of one, you are able to run effective root cause analysis processes that result in long term improvements. * Client-facing competence with enterprise IT. You can hold a technical conversation with a customer's infrastructure and security people, be trusted by them, and represent a commitment without over-promising. You will not be equally strong across all of these aspects. Strong in most, familiar with the rest, unafraid of learning fast while being aware of and drawing in support for your current limitations. Advantageous experience * Public sector, utility or other regulated operational environments, and the security review processes that come with them * Air-gapped, sovereign or disconnected deployments of data-intensive systems * Observability and monitoring stacks, and building the alerting rather than only responding to it * Serving ML models and LLM systems in production, including offline, Due to the highly collaborative and internationally distributed nature of our work, successful candidates must be comfortable operating in small teams while contributing to larger, globally coordinated efforts. A strong sense of ownership, self-motivation and discipline in maintaining clear and consistent communication through virtual collaboration tools and video conferencing is essential. ## Description Sand operates in partnership with senior executives at our customers. Your primary client within these relationships will be the customer's enterprise IT organization. You will spend as much time in front of a municipal CIO's security team, an architecture review board and a cloud governance committee as you will in a terminal. You will be asked where the data sits, who has access to it, how it is encrypted, and what happens during an incident, and you will need to answer those questions accurately and calmly without escalating every one of them. Symmetri is designed to be highly modular and extensible, supporting rapid development and the mechanisms to support reusable intelligence capabilities across clients without sharing their data. As it continues to become the standard way we deliver, the line between deploying the platform and building it converges. You will be expected to see that coming and drive the strategy and execution that helps us get there. Your role will be to lead and deliver the deployment and the operation of cloud infrastructure, including Symmetri environments, for US customers, end to end., * Provisioning: Standing up infrastructure for a new environment, where the shape of the build depends on the environment type and variant, from managed cloud through single-tenant to on premises based on established patterns and practices. * Configuration: Tuning environments to a customer's needs as we land and expand on the use cases they are consuming: performance, scale, resource ceilings, network boundaries, identity, applying their security model rather than ours. * Lifecycle: Running the upgrade and maintenance cycles, and coordinating releases across live customer production estates without breaking the operations that depend on them. * Operation: Monitoring, alerting, and being first responder for critical production outages at any time of the day or night - which we see as an exceptional occurrence to be solved for, not the norm. * Collaboration & Growth: You will be the first US member of our Delivery Infra & Enablement function, but you will not need to work in isolation. You will collaborate directly with an established and experienced remote enablement team and regional operations counterparts in other geographies. This will assist you to establish a full US-based operations team and set the standard of how US deployment and operations actually works in practice. Your first 6 months * You have taken a US customer environment from provisioning to production and you are the named owner of it. * Runbooks for that environment exist, are accurate, and someone other than you (or their agent) could follow them safely. * Monitoring and alerting are in place and tuned, so that a page means something and silence means something too. * You have been through at least one customer security review as our technical voice in the room. * You can say, with evidence, which parts of our global deployment approach need to work differently in the United States. ## Related Videos - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [Shifting Stress to Progress— Understanding DevOps to do DevOps Better](https://www.wearedevelopers.com/videos/268-shifting-stress-to-progress-understanding-devops-to-do-devops-better) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [3 Key Steps for Optimizing DevOps Workflows](https://www.wearedevelopers.com/videos/962-3-key-steps-for-optimizing-devops-workflows) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)