> Markdown version of [/jobs/ext/2634501-site-reliability-engineer-public-sector](https://www.wearedevelopers.com/jobs/ext/2634501-site-reliability-engineer-public-sector). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - Public Sector - **Company:** Trellix Public Sector LLC - **Location:** Cambridge, MA, United States (Remote available) - **Experience:** Experienced - **Salary:** $140,000.0 - $170,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Customer Data Management, DevOps, Python (Programming Language), Reliability Engineering, Software Deployment, Datadog, Data Logging, Pulumi, Cloud Platform System, AI Platforms, Kubernetes, Infrastructure Automation Frameworks, Terraform, Legacy Systems, Golang - **Published:** August 4, 2026 - **Apply:** https://jobs.ashbyhq.com/blitzy/0cd8c412-15f3-4139-a162-50c2f7957d64 ## About the Role Eligibility: U.S. citizenship required (customer badging requirement); no security clearance required, * U.S. citizenship (customer badging requirement) and ability to complete a customer background/badging process. * 3+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering roles. * Strong proficiency in Kubernetes and container orchestration; hands-on experience deploying software into customer-controlled or restricted environments. * Experience operating in isolated, restricted, or otherwise highly regulated network environments (defense, government, financial services, or similar). * Hands-on experience with infrastructure-as-code tools (Terraform, Pulumi, or equivalent) and at least one major cloud platform. * Deep expertise in observability tooling, incident management, and on-call practices. * Strong scripting and automation skills (Python, Go, Bash, or similar). * Excellent communication skills - you will work directly with customer engineering, security, and governance stakeholders and represent Blitzy on the account., * Experience deploying or operating software in government-accredited or similarly certified cloud environments. * Familiarity with handling sensitive data and security frameworks common to regulated industries. * Experience supporting AI/ML workloads or their supporting infrastructure. * Prior forward-deployed, residency, or embedded-engineer experience at an enterprise customer site. * Prior experience in a high-growth startup environment where you wore multiple hats. ## Description As a Site Reliability Engineer on Blitzy's Public Sector team, you will be the backbone of our platform's reliability and operational excellence for a dedicated enterprise customer opportunity in a highly regulated industry. You'll deploy and operate Blitzy's self-hosted platform within the customer's secure cloud environment, serving as Blitzy's embedded engineer on the account. You'll work at the intersection of software engineering, infrastructure, and customer success, ensuring our AI-powered development platform remains highly available and performant in one of the most demanding security environments in enterprise software. This is a high-impact, hands-on role for an engineer who thrives in a fast-moving environment and takes deep ownership of the systems they operate. What Success Looks Like * In 30 days: You have a deep understanding of Blitzy's self-hosted deployment architecture, have begun customer onboarding, and are actively supporting deployment planning alongside customer infrastructure teams. * In 90 days: You are operating inside the customer environment, have stood up and hardened the deployment, and have established monitoring, alerting, and incident response workflows that meet the customer's security requirements. * In 6 months: The platform is running reliably at scale with defined SLOs and validated capacity headroom, and you are the trusted technical voice for the customer's infrastructure, security, and platform teams. Areas of Ownership * Deploy, operate, and maintain Blitzy's self-hosted platform within a customer-controlled, secure cloud environment. * Own the Kubernetes-based deployment: releases, upgrades, capacity planning, and performance benchmarking for compute-intensive AI workloads. * Design and maintain observability - logging, metrics, tracing, and alerting - that operates fully within the customer's security boundary. * Serve as Blitzy's on-account technical presence: partner with customer infrastructure, security, and governance teams on provisioning, reviews, documentation, and operational escalations. * Handle sensitive customer data in accordance with customer security requirements; champion security best practices across the deployment. * Feed lessons learned back into Blitzy's product and infrastructure roadmap to strengthen our self-hosted offering for future public sector customers. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Biggest German Tech Companies](https://www.wearedevelopers.com/magazine/424-the-biggest-german-tech-companies)