> Markdown version of [/jobs/ext/2042977-site-reliability-engineer-ii](https://www.wearedevelopers.com/jobs/ext/2042977-site-reliability-engineer-ii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer II - **Company:** Kastle Systems - **Location:** Falls Church, VA, United States - **Experience:** Experienced - **Salary:** $125,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Bash Shell, C Sharp (Programming Language), Computer Networks, System Configuration, Software Debugging, Linux, Distributed Systems, Domain Name System (DNS), Hypertext Transfer Protocols (HTTP), Python (Programming Language), Linux Kernel, Performance Tuning, Role-Based Access Control, Reliability Engineering, Prometheus, Software Engineering, SQL Databases, TCP/IP, Pulumi, Scripting, Transport Layer Security, Cloud Monitoring, Delivery Pipeline, Grafana, Mttr, Kubernetes, Deployment Automation, Cosmos DB, Terraform, Key Vault - **Published:** August 13, 2026 - **Apply:** https://www.disabledperson.com/jobs/74221471-site-reliability-engineer-ii ## About the Role This is a mid-level individual contributor role. You are expected to execute technical work independently, drive reliability improvements end-to-end, and participate meaningfully in architecture discussions. You will carry on-call responsibilities as part of a shared rotation with a well-defined escalation model and a strong blameless post-incident review culture., * Experience: 4-6 years in an SRE, Platform Engineering, or Infrastructure Engineering role, with demonstrated ownership of production systems. * Cloud - Azure: Hands-on experience managing production infrastructure in Azure: AKS, Azure Container Registry, Azure Monitor, Cosmos DB, Key Vault, Azure Front Door, or equivalent services. AWS/GCP backgrounds considered with clear willingness to operate in Azure. * Kubernetes: Deep operational experience with Kubernetes in production: resource management, network policies, RBAC, HPA/VPA, persistent volumes, and debugging live workload issues. * GitOps & Release Tooling: Experience with ArgoCD, Flux, or equivalent GitOps deployment tools. Familiarity with multi-stage progressive delivery and approval gate patterns is a strong plus. * Infrastructure as Code: Proven track record with Terraform, OpenTofu, or Pulumi in a production GitOps context - not just writing HCL, but maintaining drift-free state and managing state backends safely. * Observability: Hands-on configuration of Prometheus, Grafana, OpenTelemetry, and/or ELK/OpenSearch. Ability to go from symptom to instrumentation to dashboard without hand-holding. * Programming & Scripting: Proficiency in Python or Go for automation and tooling; strong Bash scripting. Ability to read and reason about application code when debugging production issues. Proficiency in C# and SQL for reviewing deliverables and participating in triage. * Linux & Networking: Solid understanding of Linux internals, TCP/IP, DNS, TLS, and HTTP semantics. Comfortable debugging at the network and OS layer., * Experience with Crossplane or other Kubernetes-native infrastructure operators. * Familiarity with feature flag platforms (LaunchDarkly, Flagsmith, or similar) and gradual rollout strategies. * Background in IoT, physical security, access control, or other latency-sensitive, event-driven domains. * Comfort with async collaboration across distributed time zones (US + India team structure). * Experience with AI-assisted development tooling and an appetite to incorporate it into engineering workflows. * Knowledge of CMMC 2.0, SOC 2, or FedRAMP compliance postures as they apply to infrastructure and access control. ## Description The SRE II sits at the intersection of software engineering and platform operations. You will own the reliability, scalability, and operational hygiene of Kastle's core infrastructure - engineering away toil, hardening deployment pipelines, and partnering with product engineering teams to make new services production-ready from day one., The team is in the middle of a meaningful platform evolution: formalizing multi-tier release pipelines (Dev * QA * Integration * UAT * Prod) with ArgoCD-based approval gates, building out SLI/SLO frameworks, and migrating toward full GitOps. You will be a hands-on contributor to all of it., Release Engineering & GitOps * Own and evolve the multi-stage deployment pipeline using ArgoCD, including approval gates, promotion policies, and rollback mechanisms. * Maintain trunk-based branching discipline and enforce release governance standards across the engineering organization. * Manage feature flag lifecycle - from creation and gradual rollout to deprecation - in coordination with product and QA teams. * Build and maintain CI/CD pipelines that enable safe, frequent, and auditable deployments., * Provision and manage Azure infrastructure using Terraform or OpenTofu, maintaining drift-free state aligned with GitOps principles. * Own Kubernetes cluster operations including workload scheduling, resource optimization, RBAC, network policy, and cost governance. * Identify and act on infrastructure cost optimization opportunities (compute rightsizing, storage tier selection, idle resource elimination). * Support Crossplane or similar operator patterns for Kubernetes-native infrastructure management where applicable. Reliability & Observability * Define, instrument, and enforce SLIs and SLOs in partnership with product engineering teams. * Build and maintain observability infrastructure - metrics, logs, and distributed traces - using Prometheus, Grafana, OpenTelemetry, or equivalent tooling. * Conduct proactive capacity planning and performance tuning across multi-tenant, distributed environments. * Establish and maintain runbooks, dashboards, and alerting policies that reduce cognitive overhead during incidents. Incident Management * Participate in shared on-call rotation covering core platform and infrastructure services; on-call load is balanced across the team with structured handoff practices. * Lead mitigation of live production incidents with a focus on minimizing MTTR and clear stakeholder communication under pressure. * Facilitate blameless post-incident reviews and drive preventative engineering to closure - not just documentation. Engineering Partnership * Embed with product engineering teams during design and architecture phases to establish reliability, scalability, and security requirements before code is written. * Maintain clear, comprehensive documentation for infrastructure architecture, operational procedures, and onboarding guides. * Push back constructively when proposed designs compromise reliability or operability, proposing alternatives rather than just raising concerns. ## Related Videos - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) - [Building a Cloud Platform Where Everything is Just Another Kubernetes Resource](https://www.wearedevelopers.com/videos/100137-building-a-cloud-platform-where-everything-is-just-another-kubernetes-resource) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus)