> Markdown version of [/jobs/ext/1394371-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1394371-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** ANNA LLC - **Location:** Huntsville, AL, United States - **Experience:** Experienced - **Salary:** $130,000.0 - $160,000.0 - **Contract:** Permanent contract - **Skills:** DevOps, Amazon DynamoDB, Monitoring of Systems, Reliability Engineering, Prometheus, Datadog, GitHub Copilot, Grafana, Mttr, Kubernetes Helm Charts, Kubernetes, Apache Kafka, Terraform, Serverless Computing - **Published:** July 23, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=370ed146f02168df ## About the Role * 3-5+ years in a site reliability, DevOps, or platform engineering role * Strong Kubernetes operational experience (EKS preferred) - you've troubleshot pod failures, node issues, and deployment rollbacks in production * Experience with monitoring and alerting platforms (Datadog preferred; Prometheus/Grafana acceptable) * Familiarity with incident response processes - you've responded to production incidents * Ability to read and understand Terraform and Helm charts (you don't need to author complex modules, but you need to navigate them) * Comfortable with GitOps workflows (ArgoCD or Flux) * Strong written communication - runbooks, post-incident reviews, and documentation are core deliverables of this role Preferred * Kafka operational experience (consumer lag diagnosis, broker health, partition rebalancing) * Knative or serverless-on-Kubernetes experience * DynamoDB operational familiarity (throttling, capacity management, GSI patterns) * Experience with AI-assisted development tools (e.g. Kiro, GitHub Copilot, or similar) and willingness to integrate AI tooling into operational workflows * Healthcare or compliance-adjacent environment experience (HITRUST, SOC 2, HIPAA) * Experience defining and implementing SLOs/SLIs ## Description We are a customer experience technology company serving enterprise clients across regulated industries including healthcare and financial services. Our platform is growing, and we need an SRE to help us build operational maturity, reduce single-point-of-failure risk, and establish structured incident response capabilities as we scale. This is a foundational hire. You will work alongside senior infrastructure leadership to take operational ownership of our Kubernetes platform, strengthen monitoring and alerting coverage, develop runbooks, and drive continuous reliability improvement. What You'll Do Incident Response * Triage and respond to escalated platform issues, working collaboratively with engineering and infrastructure teams to identify root causes and drive resolution * Develop and maintain operational runbooks and playbooks, contributing to a growing knowledge base alongside senior staff Monitoring & Observability * Own the monitoring and alerting stack (Datadog) - identify coverage gaps, tune alert thresholds, reduce noise, and eliminate client-discovered outages * Track and improve MTTR, availability, and change failure rate metrics * Define and propose SLOs for client-facing services based on actual observability data * Build dashboards that give leadership visibility into platform health Platform Operations * Operate EKS clusters day-to-day: investigate sync failures, pod health issues, node problems, failed deployments * Operational ownership of Kafka (Strimzi - consumer lag, broker health), Knative (autoscaler, scaling events), and DynamoDB-backed services * Monitor cross-account drift as new tenant accounts come online * Support tenant onboarding from an operational readiness perspective Collaboration & Growth * Work closely with VP of Infrastructure during ramp period to absorb platform knowledge * Document operational procedures and contribute to the team knowledge base * Mentor junior team members as the team grows * Contribute to post-incident reviews and drive continuous improvement ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [The OpenTelemetry mistakes I keep seeing (and how to stop making them)](https://www.wearedevelopers.com/videos/100158-the-opentelemetry-mistakes-i-keep-seeing-and-how-to-stop-making-them) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)