> Markdown version of [/jobs/ext/1745072-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1745072-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** SecurityScorecard - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $152,000.0 - $195,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Automation of Tests, Bash Shell, Cloud Computing, Continuous Integration, DevOps, Github, Python (Programming Language), Octopus Deploy, Open Source Technology, Reliability Engineering, Prometheus, Datadog, Policy as Code, Pulumi, Large Language Models, Grafana, Gitlab-ci, Kubernetes, Apache Flink, Apache Kafka, Machine Learning Operations, Vertica, Terraform, Jenkins, Vulnerability Analysis - **Published:** July 14, 2026 - **Apply:** https://job-boards.greenhouse.io/securityscorecard/jobs/8062312 ## About the Role * 6+ years in SRE, DevOps, or Infrastructure roles, with significant production Kubernetes experience. * Hands-on experience integrating AI/LLM tooling into engineering or operational workflows (e.g., MCP servers, AI agents acting on infrastructure), and a clear grasp of the security and governance considerations of giving AI access to production. * Proven success building CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, or similar). * Strong with Kubernetes internals and managed services like EKS, GKE, or AKS. * Expertise with Infrastructure as Code (Terraform, Helm, Pulumi) and GitOps. * Proficient in Python, Bash, or Go. * Knowledge of observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry). * Production experience with Kafka, Flink, and ClickHouse. * Strong communication and cross-team collaboration skills., * Multi-region or multi-cluster Kubernetes experience. * Chaos engineering or resilience testing. * Security scanning, compliance automation, or policy-as-code. * LLM observability/tracing tooling (Langsmith, Langfuse) or MLOps workflows. * Contributions to open-source Kubernetes or CI/CD projects. ## Description As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization of our Kubernetes-based infrastructure and CI/CD systems. You will also own the infrastructure behind our AI tooling - building MCP servers and defining safe, auditable AI access patterns for production systems. You'll work hands-on with engineering teams to accelerate delivery, ensure production reliability, and embed best practices for automation, observability, and resilience., * Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications. * Build and operate AI tooling infrastructure - stand up MCP servers and establish secure, governed AI access and guardrails for production systems. * Optimize and maintain CI/CD pipelines, improving reliability, speed, and rollback safety. * Implement progressive delivery strategies such as blue/green and canary deployments. * Advance Infrastructure as Code with Terraform, Helm, and Argo CD, defining reusable patterns for the org. * Operate and optimize streaming and analytics infrastructure: Kafka, Flink, and ClickHouse. * Build automated testing into the CI/CD lifecycle. * Improve system observability - define SLOs, alerts, and dashboards. * Lead incident response and postmortems, focusing on root cause and durable fixes. * Mentor engineers across teams on Kubernetes, CI/CD, and cloud infrastructure. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Terraform for Developers](https://www.wearedevelopers.com/videos/3-terraform-for-developers) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 138 - Are you secure about this?](https://www.wearedevelopers.com/magazine/486-dev-digest-138-are-you-secure-about-this) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss)