> Markdown version of [/jobs/ext/3278588-senior-site-reliability-engineer-golang-kubernetes](https://www.wearedevelopers.com/jobs/ext/3278588-senior-site-reliability-engineer-golang-kubernetes). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer (Golang / Kubernetes) - **Company:** Mirantis, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, InfiniBand, Python (Programming Language), Reliability Engineering, Prometheus, Software Engineering, Datadog, Data Logging, Grafana, Build Management, Kubernetes, Bare Metal, Golang - **Published:** September 1, 2026 - **Apply:** https://jobs.smartrecruiters.com/Mirantis/744000146541630-senior-site-reliability-engineer-golang-kubernetes- ## About the Role * 5+ years in SRE, platform reliability, or a closely related software/infrastructure role. * Strong software engineering skills (e.g., Go or Python) with experience building and operating APIs or services in production. * Demonstrated experience defining SLIs/SLOs and error budgets for real production systems. * Hands-on experience with observability tooling - metrics, logging, and tracing (e.g., Prometheus/VictoriaMetrics, OpenTelemetry, Grafana). * Solid understanding of Kubernetes and the signals it and its workloads emit. * Strong written and verbal communication with technical audiences. Preferred * Experience instrumenting or monitoring bare-metal and NVIDIA infrastructure (BMC/Redfish, InfiniBand, NVLink, UFM). * Experience with the Mirantis K0rdent stack (K0rdent Enterprise, K0rdent AI, KOF) and Cluster API. * Familiarity with VictoriaMetrics/VictoriaLogs at scale. * Proven experience in sovereign or high-security air-gapped environments. ## Description Define what reliability means for a GPU-accelerated AI platform and make it measurable. You will own the service-level indicators and objectives for the K0rdent Observability Framework (KOF) - deriving meaningful SLIs from the signals the platform already emits, and exposing them to Platform Administrators through a clean API. Work spans hybrid, edge, and air-gapped deployments built on the Mirantis K0rdent stack., We are looking for a Senior SRE who thinks past dashboards to the contract between a platform and its operators. The right candidate can look at raw telemetry from Kubernetes, bare metal, and NVIDIA infrastructure, decide which signals actually predict user-visible reliability, and turn them into SLIs and SLOs that operators can act on. You are equally comfortable writing the service that exposes those SLIs through an API and reasoning about error budgets, alerting quality, and signal-to-noise. You should be self-directed, able to own reliability definitions end to end, and effectively communicate them across teams., * Define SLIs and SLOs based on the signals available across the platform - Kubernetes, bare-metal hosts, and NVIDIA infrastructure (BMC, InfiniBand, NVLink, UFM). * Design and build the API that exposes SLIs and reliability state to Platform Administrators and downstream systems. * Establish alerting and error-budget practices that maximize signal and minimize noise. * Partner with infrastructure, storage, and networking teams to ensure the right signals are instrumented and collected. * Diagnose reliability and performance issues across the observability stack and drive their resolution. ## Related Videos - [Monitoring as Code - Managing your dashboards at scale](https://www.wearedevelopers.com/videos/753-monitoring-as-code-managing-your-dashboards-at-scale) - [The Memory Leak That Ate Our Cluster: A Postmortem](https://www.wearedevelopers.com/videos/2057-the-memory-leak-that-ate-our-cluster-a-postmortem) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)