> Markdown version of [/jobs/ext/2000265-senior-software-engineer-infrastructure](https://www.wearedevelopers.com/jobs/ext/2000265-senior-software-engineer-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, Infrastructure - **Company:** Docker - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $160,900.0 - $260,700.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Computing, Continuous Integration, Software Debugging, Software Design Documents, Linux, Github, Octopus Deploy, Reliability Engineering, Cloud Services, Prometheus, Runbook, Software Engineering, Istio, Grafana, Core Api, Backend, Kubernetes, Terraform, Docker - **Published:** August 9, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/paepjscc6w ## About the Role * 6+ years of hands-on software engineering in backend, infrastructure, or platform engineering, though we weight real depth and impact more heavily than exact tenure. * Strong software engineering in Go or a similar language: design, testing, debugging, review, and long-term maintainability. * A track record of building, shipping, and operating cloud services or infrastructure in production. * Deep expertise in at least one of Kubernetes, networking, cloud platforms, reliability engineering, or developer platforms - plus solid Linux and production-ops fundamentals. * Experience shaping technical direction and working effectively across teams. * Clear written and verbal communication in a remote environment (RFCs, design docs, incident writeups). * Bachelor's in CS/Engineering or equivalent practical experience. Nice to have: EKS and ingress/CNI/service-mesh experience; observability with OpenTelemetry/Prometheus/Grafana; CI/CD and progressive delivery (GitHub Actions, Argo CD, canaries); driving migrations or adoption programs across teams. _You don't need every item here. We value strong systems judgment, real depth in at least one area, and curiosity across the rest. ## Description This is a senior, deeply hands-on role. You'll own significant components, drive projects from design through production adoption, and shape the platform's technical direction while helping teammates grow. Concretely, you will: * Turn ambiguous infrastructure problems into clear designs and working systems - contributing to RFCs and architecture reviews, and driving your projects to done. * Build self-service capabilities and platform APIs (primarily in Go) for onboarding, provisioning, deployment, observability defaults, and day-2 operations - with contracts and docs teams actually use. * Apply and help shape delivery standards with Terraform, GitOps on Argo CD, progressive rollout, and strong testing - including the continuous-deployment flow we're missing today. * Strengthen the multi-tenant EKS foundations for reliability, security, scale, and cost: Envoy Gateway ingress, traffic routing, and multi-region, cross-account connectivity. * Improve SLOs, alerting, and incident follow-up on Grafana Cloud so production gets safer and less dependent on heroics. We measure this work by outcomes the consuming teams feel: how fast they can provision and ship, how much they can do without us, and how reliably it all runs. AI-assisted operations We're Investing In AI-assisted And Agentic Workflows To Cut Operational Toil, And We Care That They Stay Safe, Auditable, And Human-reviewed. You'll Help Shape Where They Earn Their Place And Where They Don't. Early Targets * Alert enrichment and incident context-gathering: assembling the relevant signals, history, and runbook so the on-call engineer starts with context instead of a blank page. * Runbook-assisted diagnosis and remediation recommendations, with a human in the loop on anything that changes production. * Onboarding and readiness assistants that answer the questions our experts answer today. If you've built operational automation and have a healthy skepticism about where automation belongs, this is a place to put both to work. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Building AI Solutions with Rust and Docker](https://www.wearedevelopers.com/magazine/494-building-ai-solutions-with-rust-and-docker) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read)