> Markdown version of [/jobs/ext/2703942-software-engineer-cloud-infrastructure](https://www.wearedevelopers.com/jobs/ext/2703942-software-engineer-cloud-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Cloud Infrastructure - **Company:** DECAGON, LLC - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $200,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Continuous Integration, DevOps, Distributed Systems, Domain Name System (DNS), Fault Tolerance, Identity and Access Management, Role-Based Access Control, Prometheus, Datadog, Load Balancing, Cloud Platform System, Grafana, Amazon Virtual Private Cloud (VPC), Kubernetes, Low Latency, Deployment Automation, Terraform - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-software-engineer-cloud-infrastructure-decagon-8737420 ## About the Role * 4+ years building and operating core infrastructure, platform engineering, or infrastructure/DevOps, ideally with some customer-facing deployment experience. * Deep experience with a major cloud provider (GCP, AWS, or Azure), along with Terraform and Kubernetes at scale. * Strong grasp of cloud networking fundamentals (VPCs, IAM, DNS, load balancing) and how they surface as real deployment constraints. * A track record operating production systems reliably: monitoring, on-call, incident response, and reasoning about failure modes up front. * Comfort navigating ambiguity across a range of stakeholders, from engineers to security and compliance teams, and turning those conversations into actionable plans. * Clear technical writing and a track record of driving adoption across teams. * Comfortable in a fast-moving environment with rapid change. Even better if you have * Experience managing deployments in customer-owned cloud environments, including security reviews, compliance requirements, and change management. * Experience building internal platforms or paved roads: service templates, self-serve environments, CI/CD pipeline design, deployment automation. * Familiarity with observability and incident management in distributed systems (Prometheus, Grafana, Datadog, or similar). * Infrastructure-as-code with a security-minded approach to supply chain (provenance, secrets, least privilege). * Experience operating latency-sensitive or ML/AI-serving workloads in production. * Experience using AI-assisted tooling to make yourself and your team dramatically more effective. ## Description The Infrastructure team builds and operates the foundations that power Decagon: networking, data, ML serving, developer platform, and real-time voice. We partner closely with product, data, and ML to deliver high-scale, low-latency systems with clear SLOs and great developer ergonomics. Read more about the infra team's work here: https://decagon.ai/blog/what-an-air-gapped-ai-deployment-actually-requires About the Role Decagon builds agentic AI that resolves customer support conversations end to end, for companies ranging from fast-growing startups to some of the largest financial institutions in the world. Keeping that system fast, reliable, and secure (across our multi-tenant cloud and inside customers' own locked-down environments) is an infrastructure problem, and that's the problem this role owns. You'll build the platforms and abstractions our product teams ship on, and you'll architect and operate the deployments that run our agents inside enterprise customer clouds, where security, compliance, and operational rigor matter as much as speed. The work is core infrastructure at heart: reliability, CI/CD, deployment automation, on-call. You'll be joining a team with real infrastructure and momentum already in place, but there's no established playbook for running agentic systems at this scale, so a big part of the job is figuring it out as the technology shifts and usage grows by orders of magnitude. What you'll do * Build the platform + Design the development and production platforms that power our products, and the abstractions over cloud infrastructure, Kubernetes, and networking that let engineers ship without becoming infrastructure experts. + Make sure it all scales to the next order of magnitude as usage grows. * Own enterprise deployments + Take end-to-end ownership of deployment architecture in customer-owned cloud environments (VPC configuration, permissioning, networking, provisioning) and the full lifecycle that follows: setup, upgrades, scaling, and incident support. + Build the runbooks and automation that make it repeatable. * Keep agentic workloads reliable + Treat monitoring, alerting, and rollback as first-class parts of anything you ship, not afterthoughts. + Own the reliability of the systems our AI agents depend on in production, where latency, availability, and graceful degradation directly shape the customer experience. * Partner across boundaries + Work directly with customers' platform, security, and DevOps teams to navigate their infrastructure and compliance constraints, and with our Product, Security, Sales, and Customer Success teams to turn customer requirements into concrete deployment plans. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)