> Markdown version of [/jobs/ext/3464319-member-of-technical-staff-cloud-infrastructure](https://www.wearedevelopers.com/jobs/ext/3464319-member-of-technical-staff-cloud-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Member of Technical Staff - Cloud Infrastructure - **Company:** RUNTA INC. - **Location:** San Mateo, CA, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Business Process Model and Notation, Cloud Computing, Continuous Integration, Distributed Systems, Network Control, Web Traffics, Enterprise Software Applications, Cloud Platform System, Autoscaling, Large Language Models, Multi-Cloud, Containerization, Kubernetes, AWS Fargate, Machine Learning Operations, Terraform, Serverless Computing - **Published:** September 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=cdcaa55bc6ec9d5e ## About the Role * You've designed and operated production AWS infrastructure you were accountable for, not just worked within someone else's * Deep Terraform (or equivalent IaC) experience: complex modules built for other engineers to consume, state management at scale, infra CI/CD * Multi-cloud or cross-cloud experience: provider-spanning abstractions, workload migrations, or internal platforms built on IaC * Experience running serverless or container platforms at scale (Lambda, Fargate, ECS, EKS) * A cost-optimization track record you can talk through in specifics: RI strategy, Spot fleet design, rightsizing with real numbers * Solid distributed systems fundamentals: consistency models, replication, consensus * You understand orchestration platforms at the why level, not just the kubectl level Nice to have: background in AI/ML infrastructure, agent frameworks, or LLM serving. You Should Genuinely Care About AI ## Description Our runtime team builds the execution engine. You build the cloud platform it runs on: the production infrastructure that serves paying customers. You are the person accountable for it working, scaling, and not bankrupting us. Runta does three things: run, govern, record. This role makes run economical at scale, gives record a storage layer that's cheap at volume and fast on restore, and owns the network boundaries that make govern enforceable in production. You decide ECS vs. EKS, when Spot instances make sense, how snapshot storage should be tiered, and what the architecture costs at 10x load. You'll also own the infrastructure-as-code layer that makes our platform reproducible and portable. Writing Terraform is table stakes here. The job is designing modules other engineers consume, managing state at scale, and building abstractions that keep multi-cloud on the table without over-engineering for it today. This is an ownership role, not a support role. The infrastructure layer is part of the product: customer-facing, versioned, reproducible. What You'll Do * Design and operate Runta's production infrastructure on AWS: compute, networking, storage, observability * Own the IaC layer: Terraform modules, state management, CI/CD for infrastructure changes * Design cross-cloud abstractions that don't lock us into a single provider * Own the infrastructure side of snapshot/resume/branch: snapshot graph storage, control plane / data plane separation, local cloud state consistency * Make cost-aware architecture decisions: instance strategy, Spot fleet design, storage tiering, rightsizing * Shape the infrastructure roadmap as an early technical leader, working directly with the founder The Problems You'd Be Solving * A workload runs anywhere from 2 seconds to 2 hours, unpredictably. How do you size, schedule, and price the fleet under it? * Agents checkpoint constantly. Snapshot storage must be cheap at petabyte volume but restore in seconds. What's the tiering design? * An agent job fans out 50 concurrent sub-tasks, then the whole fleet goes idle for 10 minutes. Autoscaling policies tuned for web traffic fall over. What replaces them? * Agents retry non-deterministically on failure. Where must the infrastructure itself be idempotent, and where is eventual consistency a silent-corruption risk? * The same platform needs to deploy identically across environments, and someday across clouds. What do you abstract, and what stays provider-specific? If you read those and started sketching answers, we want to talk to you., Our customers are AI builders. The decisions you make (cold-start latency, snapshot frequency, concurrency limits) determine what AI products are possible to build on Runta. The best fit is someone who already uses AI tools daily, tinkers with agents on their own time, and has real curiosity about where agentic systems are going. If AI is "just another workload" to you, this probably isn't the right role., * Your orchestration background is enterprise application workflow (BPMN, Spring-based orchestration) rather than execution infrastructure * You're looking for a role where the architecture is already settled and the job is keeping it running How We Work Small team, high trust, high ownership. We value people who are hungry, humble, and sharp. We look for clear communicators who hold strong opinions loosely and use AI-native workflows in their daily engineering. You'll work directly with the founder, alongside the core runtime and isolation teams. ## Related Videos - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [Building Applications with Infrastructure as Code](https://www.wearedevelopers.com/videos/887-building-applications-with-infrastructure-as-code) - [Terraform for Developers](https://www.wearedevelopers.com/videos/3-terraform-for-developers) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [ I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose.](https://www.wearedevelopers.com/magazine/773-i-gave-a-video-editor-more-autonomy-than-a-trading-bot-on-purpose) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)