Member of Technical Staff - Cloud Infrastructure
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+7 more
Job description
Our runtime team builds the execution engine. You build the cloud platform it runs on: the production infrastructure that serves paying customers. You are the person accountable for it working, scaling, and not bankrupting us.
Runta does three things: run, govern, record. This role makes run economical at scale, gives record a storage layer thatâs cheap at volume and fast on restore, and owns the network boundaries that make govern enforceable in production. You decide ECS vs. EKS, when Spot instances make sense, how snapshot storage should be tiered, and what the architecture costs at 10x load.
Youâll also own the infrastructure-as-code layer that makes our platform reproducible and portable. Writing Terraform is table stakes here. The job is designing modules other engineers consume, managing state at scale, and building abstractions that keep multi-cloud on the table without over-engineering for it today.
This is an ownership role, not a support role. The infrastructure layer is part of the product: customer-facing, versioned, reproducible.
What Youâll Do
- Design and operate Runtaâs production infrastructure on AWS: compute, networking, storage, observability
- Own the IaC layer: Terraform modules, state management, CI/CD for infrastructure changes
- Design cross-cloud abstractions that donât lock us into a single provider
- Own the infrastructure side of snapshot/resume/branch: snapshot graph storage, control plane / data plane separation, local cloud state consistency
- Make cost-aware architecture decisions: instance strategy, Spot fleet design, storage tiering, rightsizing
- Shape the infrastructure roadmap as an early technical leader, working directly with the founder
The Problems Youâd Be Solving
- A workload runs anywhere from 2 seconds to 2 hours, unpredictably. How do you size, schedule, and price the fleet under it?
- Agents checkpoint constantly. Snapshot storage must be cheap at petabyte volume but restore in seconds. Whatâs the tiering design?
- An agent job fans out 50 concurrent sub-tasks, then the whole fleet goes idle for 10 minutes. Autoscaling policies tuned for web traffic fall over. What replaces them?
- Agents retry non-deterministically on failure. Where must the infrastructure itself be idempotent, and where is eventual consistency a silent-corruption risk?
- The same platform needs to deploy identically across environments, and someday across clouds. What do you abstract, and what stays provider-specific?
If you read those and started sketching answers, we want to talk to you., Our customers are AI builders. The decisions you make (cold-start latency, snapshot frequency, concurrency limits) determine what AI products are possible to build on Runta. The best fit is someone who already uses AI tools daily, tinkers with agents on their own time, and has real curiosity about where agentic systems are going. If AI is âjust another workloadâ to you, this probably isnât the right role., * Your orchestration background is enterprise application workflow (BPMN, Spring-based orchestration) rather than execution infrastructure
- Youâre looking for a role where the architecture is already settled and the job is keeping it running
How We Work
Small team, high trust, high ownership. We value people who are hungry, humble, and sharp. We look for clear communicators who hold strong opinions loosely and use AI-native workflows in their daily engineering. Youâll work directly with the founder, alongside the core runtime and isolation teams.
Requirements
- Youâve designed and operated production AWS infrastructure you were accountable for, not just worked within someone elseâs
- Deep Terraform (or equivalent IaC) experience: complex modules built for other engineers to consume, state management at scale, infra CI/CD
- Multi-cloud or cross-cloud experience: provider-spanning abstractions, workload migrations, or internal platforms built on IaC
- Experience running serverless or container platforms at scale (Lambda, Fargate, ECS, EKS)
- A cost-optimization track record you can talk through in specifics: RI strategy, Spot fleet design, rightsizing with real numbers
- Solid distributed systems fundamentals: consistency models, replication, consensus
- You understand orchestration platforms at the why level, not just the kubectl level
Nice to have: background in AI/ML infrastructure, agent frameworks, or LLM serving.
You Should Genuinely Care About AI
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Trustworthy AI Starts at Deployment: 5 Checks Before You Ship
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose.
From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path