Staff Infrastructure Engineer

Ai, Inc
San Francisco, CA, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$307,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Amazon S3 Cloud Computing Code Review Continuous Integration Data Infrastructure Identity and Access Management Python (Programming Language) Ruby Software Engineering TypeScript
+8 more
Pulumi Amazon Virtual Private Cloud (VPC) Amazon Relational Database Service Kubernetes AWS Fargate Functional Programming Software Coding Golang

Job description

Kikoff’s infrastructure team builds the systems that enable engineering teams to move quickly without sacrificing reliability, security, or cost discipline. The team owns five connected areas: Observability, Developer Productivity, Compute Infrastructure, Networking and Storage, and Data Infrastructure.

As a Staff Infrastructure Engineer, you will own high-leverage infrastructure problems from design through production operation and help run infrastructure as a product for Kikoff engineers. You will build internal products and paved paths that turn ambiguous problems into durable systems. Your work should improve delivery speed and reliability, reduce cost, and reduce the amount of operational work product teams carry.

This is a hands-on Staff IC role. You will foster relationships, write code, review designs, make architecture decisions, lead through incidents, and help engineers solve problems outside the runbook. You will have a primary area of depth and enough range to follow production problems across infrastructure boundaries., * Design and implement self-service infrastructure on AWS using reusable code and infrastructure-as-code patterns. We use Pulumi, HCL, and TypeScript heavily. The company runs on Ruby.

  • Own the systems you build in production, including reliability, security, capacity, cost, upgrades, incidents, and recovery.
  • Give critical services an SLO, actionable alerts, a useful dashboard, a runbook, and a tested recovery path.
  • Automate recurring operational work and eliminate failure-prone manual steps rather than allowing them to become permanent processes.

Run Infrastructure as a Product

  • Work directly with engineers to turn recurring friction into paved paths, self-service tools, and automated workflows that are faster and safer than one-off solutions.
  • Measure outcomes for internal customers through adoption, developer feedback, delivery speed, reliability, cost, toil, and on-call load.
  • Use the fastest responsible path when a team is blocked, then turn recurring friction into automation or a durable platform capability.

Set Technical Direction

  • Set technical direction for ambiguous infrastructure work, then carry it from problem framing and design through implementation, rollout, and production ownership.
  • Make clear trade-offs among delivery speed, reliability, security, cost, and long-term operational complexity.
  • Partner across Infrastructure, Security, Data, and Product Engineering to build security and compliance controls into normal engineering workflows.

Raise the Engineering Bar

  • Review code and designs, challenge weak assumptions, help other engineers make better technical decisions, and uphold high standards for reliability, testing, safe deployments, security, and maintainability.
  • Lead through incidents and unfamiliar failure modes, and ensure the system is better after service is restored.
  • Create standards, reusable patterns, and durable documentation that improve how teams build, deploy, observe, and operate software.

Requirements

  • 7+ years of infrastructure, platform, or software engineering experience, or an equivalent record of Staff-level technical impact.
  • Sustained ownership of a consequential production system. You have operated what you built, handled incidents and unfamiliar failure modes, and improved the system afterward.
  • Strong coding skills in TypeScript, Python, Go, Ruby, or a similar language. You can build production software, not only configure tools.
  • Production experience with AWS, infrastructure as code, containers, CI/CD, observability, and the security boundaries around them.
  • Experienced generalist judgment. You have depth in at least one infrastructure domain and can learn quickly across adjacent ones.
  • A track record of building platform capabilities that engineers adopted because they solved real problems and were easier to use than one-off alternatives.
  • Clear, direct communication. You can explain a technical decision, its evidence, and its tradeoffs to engineers and non-technical partners., * Deep experience in one or more of observability, developer productivity, cloud compute, networking, storage, data infrastructure, or security.
  • Experience with Pulumi and TypeScript on AWS, including ECS or Fargate, k8s, RDS, S3, Lambda, VPC, and IAM.
  • Experience building and operating an early-stage platform through rapid growth in fintech or another regulated environment.

About the company

Scrappy. We move quickly and build what we need. We do not cut corners when they matter, and we do not over-engineer when they do not.

Risk-oriented. We make tradeoffs deliberately. A mature team knows the difference between a risk worth taking and one that is not.

Data-obsessed. We look at the data, understand the mechanics behind it, and use that evidence to make better decisions.

Humble. We know the company has benefited from timing, circumstance, and the right people showing up. We are grateful, and we do not take it for granted. Base Range

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

50 sec

Why developer happiness matters in web frameworks

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

3:20 min

Overview of infrastructure as code tools

Alexander Bubeck · World Congress 2023

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

Videos

See all

Related articles

See all