Cloud Infrastructure Engineer

ONE STOP COLLECTIBLE CORP
San Jose, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Artificial Intelligence Amazon Web Services Bash Shell Cloud Computing Linux Domain Name System (DNS) Identity and Access Management IP Addressing Python (Programming Language) Key Management Network Architecture Routing
+13 more
Peering Software Architecture Role-Based Access Control Reliability Engineering Prometheus Management of Software Versions Private Cloud Environment Transport Layer Security Load Balancing Grafana Firewalls (Computer Science) Kubernetes Terraform

Job description

We run a production cloud platform on Google Cloud today and are expanding our AWS footprint. We need a senior infrastructure engineer to help establish a deliberate multi-cloud foundation and define how GCP and AWS should coexist as the platform grows.

You will own the Cloud Foundations layer: cloud account and project structure, infrastructure-as-code, state strategy, networking, Kubernetes foundations, and implementation of cloud access controls. You will help decide where Trener should standardize across providers and where provider-specific architecture is the right boundary.

This is a senior individual contributor role. You will not be managing people, and you will not be inheriting a greenfield.

What You Will Own

  • Cloud foundations across GCP and AWS - project and account topology, landing zones, organization policies, baseline guardrails, and enrollment of new environments.
  • Cloud and cross-cloud networking - VPCs, routing, peering, DNS, load balancing, firewall policy, IP addressing, and private connectivity between cloud environments.
  • Infrastructure as code - reusable Terraform/OpenTofu modules, environment composition, lifecycle management, and patterns that keep infrastructure understandable as the estate grows.
  • State strategy and blast-radius boundaries - how infrastructure state is partitioned, how dependencies are expressed, and how those patterns evolve across providers.
  • Cloud identity implementation - implement cloud IAM, workload identity, federation, and access controls in partnership with Security.
  • Kubernetes foundations - cluster lifecycle, baseline configuration, upgrades, and shared infrastructure services beneath application workloads.
  • Infrastructure cost visibility - implement the tagging, labeling, budgets, alerts, and reporting needed to make cloud spend understandable and surface obvious infrastructure waste.
  • Infrastructure controls supporting our SOC 2 program - resource labeling, access boundaries, configuration standards, and evidence-producing infrastructure practices.

Requirements

  • Proven experience operating production infrastructure as a cloud, infrastructure, platform, or site reliability engineer.
  • Terraform or OpenTofu at module-authoring depth - writing and versioning reusable modules, managing state across environments, and handling the lifecycle of real infrastructure over time.
  • Cloud foundation design experience - multi-account, multi-project, landing-zone, guardrail, IAM, or network architecture beyond a handful of isolated workloads.
  • Experience with infrastructure composition or orchestration - Atmos, Terragrunt, Terraspace, or an equivalent approach to managing reusable infrastructure across environments.
  • Strong networking depth - VPCs, routing, peering, DNS, TLS, load balancing, firewall policy, IP addressing, and connectivity troubleshooting.
  • Kubernetes working knowledge - cluster operations, RBAC, networking, shared services, and troubleshooting workloads that will not start or communicate correctly.
  • Helm chart authoring, not just chart installation.
  • Production cloud experience across at least two major providers, or deep experience with one plus substantial hands- on involvement extending an organization into another.
  • Python and Bash for automation, CLIs, diagnostics, and glue.
  • Linux fluency.
  • Experience with controlled infrastructure environments - SOC 2, HIPAA, HITRUST, FedRAMP, PCI, or similar security/compliance expectations.
  • Fluency in English and strong written communication. Architectural decisions here are expected to be documented and reviewed in writing

Experience That Will Help You Succeed

  • GitOps workflows across multiple environments and comfort operating infrastructure through reviewed, auditable changes.
  • Externalized secrets management and Kubernetes-native secret delivery patterns.
  • Operating Prometheus/Grafana/Loki or comparable observability tooling as an infrastructure consumer and operator.
  • Experience introducing or formalizing a second cloud, including a clear account of what you would repeat and what you would change.
  • Cloud identity federation and workload identity across providers.
  • Disaster-recovery, regional resilience, or high-availability infrastructure work.
  • Private cloud networking or overlay connectivity.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

5:28 min

Bitcoin scaling and the transition to peer-to-peer transactions

Jad Wahab · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:44 min

Career transition into cloud native and data management

Michael Cade · LIVE

Videos

See all

Related articles

See all