Platform Engineer

CLERA, LLC
San Francisco, CA, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$150,000.0 - $250,000.0
Working hours
Regular working hours
Job source

Tech stack

Clean Code Principles Training Data Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Application Release Automation Cloud Computing Databases Continuous Integration Identity and Access Management
+12 more
Key Management Azure Machine Learning Service-Oriented Architecture Software Engineering Reinforcement Learning Autoscaling Delivery Pipeline Caching Backend Kubernetes Terraform Docker

Job description

We’re a fast-moving AI/ML platform startup building infrastructure for reinforcement learning environments, post-training data pipelines, and large-scale agent evaluation. Our engineering team of ~15 includes exceptional technical talent - competition medalists, serial AI startup founders, and published researchers., * Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.

  • Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
  • Design and improve backend and platform systems for scale - including capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.
  • Define and improve dashboards, alerts, logs, traces, SLOs, runbooks, and on-call workflows so failures are detected, debugged, and resolved quickly.
  • Build reliable CI/CD pipelines, release automation, environment management, and deployment workflows that improve developer productivity and reduce production risk.
  • Write clean, maintainable code to automate systems, improve backend services, and create internal tooling.

Requirements

  • 2-4 years of experience owning production cloud infrastructure for a high-availability, user-facing platform, with responsibility for uptime, performance, deployment safety, and cost.
  • Deep hands-on experience with AWS and containerized systems; strong familiarity with Terraform, Kubernetes/EKS, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
  • Proven track record building or operating CI/CD, release automation, observability, alerting, and incident response systems.
  • Strong backend engineering judgment - ability to reason about service architecture, APIs, databases, async systems, queues, scaling limits, and production failure modes.
  • Ability to write clean, maintainable code and apply software engineering judgment across infrastructure, backend systems, and developer workflows.
  • High ownership mindset; comfortable being accountable for production systems end-to-end.

Nice to Have

  • Experience designing systems for bursty workloads, long-running jobs, sandboxed execution, distributed workers, or high-concurrency services.
  • Background operating infrastructure for data-heavy, ML/AI, workflow, marketplace, developer-tools, or enterprise platforms.
  • Demonstrated focus on reducing cloud spend through better architecture, autoscaling, workload placement, caching, cleanup systems, or observability.

Benefits & conditions

  • Salary: $150,000 - $250,000 USD annually (full-time, US-based).
  • Equity participation in an early-stage, well-funded AI startup.
  • Work alongside a world-class technical team on infrastructure that operates at real scale.
  • High degree of autonomy and direct impact on product and platform direction.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all