Software Engineer III – Platform Software

ClearCompany
United States
3 days ago
Apply on alleninstitute.hrmdirect.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$146,600.0 - $183,250.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Amazon S3 Automation of Tests Command-Line Interface Code Review Continuous Integration Data Security Cloud Services Prometheus Webui Software Engineering
+5 more
Software Application Programming Multi-Cloud Kubernetes Information Technology Machine Learning Operations

Job description

  • Own the control plane surface scientists and admins touch daily (API, web UI, and CLI)
  • Own Kubernetes-native platform software - operators, CRDs, and controllers for tenancy, storage, and job lifecycle across large workloads and interactive compute
  • Telemetry as a product: provide data for understanding capacity, utilization and improvement opportunities
  • Ship with production discipline - CI/CD, automated testing, code review, and quality standards - and hold that bar for work delivered by partner and vendor teams
  • Own security and compliance of what you build
  • Work across the AWS and GCP services the platform composes (eg EKS, GKE, FSx, S3, and GCS)
  • Implement experiences for the whole Institute for cutting-edge AI methods - learn emerging practices and build or scale toolkits for agentic science, data access, and AI-augmented research
  • Collaborate with Allen Institute and vendor teams spanning scientists, engineers, and service

Note: Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions. This description reflects management’s assignment of essential functions; it does not proscribe or restrict the tasks that may be assigned.

Requirements

  • Bachelor’s Degree in Computer Science, or related technical field; or equivalent combination of degree and experience
  • Minimum of 3 years of experience building ML infrastructure (compute and storage)
  • Strong software engineering in Go, with Kubernetes operators or controllers shipped to production

Preferred Education and Experience

  • Proficiency with building applications on Kubernetes
  • Telemetry fluency with Prometheus, and log pipelines
  • Cloud expertise, multi-cloud preferred (AWS and GCP)
  • Experience with testing, CI/CD, and quality standards as habits
  • Comfort running security reviews of your own systems, including compliance-scoped workloads
  • Empathy for researchers as users, and for their workflows
  • GPU scheduler internals, NVIDIA stack, JupyterHub internals, HPC-to-Kubernetes migrations, FinOps / billing system exposure

Physical Demands

  • Fine motor movements in fingers/hands to operate computers and other office equipment

Position Type/Expected Hours of Work

  • This role is currently able to work both remotely and onsite in a hybrid work environment. We are a Washington State employer, and the primary work location for all Allen Institute employees is 615 Westlake Ave N.; any remote work must be performed in Washington State.

Benefits & conditions

  • Employees (and their families) are eligible to enroll in benefits per eligibility rules outlined in the Allen Institute’s Benefits Guide. These benefits include medical, dental, vision, and basic life insurance. Employees are also eligible to enroll in the Allen Institute’s 401k plan. Paid time off is also available as outlined in the Allen Institutes Benefits Guide. Details on the Allen Institute’s benefits offering are located at the following link to the Benefits Guide: https://alleninstitute.org/careers/benefits.

About the company

The Allen Institute accelerates science for a healthier world through large-scale research designed to answer some of the most complex questions in biology. Our multi-disciplinary teams generate foundational knowledge, tools, and data to understand how our brain, cells, and immune system work. We share our work openly so others can build on it, move faster, and ask bigger questions. We drive discovery forward and create new possibilities for improving human health.

The Office of the Chief Technology Officer (OCTO) drives computational and technological innovation across the Allen Institute by advancing engineering excellence, developing scalable technology solutions, and enabling scientific discovery. Through strategic partnerships and technical leadership, OCTO accelerates research, enhances operational efficiency, and fosters collaboration in support of the Institute’s mission.

We are seeking an independent and motivated Platform Software Engineer in the Office of the CTO to drive the Institute’s AI compute platform. Reporting to the Director of AI Infrastructure, you will own the platform as a product: the control plane, operators, and user experience over multi-cloud Kubernetes GPU infrastructure serving scientific AI workloads. You will set and hold the production bar — for your own work and for what partner teams deliver — as the platform’s durable owner. Beyond the platform, you will catalyze cutting-edge AI methods across the Institute — exchanging practices and creating toolkits for agentic science and AI-augmented research.

At the Allen Institute, we believe that science is for everyone – and should be open to everyone. We are dedicated to combating biases and reducing barriers to STEM careers more broadly.

We also believe that science is better when it includes different perspectives and voices. We strive to make the Allen Institute a place where everyone feels like they belong and are empowered to do their best work in a supportive environment.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on alleninstitute.hrmdirect.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

1:05 min

Measuring system availability utilizing Prometheus and straightforward PromQL

Alexander Schwartz Alexander Schwartz · World Congress 2025

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

3:44 min

Automating storage savings with S3 intelligent tiering

Sébastien Stormacq · World Congress 2021

13:07 min

Configuring application observability with Micrometer and Prometheus

Aleksandr Kalikov · LIVE

Videos

See all

Related articles

See all