Infrastructure Software Engineer

Normal Computing
New York, NY, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$185,000.0 - $285,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Software as a Service Continuous Integration Software Debugging Distributed Systems Job Scheduling PostgreSQL Operational Databases Software Product Management Redis Reliability Engineering
+14 more
Software Engineering Autoscaling Large Language Models Concurrency State Machines Kubernetes Helm Charts Backend Containerization Kubernetes Machine Learning Operations Virtual Agents Terraform Automation Anywhere Docker

Job description

This is an application engineering role focused on infrastructure-shaped software: orchestration services, execution runtimes, internal APIs, persistence layers, observability, and developer experience. You’ll help define the runtime layer for a new class of AI products: systems where agents execute long-running work, coordinate across distributed environments, interact with code and tools, and need to be reliable enough for real customer workflows.

This role sits between product engineering, AI engineering, and platform engineering. You will not primarily be managing Terraform, Helm charts, CI/CD, or company-wide SaaS infrastructure. Instead, you’ll own the application-level infrastructure that powers long-running AI workflows: session lifecycle, sandboxed execution, workload orchestration, persistence, observability, reliability, and the internal interfaces other engineers build on.

The systems you build will be used directly by product, AI, research, and platform teams as new capabilities move from early ideas into production. Developer experience matters: APIs should be understandable, failure modes should be debuggable, and abstractions should make the right thing easy.

This is a highly cross-functional role for someone who enjoys ambiguity, cares about clean abstractions, and wants to help shape how a frontier AI company builds and operates production systems. Strong engineering judgment and ownership matter more than rigid specialization.

On any given day, you might design the runtime architecture for a new AI product capability, build the orchestration layer for long-running autonomous workflows, improve how workloads are scheduled and isolated across distributed environments, or create the systems abstractions that let engineers turn ambitious AI prototypes into reliable production products., * Build and maintain production software infrastructure for Normal’s AI products, especially orchestration, execution, and runtime systems.

  • Design internal backend services and APIs used by product engineers, AI engineers, execution services, and other internal systems.
  • Improve the operational maturity of rapidly evolving systems through better state management, failure handling, metrics, tracing, and debugging tools.
  • Work with Kubernetes-backed execution environments, including container lifecycle, scheduling behavior, autoscaling, resource isolation, and runtime reliability.
  • Build developer-facing tools and abstractions that make it easier for other engineers to use and extend the systems you own.
  • Turn promising prototypes into durable production systems by designing clear abstractions, hardening critical paths, and creating operational patterns that scale with the product.
  • Collaborate closely with product, AI, research, and platform engineers to define the right interfaces between product features, AI workloads, and production infrastructure.
  • Lead design discussions for core runtime and orchestration systems, including API boundaries, state management, execution models, and operational tradeoffs.

Requirements

Do you have experience in Design (software development lifecycle)?, * 4+ years of experience in infrastructure software, backend infrastructure, production infrastructure, platform engineering, distributed systems, or related areas.

  • Strong software engineering fundamentals, including backend programming, APIs, data modeling, concurrency, debugging, and testing.
  • Experience building or operating production services where reliability, observability, and maintainability matter.
  • Practical experience with Docker and Kubernetes, including debugging containerized workloads, deployments, networking, resource limits, and lifecycle issues.
  • Comfort working with persistence systems such as Postgres, Redis/Valkey, object storage, or similar production data stores.
  • Experience building orchestration systems, job schedulers, workflow engines, sandboxes, developer platforms, or distributed execution systems.
  • Experience designing internal APIs and developer-facing abstractions that other engineers can use confidently.
  • Strong systems thinking: you can reason about state machines, failure modes, retries, queues, leases, scheduling, and long-running workflows.
  • Pragmatism in fast-moving environments: you know when to improve an abstraction, when to delete one, and when to ship the simple version.
  • Ownership mindset: you care whether the systems you build work in production and are usable by other engineers.
  • Clear communication and good technical judgment across product, AI, and infrastructure boundaries.

Nice to Have

  • Deep Kubernetes experience, such as controllers/operators, networking, storage, scheduling, autoscaling, or resource isolation.
  • Experience with AI agent infrastructure, ML infrastructure, model orchestration, or LLM-based product systems.
  • Background in production infrastructure, reliability engineering, or infrastructure software at meaningful scale.
  • Experience in high-growth startups or engineering teams where ownership boundaries are still being defined.
  • Experience with Chips, EDA or Device Verification

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · WWC Europe 2026

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · WWC 2022

Videos

See all

Related articles

See all