Software Engineer - Infra Agent Systems

Together Ai
Amsterdam, Netherlands
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Automated Storage and Retrieval Systems Intelligent Platform Management Interface Command-Line Interface Programming Tools Distributed Systems Graph Database Python (Programming Language) Enterprise Messaging Systems Cloud Services Prometheus
+16 more
Software Systems TypeScript AI Infrastructure Large Language Models Grafana Event Driven Architecture Build Management Kubernetes Slack Bare Metal Apache Kafka Hardware Infrastructure Virtual Agents Api Design Software Version Control Golang

Job description

Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure.

We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling.

You’ll work across two areas:

Infrastructure Agent Systems - Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack.

Core Agent Platform - Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve.

We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure.

This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation, solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure.

responsible for delivering the software but also for operating and supporting it in production.

Why this Role

You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective.

You’ll have the opportunity to build foundational systems from the ground up, work on infrastructure at massive scale, and help define how self-improving AI agents operate real-world AI infrastructure., * Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.

  • Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
  • Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions.
  • Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs.
  • Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.
  • Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops.
  • Turn what agents learn in production into reliable, reviewed software and automation.

Requirements

  • 5+ years of experience building production backend systems, distributed systems, or infrastructure platforms.
  • Strong systems design skills and experience owning significant systems from design through production.
  • Depth in at least one of the following:

  • AI agent systems, orchestration, tool use, evaluation, or grounding
  • Knowledge graphs or graph data modeling
  • Search, retrieval, ranking, RAG, or semantic search systems

Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems.

Experience with Kubernetes, GitOps such as ArgoCD, infrastructure-as-code, and cloud platforms.

Comfortable working across languages such as Go, TypeScript, Python, or Rust.

Experience in the following is a plus:

  • GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers
  • Graph databases
  • Event-driven systems and messaging platforms such as NATS or Kafka
  • Observability platforms such as Prometheus and Grafana
  • Building evaluation frameworks or improving the quality and reliability of LLM-powered systems

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:45 min

Fusing developer experience and platform engineering for agentic SDLC

Julia Kordick Julia Kordick · World Congress 2026 Europe

1:14 min

Automating user bug triage and resolutions using Slack agents

Brian Lovin Brian Lovin · World Congress 2026 Europe

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:52 min

Orchestrating autonomous specialist agents for complex system delivery

Ayotunde Obasa Ayotunde Obasa · Europe 2026 Virtual

5:36 min

Integrating agentic coding models into Slack

Chris Heilmann +2 · LIVE

Videos

See all

Related articles

See all