Member of Technical Staff - Research Infrastructure Engineer

Black Forest Labs
Freiburg im Breisgau, Germany
1 day ago
Apply on www.adzuna.de
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
€100,000.0 - €230,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Bash Shell Cloud Engineering Software Debugging Distributed Systems Monitoring of Systems Python (Programming Language) Prometheus Graphics Processing Unit (GPU) Google Cloud Reliability of Systems Kubernetes
+3 more
Infrastructure Automation Frameworks Slurm Golang

Job description

  • Maintain research infrastructure, ensuring health, and optimizing components to extract peak performance from the system (both on application, and infrastructure side)
  • Scale infrastructure to meet growing research demands while maintaining reliability and performance
  • Collaborate with research teams to deeply understand their infrastructure needs, and design solutions that balance performance with cost efficiency.
  • Identify and resolve performance bottlenecks and capacity hotspots through deep analysis of distributed systems at scale.
  • Build and evolve telemetry and monitoring systems to provide deep visibility into infrastructure performance, utilization, and costs across our cloud and datacenter fleets.
  • Participate in on-call rotations and incident response to maintain system reliability

Technical Focus

  • Python, Bash, Go
  • Kubernetes
  • Nvidia GPU drivers, and operators
  • OTel, Prometheus, We’re a distributed team with real offices that people actually use. Depending on your role, you’ll either join us in Freiburg or SF at least 2 days a week (or one full week every other week), or work remotely with a monthly in-person week to stay connected. We’ll cover reasonable travel costs to make this possible. We think in-person time matters, and we’ve structured things to make it accessible to all. We’ll discuss what this will look like for the role during our interview process.

Requirements

  • Experience building or operating large-scale training platforms
  • Worked with large scale compute clusters (GPUs)
  • Proven ability to debug performance and reliability issues across large distributed fleets
  • Strong problem-solving skills and ability to work independently
  • Strong communication skills and the ability to work effectively with both internal and external partners
  • Deep knowledge of modern cloud infrastructure including Kubernetes, Infrastructure as Code, AWS, and GCP
  • Experience with SLURM
  • Experience building or operating large-scale training platforms

Benefits & conditions

  • Obsessed. We are a frontier research lab. The science has to be right, the understanding deep, the product beautiful.
  • Low Ego. The work speaks. The best idea wins, no matter who said it. Credit is shared. Nobody is above any task.
  • Bold. We take the ambitious bet. We ship, we do not wait for conditions to be perfect.
  • Kind. People over politics. We treat each other with genuine warmth. Agency without empathy creates chaos.

If this sounds like work you’d enjoy, we’d love to hear from you.

Base Annual Salary:

EU €100,000 - €230,000 + Equity

US $150,000 - $300,000 + Equity

About the company

About Black Forest Labs

We’re the team behind Latent Diffusion, Stable Diffusion, and FLUX-foundational technologies that changed how the world creates images and video. We’re creating the generative models that power how people make images and video-tools used by millions of creators, developers, and businesses worldwide. Our FLUX models are among the most advanced in the world, and we’re just getting started.

Headquartered in Freiburg, Germany with a growing presence in San Francisco, we’re scaling fast while staying true to what makes us different: research excellence, open science, and building technology that expands human creativity.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.de
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

3:52 min

Avoiding remote code execution from unsanitized inputs

Alexander Pirker · World Congress 2022

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:32 min

Career evolution from content management to cloud infrastructure

Maurice Brinkmann · World Congress 2023

Videos

See all

Related articles

See all