> Markdown version of [/jobs/ext/2725493-distributed-systems-engineers](https://www.wearedevelopers.com/jobs/ext/2725493-distributed-systems-engineers). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Distributed Systems Engineers - **Company:** Langdock - **Location:** Berlin, Germany - **Salary:** €90,000.0 - €140,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Systems Engineering, Automation of Tests, Microsoft Azure, Databases, Data Loss, Data Systems, Software Design Documents, Linux, File Systems, Distributed Systems, Protocol Buffers, PostgreSQL, Open Source Technology, Redis, TypeScript, Google Cloud, Kubernetes, Storage Technologies, Terraform - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/senior-distributed-systems-engineer-m-f-d-langdock-8966401 ## About the Role This is an experienced individual-contributor role. You should be able to independently investigate an unfamiliar problem, design a solution, ship it, and operate it in production. We generally expect this judgment to come from at least three to six six of full-time experience working on production systems., We will figure out the right level together based on your experience and scope. Levels are about the work you own, not your title or years of experience. ## Description Distributed Systems Engineers solve our most complex technical problems inside and across Langdock's shared services when existing building blocks or standard architectures are insufficient. These are problems where correctness under failure, performance, scale, security, and cost interact. This is a project-oriented role rather than ownership of one permanent layer. You might spend a period on distributed execution, then move into storage architecture, workload isolation, or model serving as company priorities change. The constant is sustained investigation of unfamiliar systems and responsibility for turning that understanding into a production system. Systems Engineers own the difficult mechanism and the measurable improvement it creates, whether in correctness, capability, performance, reliability, or cost. They operate what they build, while the Platform owner remains accountable for the service contract and long-term lifecycle around it. What you might work on The systems agenda includes: * Distributed state and execution. Design systems that remain correct when work is concurrent, long-running, retried, resumed, or moved between processes. Define explicit invariants for ordering, idempotency, recovery, and tenant isolation. * Storage and data systems. Improve how large volumes of customer and agent-generated data are stored, versioned, moved, and recovered while preserving consistency, data residency, and predictable performance. * Secure compute for agents. Develop the isolation, runtime, and scheduling mechanisms behind a shared sandbox service for untrusted code. The system needs persistent filesystems, deny-by-default networking, scoped mounts, fast startup, suspension, resumption, and predictable scheduling without weakening the security boundary. * Inference systems. Improve the serving systems beneath the Model Gateway, optimizing throughput, latency, accelerator utilization, reliability, and cost. The work is measurement driven: understand the workload, identify the actual bottleneck, and decide where changes to serving, scheduling, caching, model format, or hardware create structural advantage. * Resource scheduling and isolation. Make CPU, memory, storage, network, and accelerator capacity explicit so one workload cannot degrade another. Improve placement, admission control, backpressure, autoscaling, and recovery across deployment environments. * Billing and usage metering. Build the concurrent metering mechanism behind the billing capability, accounting for heterogeneous units such as model tokens and sandbox compute time. It must stay correct when many workloads report concurrently, events arrive late or more than once, and long-running jobs reserve capacity before their final usage is known. This involves idempotent ingestion, atomic reservations, reconciliation, and auditable records. You will start with one focused problem based on your experience and the team's priorities. The expectation is not broad activity across every domain; it is a substantial improvement in the capability, performance, reliability, or economics of the system you take on. Tech stack * Go and TypeScript in one Bazel monorepo, with implementation choices driven by the system * Linux, containers, microVMs, filesystems, and networking * Protobuf and gRPC for service contracts * Kubernetes across GCP, AWS, Azure, and on-premises deployments * Terraform and Terragrunt for infrastructure orchestration * PostgreSQL and Redis where durable metadata or coordination requires them * Open-source model serving and accelerator infrastructure You do not need prior experience with every item. You do need enough systems depth to enter an unfamiliar part of the stack, understand its behavior, and make consequential changes safely., * We operate with high trust and autonomy in squads of 3 to 4 engineers. A squad owns its roadmap, prioritization, technical decisions, and operation in production. Engineers are expected to find the context they need, ask for input when it improves the outcome, and move work forward without waiting for every next step to be assigned. * We align asynchronously before scheduling a meeting. Product requirement documents (PRDs) define the user problem, intended outcome, and constraints. Design documents make architectural boundaries, tradeoffs, failure modes, migrations, and rollouts explicit. People read and challenge the thinking asynchronously; once the context is shared, a short in-office discussion or whiteboard session usually resolves the remaining questions quickly. * We optimize for leverage. Engineers choose the AI tools that work for them, supported by clear ticket context, focused branches, automated tests, and AI review before human review. We also invest in observability, migration tooling, automated recovery, and runbooks so recurring product maintenance does not depend on someone remembering a manual step. * The engineer who ships a change owns it in production. If something breaks, you lead the fix. What we are looking for * You have owned technically difficult production systems in areas such as distributed systems, databases, runtimes, inference, networking, or storage. You operate what you build and remain responsible for it after deployment. * You reason from first principles, form testable hypotheses, and use measurement to understand unfamiliar behavior. You investigate beyond the first working solution until you understand the underlying mechanism. * Given an underspecified problem, you identify the most important invariant, decide where depth will change the outcome, and defend what you deliberately leave out. * You reason precisely about concurrency, isolation, ordering, data loss, and failure, while finding pragmatic ways to improve performance, reliability, or cost without unnecessary complexity. * You use AI tools as leverage while verifying their output and retaining ownership of the result. You communicate complex systems clearly, expose uncertainty, and work constructively with others. ## Related Videos - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) ## Related Articles - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)