Senior Software Engineer - Messaging & Async Platform Infrastructure

DeepL
Municipality of Valencia, Spain
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English

Job location

Municipality of Valencia, Spain

Tech stack

Amazon Web Services (AWS)
Azure
Cloud Computing
Distributed Systems
Fault Tolerance
Protocol Buffers
JSON
Python
RabbitMQ
Prometheus
Data Streaming
TypeScript
Datadog
Pulumi
Google Cloud Platform
Istio
Grafana
Backend
Event Driven Architecture
Containerization
Kubernetes
Infrastructure Automation Frameworks
Low Latency
Avro
Kafka
Asynchronous Programming
Api Gateway
Terraform
Go
Programming Languages

Job description

About the RoleWe are looking for a Senior Software Engineer to join our Infrastructure - Messaging & Async Platform team, focused on building and scaling our messaging and asynchronous communication platform. This role is ideal for someone who has strong experience designing, operating, and improving distributed systems that power event-driven architectures. You will work on the core infrastructure that enables reliable communication between services, supports high-throughput workloads, and helps engineering teams build resilient asynchronous workflows. You should have hands-on experience with technologies such as NATS, Kafka, RabbitMQ, Pulsar, or similar messaging and streaming platforms, as well as a strong understanding of distributed systems, observability, reliability, and performance at scale.ResponsibilitiesOperate and improve our async platform based on NATS; design, build, and maintain messaging and event-driven infrastructure used across multiple engineering teams. Develop tooling, libraries, and platform capabilities that make asynchronous communication easier, safer, and more reliable for engineering teams.Improve reliability, scalability, latency, throughput, and operational visibility of messaging systems.Define best practices for event publishing, consuming, schema evolution, retries, dead-letter queues, backpressure, and message durability.Partner with application teams to migrate, adopt, or optimize their use of messaging and async workflows.Troubleshoot complex production issues in distributed systems and drive long-term improvements.Contribute to architecture decisions around event-driven systems, service communication, and platform reliability.Participate in incident response, postmortems, and continuous improvement of operational processes.RequirementsStrong professional experience as a software engineer, ideally in infrastructure, platform engineering, backend systems, or distributed systems.Hands-on experience with messaging, streaming, or pub/sub technologies such as NATS, Kafka, RabbitMQ, Pulsar, or similar.Solid understanding of distributed systems concepts, including consistency, ordering, delivery guarantees, fault tolerance, retries, idempotency, and backpressure.Experience operating production systems at scale, including monitoring, alerting, and incident response.Proficiency in one or more programming languages such as Typescript, Go, Python, or similar.Experience with cloud infrastructure and containerized environments, such as Kubernetes, AWS, GCP, or Azure.Ability to design pragmatic, reliable systems and communicate technical trade-offs clearly.Strong ownership mindset and comfort working across teams on shared infrastructure.Nice to HaveDeep experience operating NATS, KAFKA, or another messaging platform in high-throughput production environments.Experience building internal developer platforms, SDKs, or shared infrastructure services.Knowledge of schema registries, event contracts, protobuf, Avro, or JSON Schema.Experience with service mesh, API gateways, or service-to-service communication patterns.Familiarity with infrastructure as code tools such as Terraform, Pulumi, or similar.Experience with observability platforms such as Prometheus, Grafana, OpenTelemetry, Datadog, or similar.Prior experience defining reliability standards, SLOs, and operational playbooks.#J-*****-Ljbffr

Requirements

Strong professional experience as a software engineer, ideally in infrastructure, platform engineering, backend systems, or distributed systems. Hands-on experience with messaging, streaming, or pub/sub technologies such as NATS, Kafka, RabbitMQ, Pulsar, or similar. Solid understanding of distributed systems concepts, including consistency, ordering, delivery guarantees, fault tolerance, retries, idempotency, and backpressure. Experience operating production systems at scale, including monitoring, alerting, and incident response. Proficiency in one or more programming languages such as Typescript, Go, Python, or similar. Experience with cloud infrastructure and containerized environments, such as Kubernetes, AWS, GCP, or Azure. Ability to design pragmatic, reliable systems and communicate technical trade-offs clearly. Strong ownership mindset and comfort working across teams on shared infrastructure. Nice to Have Deep experience operating NATS, KAFKA, or another messaging platform in high-throughput production environments. Experience building internal developer platforms, SDKs, or shared infrastructure services. Knowledge of schema registries, event contracts, protobuf, Avro, or JSON Schema. Experience with service mesh, API gateways, or service-to-service communication patterns. Familiarity with infrastructure as code tools such as Terraform, Pulumi, or similar. Experience with observability platforms such as Prometheus, Grafana, OpenTelemetry, Datadog, or similar. Prior experience defining reliability standards, SLOs, and operational playbooks. #J-*****-Ljbffr

About the company

Helping people overcome communication barriers is the heart of what we do. Founded in Germany in 2017 by a team of engineers and researchers, DeepL has developed the world’s most accurate AI translation technology—enabling real-time, human-sounding translation.

Accessible via a web translator, browser extensions, desktop and mobile apps, and an API, DeepL supports a best-in-class translation experience in 34 languages and counting. Our 550-person team operates across four European hubs in Germany, the Netherlands, the UK, and Poland.

Apply for this position