Senior Software Engineer (Serverless)

Jobgether
Málaga, Spain
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Málaga, Spain

Tech stack

Clean Code Principles
API
Artificial Intelligence
Cloud Computing
Code Review
Computer Programming
Software Design Documents
Distributed Systems
Open Source Technology
Performance Tuning
Software Engineering
AI Infrastructure
Graphics Processing Unit (GPU)
System Availability
Large Language Models
Deep Learning
AWS Lambda
AI Platforms
Kubernetes
Optimization Algorithms
Cloudflare
TensorRT
Serverless Computing
Go

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer (Serverless) based in Spain. This role offers the opportunity to build the next generation of AI cloud infrastructure powering advanced machine learning workloads worldwide. You will work on a high-impact serverless platform designed to help developers deploy and scale AI applications without managing complex infrastructure. As a senior engineer, you will take ownership of critical distributed systems challenges, from GPU scheduling and runtime performance to customer-facing APIs and platform reliability. Working in a highly technical environment, you will influence architecture decisions, mentor engineers, and help define engineering standards. This position is ideal for an experienced software engineer passionate about large-scale systems, cloud technologies, and solving complex infrastructure challenges at the forefront of AI innovation.AccountabilitiesDesign, develop, and maintain core components of a GPU-native serverless AI platform, including control planes, schedulers, runtimes, autoscaling systems, and customer-facing APIs.Solve complex engineering challenges related to cold-start optimization, GPU scheduling, multi-tenant isolation, fair resource allocation, request routing, and platform scalability.Own technical architecture decisions for key platform areas by creating design documents, evaluating solutions, and aligning engineering teams around effective approaches.Establish and maintain high engineering standards through code reviews, design reviews, technical guidance, and active collaboration with team members.Operate services with an SRE mindset by defining reliability objectives, improving observability, supporting incident response, and driving continuous platform improvements.Collaborate directly with customers on architecture discussions, performance optimization, and complex production issues requiring deep technical expertise.Partner with product, infrastructure, and go-to-market teams to translate customer needs into scalable technical roadmaps.Contribute to improving platform performance, reliability, and developer experience through innovative engineering solutions.Support knowledge sharing and technical mentorship to help raise the overall engineering capability of the team.Requirements7+ years of professional software engineering experience with a proven track record of building and operating large-scale distributed systems.Strong programming experience with Golang, or the ability and willingness to quickly become proficient in the language.Deep experience with Kubernetes and container orchestration systems, including real-world operation of production environments.Strong understanding of distributed systems concepts such as consistency, availability trade-offs, queueing, backpressure, retries, idempotency, and multi-tenancy.Experience designing and operating high-throughput, low-latency services with a strong focus on performance optimization and reliability.Proven ability to take ownership of complex technical challenges, lead design discussions, unblock teams, and deliver impactful solutions.Strong software engineering fundamentals with the ability to write reliable, maintainable code and investigate complex technical problems.Collaborative mindset with excellent communication skills and the ability to work effectively across engineering and business teams.Experience with serverless platforms or function-as-a-service technologies such as Knative, AWS Lambda, GCP Cloud Run, Cloudflare Workers, or similar solutions is a strong advantage.Familiarity with GPU scheduling technologies, including Kubernetes device plugins, MIG, MPS, time-slicing, or NVIDIA GPU Operator, is highly desirable.Experience with ML inference technologies such as vLLM, TensorRT-LLM, Triton Inference Server, SGLang, or similar platforms is a plus.Knowledge of runtime optimization techniques including cold-start reduction, image streaming, checkpoint/restore, or sandboxing technologies is beneficial.Experience developing Kubernetes operators using Go and frameworks such as controller-runtime or kubebuilder is advantageous.Contributions to open-source projects related to serverless infrastructure, scheduling, or AI inference are a plus.BenefitsCompetitive compensation package.Flexible working environment with hybrid opportunities.High level of ownership and autonomy over technical decisions.Career development opportunities and continuous learning support.Opportunity to work on impactful AI infrastructure projects shaping the future of machine learning.Collaborative and innovative culture with highly skilled international teams.Exposure to cutting-edge technologies across cloud infrastructure, distributed systems, GPUs, and AI workloads.Opportunity to contribute to ambitious projects in a fast-moving, high-growth environment.#J-*****-Ljbffr

Requirements

7+ years of professional software engineering experience with a proven track record of building and operating large-scale distributed systems. Strong programming experience with Golang, or the ability and willingness to quickly become proficient in the language. Deep experience with Kubernetes and container orchestration systems, including real-world operation of production environments. Strong understanding of distributed systems concepts such as consistency, availability trade-offs, queueing, backpressure, retries, idempotency, and multi-tenancy. Experience designing and operating high-throughput, low-latency services with a strong focus on performance optimization and reliability. Proven ability to take ownership of complex technical challenges, lead design discussions, unblock teams, and deliver impactful solutions. Strong software engineering fundamentals with the ability to write reliable, maintainable code and investigate complex technical problems. Collaborative mindset with excellent communication skills and the ability to work effectively across engineering and business teams. Experience with serverless platforms or function-as-a-service technologies such as Knative, AWS Lambda, GCP Cloud Run, Cloudflare Workers, or similar solutions is a strong advantage. Familiarity with GPU scheduling technologies, including Kubernetes device plugins, MIG, MPS, time-slicing, or NVIDIA GPU Operator, is highly desirable. Experience with ML inference technologies such as vLLM, TensorRT-LLM, Triton Inference Server, SGLang, or similar platforms is a plus. Knowledge of runtime optimization techniques including cold-start reduction, image streaming, checkpoint/restore, or sandboxing technologies is beneficial. Experience developing Kubernetes operators using Go and frameworks such as controller-runtime or kubebuilder is advantageous. Contributions to open-source projects related to serverless infrastructure, scheduling, or AI inference are a plus.

Benefits & conditions

Competitive compensation package. Flexible working environment with hybrid opportunities. High level of ownership and autonomy over technical decisions. Career development opportunities and continuous learning support. Opportunity to work on impactful AI infrastructure projects shaping the future of machine learning. Collaborative and innovative culture with highly skilled international teams. Exposure to cutting-edge technologies across cloud infrastructure, distributed systems, GPUs, and AI workloads. Opportunity to contribute to ambitious projects in a fast-moving, high-growth environment. #J-*****-Ljbffr

Apply for this position