Senior Site Reliability Engineer

Tamarind Intelligence
Madrid, Spain
9 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English, Spanish

Tech stack

Java (Programming Language) Backup Devices BigQuery Cloud Computing Cloud Computing Security Continuous Integration Disaster Recovery Distributed Systems Failover Github Identity and Access Management Java Virtual Machine (JVM)
+23 more
Key Management PostgreSQL Network Segmentation Role-Based Access Control Reliability Engineering Prometheus Software Engineering Software Vulnerability Management YAML Data Logging Cloud Platform System Cloud Monitoring Autoscaling Grafana Backend Rate Limiting Kotlin Kubernetes Production Code Apache Kafka Terraform Dynatrace Docker

Job description

Who we areTRLLNis building the future of asset intelligence.OurTracking-as-a-Service (TaaS)platform combines a proprietaryIoT mesh network, telemetry, and advanced analyticsto deliver real-time visibility over physical assets — track any asset, anywhere — with no gates, readers, or fixed infrastructure.Born insideIFCO, the global market leader in reusable packaging containers for fresh food, TRLLN is now astandalone subsidiary of IFCO: engineered for the world’s largest reusable-asset network —400 million assets across 50+ countries— and combining that proven, industrial scale with the agility and focus of a tech company, from ourBarcelona HQ .What environment will you be joining?As aSenior Site Reliability EngineeratTRLLN, you’ll join the product engineering team that builds and runs thecore of our platform: the backend and infrastructure that ingest, process, and servehigh-volume IoT telemetryin near real time, on amulti-tenant, cloud-native architectureonGoogle Cloud— GKE, Cloud Run, Managed Kafka, Pub/Sub, AlloyDB and BigQuery, all managed with Terraform.You will be ourfirst dedicated reliability role.Today reliability is everyone’s job and nobody’s specialty: the team has strong software engineering foundations, an observability stack we are actively maturing, and availability commitments to customers that we want to back with realSLOs, alerting and runbooks.Your job is to turn that into areliability practice— and to make the engineers around you better at operating what they build.Security is a big part of this role.The platform goes through external security assessments and is preparing for theEU Cyber Resilience Act ; the operational side of that — cloud security posture, vulnerability and supply-chain management, incident readiness, business continuity — will sit with you.TRLLN is astartup-like environment backed by a global enterprise: we move fast, iterate with purpose, and care deeply about doing things right, always balancingpragmatism, speed and quality .And we face real scale challenges: seasonal traffic peaks from massive IoT device rollouts, a growing hardware fleet in the field, demanding reliability targets, and a platform engineered to keep growing.What we’re looking forStronganalytical and problem-solvingskills, with a focus on delivery and quality.Demonstrated leadership intechnical decision-making,mentoring, andteam collaboration .Excellentcommunication skills.Able to explain complex systems and trade-offs clearly, both verbally and in writing.Anenabler mindset .You make the team better.What you bringProven experienceoperating high-volume / high-traffic distributed systems in production— event streaming (Kafka, Pub/Sub or similar), telemetry or time-series data, performance under real load, and capacity planning for traffic peaks.Strongreliability engineeringbackground: experience working with SLIs/SLOs and error budgets, alerting design, incident management, blameless postmortems.Hands-on with theobservability stack : Prometheus / OpenTelemetry / Grafana or a cloud-native equivalent (we use Cloud Monitoring + Managed Prometheus) — metrics, logging and distributed tracing, including their cost.SolidKubernetes,DockerandTerraform(or similar IaC) experience — networking, autoscaling, workload identity, private clusters.We run onGCP , but experience on any major cloud transfers.Production databases : PostgreSQL (AlloyDB / Cloud SQL or equivalent) — performance, backups, restore and failover.Platform security, hands-on : cloud security posture (IAM and least privilege, network segmentation, secrets management, encryption / KMS), WAF and rate limiting (Cloud Armor or similar), and remediating findings from external security assessments and pentests.Vulnerability management and supply chain : image and dependency scanning, SBOMs, signed and attested builds, patch cadence — plus the operational side of security incident response and business continuity / disaster recovery. Compliance as an engineer : comfortable turning regulatory and audit requirements (EU Cyber Resilience Act, ISO *** / SOC 2) into controls, evidence and runbooks — you’ll help us get audit-ready.CI/CD and GitOps : pipelines as code, progressive delivery, safe rollbacks (we use GitHub Actions).Youwrite production code : comfortable reading and contributing to Go and/or JVM (Kotlin / Java) services, not just YAML.Automation and instrumentation are part of the job.Comfortable working in both English and Spanish.Nice to haveCloudFinOps/ cost optimisation experience.Load and chaos testing.Exposure toIoT / device fleetsor telemetry-heavy products.Security certifications (e.g. CKS, GCP Professional Cloud Security Engineer) or prior experience on an ISO *** / SOC 2 certification effort.What we offerHybrid work modelwith aflexible schedule Based in ourBarcelona HQ Comprehensive health insurancefor your family unit Salary range:**€ – **€ Annual training budgetfor continuous learning and growth#J-*****-Ljbffr

Requirements

Strong analytical and problem-solving skills, with a focus on delivery and quality. Demonstrated leadership in technical decision-making , mentoring , and team collaboration . Excellent communication skills. Able to explain complex systems and trade-offs clearly, both verbally and in writing. An enabler mindset . You make the team better. What you bring Proven experience operating high-volume / high-traffic distributed systems in production — event streaming (Kafka, Pub/Sub or similar), telemetry or time-series data, performance under real load, and capacity planning for traffic peaks. Strong reliability engineering background: experience working with SLIs/SLOs and error budgets, alerting design, incident management, blameless postmortems.

About the company

Who we are
TRLLN
is building the future of asset intelligence.
Our
Tracking-as-a-Service (TaaS)
platform combines a proprietary
IoT mesh network, telemetry, and advanced analytics
to deliver real-time visibility over physical assets — track any asset, anywhere — with no gates, readers, or fixed infrastructure.
Born inside
IFCO
, the global market leader in reusable packaging containers for fresh food, TRLLN is now a
standalone subsidiary of IFCO
engineered for the world’s largest reusable-asset network — 400 million assets across 50+ countries — and combining that proven, industrial scale with the agility and focus of a tech company, from our Barcelona HQ . What environment will you be joining? As a Senior Site Reliability Engineer at TRLLN , you’ll join the product engineering team that builds and runs the core of our platform
the backend and infrastructure that ingest, process, and serve high-volume IoT telemetry in near real time, on a multi-tenant, cloud-native architecture on Google Cloud — GKE, Cloud Run, Managed Kafka, Pub/Sub, AlloyDB and BigQuery, all managed with Terraform. You will be our first dedicated reliability role .

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Loading talks and stories from around this role…