Site Reliability Engineer
Role details
Job location
Tech stack
Job description
At Tinybird, we help developers and data teams unlock the power of real-time data. This enables them to build data pipelines and innovate data products quickly. With Tinybird, you can seamlessly ingest multiple data sources at scale, query them using SQL you already know, and publish results as low-latency, high-concurrency APIs for your applications. Developers can create fast APIs; time it used to take hours or days now takes only minutes. Tinybird is the essential tool data engineers and software developers have been waiting for, making it easier to drive innovation., We seek an experienced Site Reliability Engineer who enjoys keeping large-scale distributed systems reliable and adaptable as they grow. You should understand how to make hardware and software work well together and be eager to grasp both our product and the real challenges our customers and internal teams face., * Zookeeper: for coordinating ClickHouse replicas.
- ArgoCD: for GitOps-based continuous delivery.
- Grafana, Loki, Mimir, and OpenTelemetry (OTEL): for monitoring, alerting, and telemetry (we're increasingly standardizing on OTEL).
Hiring Process
- Initial contact meeting with the Hiring Manager to discuss the process.
- Live Technical Assessment (1 hour, live, screen sharing, with two team members) to demonstrate your skills.
- Team Alignment meeting (45 minutes with the other two team members) to discuss technology and teamwork.
- Final meeting with our CEO to review culture fit, long-term company vision, and any remaining questions.
Requirements
- You have strong experience in designing, building, and running distributed cloud architectures and large-scale web-based production systems.
- You have deep knowledge of Kubernetes, which is essential for this role. You should be comfortable designing and operating production-grade clusters, writing custom controllers or operators as needed, and tuning autoscaling mechanisms (KEDA, Karpenter, and similar) to respond to real-time workloads. You know how Kubernetes manages networking, storage, scheduling, and resources, and you can analyze performance and failure scenarios at scale.
- You are skilled in AWS and GCP.
- Coding skills are required. We are not looking for a software developer, but baseline. You should be able to explore our codebase, ClickHouse source code, or any other software we use to understand how things work. Our primary languages are Python and some C++.
- You are comfortable operating close to production: debugging incidents, understanding system behavior, improving observability, and enhancing service reliability.
- You think in systems and pay attention to edge cases, failure modes, and specific implementation details.
- You care about performance, reliability, cost efficiency, and operational simplicity. You prioritize action, iteration, and delivery. You know many decisions can be reversed quickly, and that speed is important in business and technology.
- You take ownership, follow through, and are willing to tackle issues that may be broken, because you can fix them if necessary.
- You enjoy data and SQL, and you are curious about how real-time analytical systems work. We use our own product, so you'll need some SQL experience to query our own data. Experience with ClickHouse and/or launching database systems at scale would be a big plus.
- Familiarity with Traefik, Varnish, Redis, Terraform, or Ansible is not mandatory, but it's helpful and will get you up to speed quickly. We don't expect anyone to know the full stack upon arrival.
- You communicate clearly in writing. This is important because we work asynchronously, document decisions, write operational notes, and share context across teams.
- You use AI tools such as Claude Code, Cursor, ChatGPT, and others to enhance efficiency and improve your workflows.
- You are fluent in English and Spanish. English is the primary language we use at Tinybird, and Spanish is commonly spoken within the Platform team.
- You are willing to participate in on-call rotations, not only to maintain service health but also to understand the real problems our customers and internal teams face.
- You are located in an EU timezone.