> Markdown version of [/jobs/ext/3044362-software-engineer-grafana-cloud-observability-provider](https://www.wearedevelopers.com/jobs/ext/3044362-software-engineer-grafana-cloud-observability-provider). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - Grafana Cloud Observability Provider - **Company:** - Tackle meaningful work in a high-growth, ever-evolving environment. - **Location:** Ventas, Spain (Remote available) - **Salary:** €94,025.0 - €112,830.0 - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), Amazon Web Services, Automation of Tests, Software as a Service, Cloud Computing, Code Review, Computer Programming, Programming Tools, Distributed Systems, Python (Programming Language), Load Testing, Reliability Engineering, Site Reliability Engineering Practices, Software Engineering, Datadog, Performance Testing, Grafana, Kubernetes, OPUS (Software), GPT, Docker, Golang - **Published:** September 24, 2026 - **Apply:** https://www.adzuna.es/contact-us.html ## About the Role Strong programming background in a modern language (Python and Go are primary, but prior experience is not required) Experience designing, building, and operating large-scale distributed systems Strong experience with SRE practices, including operating and evolving production systems at scale Strong understanding of reliability engineering concepts (e.g. incident management, observability, and failure modes) Strong experience of defining or applying SLIs/SLOs, error budgets, or reliability metrics Experience with test automation, including performance and functional testing Ability to influence engineering practices through clear technical communication, reviews, and collaboration Strong interpersonal skills and ability to work effectively across teams Familiarity with modern software engineering processes and delivery practices Self-driven and comfortable operating with a high degree of autonomy and ambiguity Experience participating in blameless incident response and writing high-quality post-incident reviews. Bonus Points for: Experience with containerized and cloud-native systems (Docker, Kubernetes, AWS) Familiarity with observability tooling and platforms (e.g. the Grafana stack) Experience working with Python, Go, JavaScript and/or Jsonnet Experience building or operating event-driven or asynchronous systems Interest in, or experience with, building testing frameworks or developer tooling ## Description We are the team behind Grafana k6, Grafana Cloud k6, and Grafana Cloud Synthetics, used by teams globally to ensure resilient, high-performing systems. This opportunity is with the Grafana Cloud k6 squad, who build and operate our performance testing product. Grafana Cloud k6 is built around the OSS k6 and targeted at users looking to run performance tests at scale. Our enterprise and SaaS offerings allow customers to load test their systems by running distributed tests from 15+ regions worldwide, using hundreds of thousands of virtual users sending millions of requests per second. We ingest huge volumes of data generated by k6, which can be used to view, correlate and analyze metrics from each test. k6 is a product used by other engineers, and as such, we are looking for people enthusiastic about building high-quality tools they would want to use themselves. Due to our small teams and fast development pace, you will have a substantial and immediate impact on how the end product is architected, developed, and how the engineering team operates. Your role will focus on establishing and scaling a cross-team culture of engineering excellence by setting standards and guiding adoption of strong engineering practices that improve reliability and operational ownership. As this foundation matures, the role is expected to expand into broader application and product development leadership, contributing architectural and technical depth beyond operational excellence. What will you be doing? Contribute hands-on to the codebase by designing and implementing production-quality software. Guide teams in the design, development, evolution, and operation of large-scale, distributed cloud systems. Build and scale a strong culture of operational excellence by defining standards and coaching teams to own reliability and availability. Help mature SRE practices, including incident response and PIRs, on-call readiness, runbooks, alerting, observability, and release/change management. Establish reliability frameworks such as SLIs/SLOs and error budgets, and use them to guide prioritization and engineering trade-offs. Provide visibility into system health through clear operational metrics and reliability reporting. Participate in the on-call rotation as a primary escalation point and contribute to incident resolution. Influence product and system direction through design reviews, architectural discussions, and cross-team collaboration. Share knowledge through clear, high-quality documentation and technical communication-internally and, where appropriate, externally-to help teams build and operate systems more effectively. As the reliability foundation matures, grow into broader application and product development leadership, contributing architectural and technical depth beyond operations. We invest heavily in developer productivity. You can use modern AI coding assistants as part of your daily workflow (your choice of tools, within security guidelines), backed by a company-funded usage budget so you can iterate quickly without unnecessary friction. We encourage pragmatic AI-assisted development: faster prototyping, test generation, refactors, documentation, and incident follow-ups-always paired with strong code review and quality standards. You'll also have access to frontier models (e.g., GPT-Codex 5/3, Claude Opus 4.6, Gemini 3 Pro). ## Related Videos - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Speak, Code, Deploy: Transforming Developer Experience with Voice Commands](https://www.wearedevelopers.com/videos/1159-speak-code-deploy-transforming-developer-experience-with-voice-commands) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers)