Data Expert (Analytics + Monitoring + Observability

Runware
UK
17 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Data Analysis BigQuery Data Infrastructure Software Debugging DevOps Distributed Systems Error Codes Monitoring of Systems Python (Programming Language) Machine Learning
+19 more
Node.Js Prometheus Datadog Data Logging Graphics Processing Unit (GPU) Real Time Systems Autoscaling Delivery Pipeline Grafana Caching Backend Fastapi Low Latency Data Analytics Performance Monitor Machine Learning Operations New Relic (SaaS) Data Pipelines Web Api

Job description

Runware is building a high-performance, full-stack AI media-creation platform - empowering developers and companies to generate any type of media instantly. As we scale fast and integrate increasingly complex models, we need stronger visibility, analytics, and monitoring across the whole platform stack.

We’re looking for a Data Expert (Analytics + Monitoring + Observability) to help us better understand, measure, and optimize how the Runware platform performs at scale - internally and for our clients.

Mission

Your main goal is to give Runware full visibility over:

  • End-to-end inference performance
  • Integration usage and model activity
  • Errors, delays, bottlenecks, regressions
  • Internal and client-facing analytics dashboards
  • Health and performance of production pipelines

You will provide the data insights that allow engineering, ML, backend, DevOps, and leadership to make informed decisions - and to continuously improve performance and reliability., * Build and maintain E2E inference time tracking (global and per-model).

  • Monitor how implementation changes impact total request latency.
  • Detect regressions introduced by suboptimal code paths.
  • Provide automated alerts & historical trends.

Usage & Analytics Reporting

  • Build dashboards for internal use (engineering, product, leadership).
  • Provide client-facing usage dashboards (requests, errors, success rate, performance).
  • Support clients who need visibility to debug their integrations.
  • Track model-level usage, API endpoints usage, adoption metrics, etc., * Implement metrics, logs, and traces that help the entire platform scale smoothly.
  • Work closely with DevOps & backend teams to improve system observability.
  • Provide insights that guide infra decisions (GPU allocation, autoscaling, caching, batching, etc.)., * Select and maintain tooling (e.g., Prometheus/Grafana, Datadog, OpenTelemetry, ELK, BigQuery, etc.).
  • Ensure data pipelines are reliable, accessible, and always up-to-date.
  • Build simple, easy-to-read dashboards for both technical and non-technical teams.

Requirements

  • Strong experience with data analytics, observability, or monitoring
  • Hands-on with metrics/logging/tracing frameworks (Prometheus, Grafana, Datadog, New Relic, etc.)
  • Good understanding of backend systems and distributed architectures
  • Ability to turn raw metrics into actionable insights
  • Experience building dashboards for internal and external stakeholders
  • Familiarity with AI model monitoring (latency, throughput, error codes, GPU utilization)

Nice-to-Have

  • Experience with AI/ML infrastructure, inference pipelines, GPUs
  • Understanding of Python APIs, FastAPI, or Node environments
  • Experience working with high-throughput real-time systems
  • Startup or scale-up experience, * A problem-solver mindset
  • Proactivity - you like digging into the data and flagging problems before anyone else sees them
  • Ability to work with ML, backend, DevOps, and product teams
  • Comfort with autonomous ownership, Please note: We are unable to offer visa sponsorship in the UK at this time. Candidates must have existing right to work in the UK.

Benefits & conditions

Our release cycles are fast and intense, but they’re followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap.

  • Generous paid time off - vacation, sick days, public holidays
  • Meaningful stock options - share in the upside you create
  • Remote-first setup - work from home anywhere we can employ you
  • Flexible hours - own your schedule outside core collaboration blocks
  • Family leave - paid maternity, paternity, and caregiver time
  • Company retreats - twice-yearly gatherings in inspiring locations

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

9:56 min

Expanding browser capabilities with modern web APIs

Ire Aderinokun · JS Congress

3:33 min

Connecting frontends via a FastAPI proxy backend layer

Saoussen Chaabnia Saoussen Chaabnia · Europe 2026 Virtual

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

1:42 min

Introduction to the fast API web framework

Sebastián Ramírez · World Congress 2022

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all