Senior Engineer, Network Observability

CoreWeave
Greater London, UK
26 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Adobe InDesign Bash Shell Linux Internet Protocol Jinja (Template Engine) Python (Programming Language) Network Troubleshooting Machine Learning Routing Reliability Engineering Ansible Tensorflow
+11 more
Prometheus Simple Network Management Protocols Grafana Scikit Learn Kubernetes Infrastructure Automation Frameworks Information Technology ArcSight Event Correlation Data Pipelines Dynatrace Golang

Job description

  • Develop, optimize, and maintain network observability platforms. Use Python and Golang to create collectors, exporters, and dashboards that provide deep visibility into network health and performance.
  • Collaborate with Network Engineering and Platform teams to ingest and unify logs, metrics, and events from various platforms (Arista EOS, NVIDIA Cumulus Linux, Nokia SR OS, SR Linux, etc.) into a single observability pipeline.
  • Design and implement scalable telemetry solutions using protocols like gNMI, SNMP, and streaming analytics. Ensure advanced alerting and anomaly detection with Prometheus, Grafana, Alertmanager.
  • Work closely with network developers, site reliability engineers, and security teams to integrate observability solutions across the broader infrastructure. Participate in design discussions, RFCs, and architectural decisions.
  • Join a rotating on-call schedule to troubleshoot and resolve observability-related issues. Provide timely support to operations teams, quickly isolating and fixing problems when they arise.
  • Guide junior team members, share best practices, and foster a culture of continuous learning and improvement within the observability domain.

Requirements

  • Deep familiarity with Prometheus, Grafana, Alertmanager, gNMI, SNMP. Experience writing or extending custom metric collectors/exporters.
  • Experience as a Network Engineer, SRE, Software Developer, or Systems Administrator in large-scale environments. Track record of building and operating robust telemetry and monitoring solutions.
  • Passion for automating tasks and processes.
  • Comfortable containerizing solutions in Kubernetes and deploying container-based workloads efficiently.
  • Proficient with Python, Go, Bash, and familiar with configuration management tools (Ansible, Jinja2).
  • Strong knowledge of Linux systems and IP networking concepts, including routing, switching, and network troubleshooting.
  • Practical knowledge with platforms such as Arista EOS, NVIDIA Cumulus Linux, Nokia SR OS, and SR Linux.
  • Collaborative, humble, and open to learning from senior colleagues., * Bachelor’s degree in Computer Science or related field.
  • Experience applying machine learning for anomaly detection (TensorFlow, scikit-learn).
  • Network certifications (CCNA, CCNP, etc.).
  • Experience with data pipelines, event correlation, or anomaly detection in large-scale environments.
  • Familiarity with OpenTelemetry, Jaeger, or Zipkin for distributed tracing.

Benefits & conditions

  • Family-level medical insurance.
  • Family-level dental insurance.
  • Generous pension contribution.
  • Life assurance at 4× salary.
  • Critical illness cover.
  • Employee assistance programme.
  • Tuition reimbursement.
  • Work culture focused on innovative disruption.

All candidates must undergo a basic criminal record check in compliance with GDPR. Employment offers are conditional upon satisfactory results.

About the company

CoreWeave is a cloud platform that empowers AI innovation. Founded in 2017, it provides infrastructure, tools, and expertise to improve performance.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

Videos

See all

Related articles

See all