> Markdown version of [/jobs/ext/3170588-senior-engineer-network-observability](https://www.wearedevelopers.com/jobs/ext/3170588-senior-engineer-network-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Engineer, Network Observability - **Company:** CoreWeave - **Location:** Greater London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Adobe InDesign, Bash Shell, Linux, Internet Protocol, Jinja (Template Engine), Python (Programming Language), Network Troubleshooting, Machine Learning, Routing, Reliability Engineering, Ansible, Tensorflow, Prometheus, Simple Network Management Protocols, Grafana, Scikit Learn, Kubernetes, Infrastructure Automation Frameworks, Information Technology, ArcSight Event Correlation, Data Pipelines, Dynatrace, Golang - **Published:** September 8, 2026 - **Apply:** https://www.collegerecruiter.com/job/2840620181-senior-engineer-network-observability ## About the Role * Deep familiarity with Prometheus, Grafana, Alertmanager, gNMI, SNMP. Experience writing or extending custom metric collectors/exporters. * Experience as a Network Engineer, SRE, Software Developer, or Systems Administrator in large-scale environments. Track record of building and operating robust telemetry and monitoring solutions. * Passion for automating tasks and processes. * Comfortable containerizing solutions in Kubernetes and deploying container-based workloads efficiently. * Proficient with Python, Go, Bash, and familiar with configuration management tools (Ansible, Jinja2). * Strong knowledge of Linux systems and IP networking concepts, including routing, switching, and network troubleshooting. * Practical knowledge with platforms such as Arista EOS, NVIDIA Cumulus Linux, Nokia SR OS, and SR Linux. * Collaborative, humble, and open to learning from senior colleagues., * Bachelor's degree in Computer Science or related field. * Experience applying machine learning for anomaly detection (TensorFlow, scikit-learn). * Network certifications (CCNA, CCNP, etc.). * Experience with data pipelines, event correlation, or anomaly detection in large-scale environments. * Familiarity with OpenTelemetry, Jaeger, or Zipkin for distributed tracing. ## Description * Develop, optimize, and maintain network observability platforms. Use Python and Golang to create collectors, exporters, and dashboards that provide deep visibility into network health and performance. * Collaborate with Network Engineering and Platform teams to ingest and unify logs, metrics, and events from various platforms (Arista EOS, NVIDIA Cumulus Linux, Nokia SR OS, SR Linux, etc.) into a single observability pipeline. * Design and implement scalable telemetry solutions using protocols like gNMI, SNMP, and streaming analytics. Ensure advanced alerting and anomaly detection with Prometheus, Grafana, Alertmanager. * Work closely with network developers, site reliability engineers, and security teams to integrate observability solutions across the broader infrastructure. Participate in design discussions, RFCs, and architectural decisions. * Join a rotating on-call schedule to troubleshoot and resolve observability-related issues. Provide timely support to operations teams, quickly isolating and fixing problems when they arise. * Guide junior team members, share best practices, and foster a culture of continuous learning and improvement within the observability domain. ## Related Videos - [Monitoring as Code - Managing your dashboards at scale](https://www.wearedevelopers.com/videos/753-monitoring-as-code-managing-your-dashboards-at-scale) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)