> Markdown version of [/jobs/ext/2708729-software-engineer-network-observability](https://www.wearedevelopers.com/jobs/ext/2708729-software-engineer-network-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - Network Observability - **Company:** CLOCKWORK SYSTEMS, INC. - **Location:** Palo Alto, United States - **Experience:** Expert - **Salary:** $140,000.0 - $210,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, C++ (Programming Language), Cloud Computing, Configuration Management, Computer Programming, Computer Networks, Computer Engineering, Network Congestion, Software Debugging, Distributed Computing Environment, Distributed Systems, Domain Name System (DNS), Hypertext Transfer Protocols (HTTP), InfiniBand, Python (Programming Language), Network Layer, Linux Kernel, Linux System Administration, Network Diagnostics, Network Monitoring, Routing, Performance Tuning, Remote Direct Memory Access, Prometheus, Service Discovery, System Programming, TCP/IP, Transmission Control Protocol (TCP), Tcpdump, Traffic Analysis, Wireshark, Rust (Programming Language), Datadog, Grafana, Backend, Perf (Linux), Kubernetes, Infrastructure Automation Frameworks, Information Technology, Low Latency, Splunk, Golang - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-software-engineer-network-observability-clockworkio-8129231 ## About the Role * Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field. * Strong hands-on programming experience in C++, Go, Python, Rust, or similar systems programming languages. * Proven experience delivering complex infrastructure projects and contributing to the design and implementation of distributed systems. * Experience building distributed systems, backend services, telemetry pipelines, or observability platforms. * Hands-on experience with RDMA, RoCE, InfiniBand, or other high-performance network fabrics. * Familiarity with libibverbs, RDMA verbs, RDMA CM, queue pairs, completion queues, memory registration, and related RDMA concepts. * Strong knowledge of Linux networking, TCP/IP, DNS, HTTP, routing, MTU, congestion control, packet loss, latency, and performance tuning. * Experience with traceroute-style diagnostics, path discovery, network reachability checks, synthetic probes, or active network measurements. * Experience with monitoring and visualization platforms such as Prometheus, Grafana, Datadog, Splunk, OpenTelemetry, or similar tools. * Strong debugging skills across software, operating system, server, and network layers. * Experience operating production systems in Linux-based environments. * Strong technical judgment and ability to build reliable, scalable, and maintainable systems., * Experience supporting AI/ML, HPC, storage, or GPU cluster infrastructure workloads. * Experience with large-scale RoCE or InfiniBand deployments. * Experience with NCCL, distributed training infrastructure, or AI cluster diagnostics. * Experience with eBPF, XDP, DPDK, perf, tcpdump, Wireshark, ethtool, iproute2, rdma-core, or Linux kernel networking tools. * Experience with cloud infrastructure on AWS, GCP, or Azure. * Experience with Kubernetes, service discovery, configuration management, and infrastructure automation. * Knowledge of security, compliance, and infrastructure best practices. * Experience designing time-series data systems, alerting pipelines, or high-cardinality telemetry platforms. ## Description We are seeking an experienced Senior Software Engineer to contribute to the architecture, development, and scaling of a high-performance network monitoring and observability platform. This role will focus on building systems that provide deep visibility into RDMA, RoCE, InfiniBand, and TCP/IP networks. The ideal candidate has strong experience in distributed systems, Linux networking, and modern observability stacks (e.g., Grafana/Prometheus)., * Design, develop, and scale high-performance network monitoring platforms for RDMA, RoCE, InfiniBand, and TCP/IP infrastructure. * Build backend telemetry services, observability dashboards, alerts, diagnostics, anomaly detection, SLA monitoring, and traffic analysis workflows. * Troubleshoot complex production issues across application, OS, server, RDMA, and network layers while optimizing low-latency collection, aggregation, and alerting. * Collaborate with cross-functional teams to deliver scalable solutions, improve engineering practices, automate operational workflows, and contribute to the technical direction of the platform. ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries)