> Markdown version of [/jobs/ext/2253623-senior-software-engineer-fleet-intelligence-backend](https://www.wearedevelopers.com/jobs/ext/2253623-senior-software-engineer-fleet-intelligence-backend). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, Fleet Intelligence Backend - **Company:** NVIDIA Corporation - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $152,000.0 - $241,500.0 - **Contract:** Permanent contract - **Skills:** Query Performance, Application Programming Interfaces (APIs), Artificial Intelligence, Amazon S3, User Authentication, Cloud Computing, Databases, Computer Engineering, Software Debugging, Linux, Distributed Systems, PostgreSQL, Operational Data Store, Query Optimization, Cloud Services, Prometheus, Software Engineering, Graphics Processing Unit (GPU), Cloud Platform System, Grafana, Backend, Kubernetes, Information Technology, Apache Kafka, Cloudwatch, Restful APIs, Amazon Simple Queue Service (SQS), Docker - **Published:** August 26, 2026 - **Apply:** https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Software-Engineer--Fleet-Intelligence-Backend_JR2023946 ## About the Role * 5+ years of industry software engineering experience with a Bachelor's or Master's degree in Computer Science, Computer Engineering, or equivalent experience. * Strong Go backend development experience. * Experience building REST APIs and production services using frameworks such as Gin or similar. * Strong PostgreSQL experience, including schema design, query tuning, transactions, migrations, and operational data modeling. * Experience with distributed systems, event pipelines, telemetry ingestion, or operational analytics. * Experience with cloud infrastructure, Docker, Kubernetes, and Helm. * Strong debugging skills across APIs, databases, cloud services, and production systems. * Familiarity with authentication and authorization patterns such as JWT, service account keys, trusted-edge proxies, or customer-scoped APIs. * Experience with Linux-based development and production environments. Ways To Stand Out From The Crowd: * Background with telemetry, monitoring, observability, health, or fleet-management platforms. * Experience with AWS services such as MSK, Aurora, S3, SQS, AMP, or CloudWatch. * Experience with OpenTelemetry, Prometheus, LightStep, Grafana, or production tracing/metrics systems. * Background with NVIDIA GPUs, DGX systems, DCGM, XID analysis, attestation, or AI datacenter operations. * Experience designing APIs and storage systems that support real-time operational workflows at fleet scale. ## Description This role focuses on backend services for GPU health, Fleet Intelligence, telemetry ingestion, inventory, attestation, alerting, reporting, and cloud operations automation, in addition, you will: * Design and develop Go backend services, REST APIs, and data models for GPU Health and Fleet Intelligence. * Build customer-facing and agent-facing APIs for compute zones, node groups, nodes, alerts, events, metrics, inventory, attestation, retention policies, and reports. * Develop high-volume ingestion and persistence paths for in-band agents and out-of-band collectors. * Work with Aurora PostgreSQL, Kafka/MSK, S3, SQS, Prometheus/AMP, and OpenTelemetry-based observability. * Build and operate scheduled backend services for liveness tracking, alerting, rollups, cleanup, notification delivery, attestation, and XID analysis. * Optimize database schemas, partitioned time-series storage, query performance, CTE-heavy queries, and connection pooling for reliable service behavior. * Collaborate with agent, infrastructure, SRE, UI, and cloud operations teams to turn operational workflows into scalable backend systems. * Improve service reliability, security, observability, testing, and deployment quality across Docker, Kubernetes, Helm, and cloud environments. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Backend Developer Salary in Germany [2023]](https://www.wearedevelopers.com/magazine/196-backend-developer-salary-in-germany-2023)