Senior Software Engineer, Fleet Intelligence Backend

NVIDIA Corporation
Santa Clara, CA, United States
3 days ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$152,000.0 - $241,500.0
Working hours
Regular working hours

Tech stack

Query Performance Application Programming Interfaces (APIs) Artificial Intelligence Amazon S3 User Authentication Cloud Computing Databases Computer Engineering Software Debugging Linux Distributed Systems PostgreSQL
+16 more
Operational Data Store Query Optimization Cloud Services Prometheus Software Engineering Graphics Processing Unit (GPU) Cloud Platform System Grafana Backend Kubernetes Information Technology Apache Kafka Cloudwatch Restful APIs Amazon Simple Queue Service (SQS) Docker

Job description

This role focuses on backend services for GPU health, Fleet Intelligence, telemetry ingestion, inventory, attestation, alerting, reporting, and cloud operations automation, in addition, you will:

  • Design and develop Go backend services, REST APIs, and data models for GPU Health and Fleet Intelligence.
  • Build customer-facing and agent-facing APIs for compute zones, node groups, nodes, alerts, events, metrics, inventory, attestation, retention policies, and reports.
  • Develop high-volume ingestion and persistence paths for in-band agents and out-of-band collectors.
  • Work with Aurora PostgreSQL, Kafka/MSK, S3, SQS, Prometheus/AMP, and OpenTelemetry-based observability.
  • Build and operate scheduled backend services for liveness tracking, alerting, rollups, cleanup, notification delivery, attestation, and XID analysis.
  • Optimize database schemas, partitioned time-series storage, query performance, CTE-heavy queries, and connection pooling for reliable service behavior.
  • Collaborate with agent, infrastructure, SRE, UI, and cloud operations teams to turn operational workflows into scalable backend systems.
  • Improve service reliability, security, observability, testing, and deployment quality across Docker, Kubernetes, Helm, and cloud environments.

Requirements

  • 5+ years of industry software engineering experience with a Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or equivalent experience.
  • Strong Go backend development experience.
  • Experience building REST APIs and production services using frameworks such as Gin or similar.
  • Strong PostgreSQL experience, including schema design, query tuning, transactions, migrations, and operational data modeling.
  • Experience with distributed systems, event pipelines, telemetry ingestion, or operational analytics.
  • Experience with cloud infrastructure, Docker, Kubernetes, and Helm.
  • Strong debugging skills across APIs, databases, cloud services, and production systems.
  • Familiarity with authentication and authorization patterns such as JWT, service account keys, trusted-edge proxies, or customer-scoped APIs.
  • Experience with Linux-based development and production environments.

Ways To Stand Out From The Crowd:

  • Background with telemetry, monitoring, observability, health, or fleet-management platforms.
  • Experience with AWS services such as MSK, Aurora, S3, SQS, AMP, or CloudWatch.
  • Experience with OpenTelemetry, Prometheus, LightStep, Grafana, or production tracing/metrics systems.
  • Background with NVIDIA GPUs, DGX systems, DCGM, XID analysis, attestation, or AI datacenter operations.
  • Experience designing APIs and storage systems that support real-time operational workflows at fleet scale.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

About the company

Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPUs act as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent.

As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA is widely recognized as one of the most desirable employers, with some of the most versatile people in the world working for us. If you’re passionate about building scalable, efficient backend systems to power cloud operations, we invite you to join our team.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

10:40 min

Evaluating automotive software architectures and backend technologies

Georg Kühberger +1 · LIVE

1:38 min

Transitioning into backend engineering from web development

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all