Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
NVIDIA Ltd.
Santa Clara, United States of America
3 days ago
Role details
Contract type
Permanent contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
English Experience level
Senior Compensation
$ 288KJob location
Santa Clara, United States of America
Tech stack
API
Artificial Intelligence
Amazon Web Services (AWS)
Azure
Cloud Computing
Cloud Database
Computer Clusters
Continuous Integration
Data Centers
Cursor (Graphical User Interface Elements)
DevOps
Python
Octopus Deploy
Ansible
Prometheus
Software Engineering
Vault (Revision Control System)
Datadog
Cloud Platform System
GitHub Copilot
Autoscaling
Grafana
Multi-Cloud
Backend
Gitlab-ci
Kubernetes
Information Technology
Bare Metal
REST
Terraform
gRPC
Jenkins
Microservices
Job description
- Build and develop backend microservices and REST/gRPC/MCP APIs that power the deployment platform, zone reservation/lease system, and developer self-service tooling
- Extend the platform with dynamic delivery, automatic rollback, drift detection, and automated zone bootstrapping
- Build and stabilize prod-like staging environments/zones so teams catch regressions early
- Develop and sustain GitOps pipelines (Flux CD, Argo CD) across on-premises and Nvidia GFN Cloud/AWS/Azure/GCP
- Develop Kubernetes CRDs and operators in Go for scheduling, auto-scaling, and compliance across data centers
- Build backend integrations and control-plane services connecting CI/CD, observability, and automation systems into a unified platform experience
- Automate dedicated hardware and multiple cloud platform configurations using Terraform, Ansible, and Vault
- Implement monitoring solutions including Prometheus, Grafana, Datadog, and ELK, paired with SLO/alerting for early detection
- Integrate automation tools - runbooks, StackStorm bots, anomaly-triggered remediation, and Slack self-service release bots
Requirements
- Bachelor or higher degree in computer science, engineering, or equivalent experience.
- 10+ years of experience in Cloud Infrastructure and DevOps, with deep expertise in Kubernetes, GitOps (or equivalent), and production-grade cloud-native CI/CD pipelines
- Expert Kubernetes: CRDs, operators, multi-cluster management, and security hardening (CIS, PCI/SOC 2)
- Proficiency in Flux CD or Argo CD, GitLab CI, and Jenkins; Go or Python for control plane development
- Experience handling Vault, Terraform, Ansible, and Helm in both on-premises and cloud environments
- Experience in developing and scaling RESTful, gRPC, MCP APIs and backend services.
- Experience running hybrid multi-cloud and bare-metal at production scale, and owning a platform roadmap end-to-end .
Ways to Stand Out from the crowd:
- Backstage or similar internal developer portals for self-service tooling
- Comprehensive expertise covering backend services as well as platform and infrastructure engineering
- Daily use of AI-assisted tools (Claude Code, GitHub Copilot, Cursor) - we use these daily
- Hybrid infrastructure spanning on-premises GPU clusters and public cloud
Benefits & conditions
With competitive salaries and a generous benefits package, NVIDIA is widely considered one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.
You will also be eligible for equity and benefits .
About the company
GeForce NOW is NVIDIA's Cloud Gaming service, streaming games at the highest quality to any and every user, regardless of their device type and capabilities - low-end PCs, Macs, TV or mobile devices. Using the most sophisticated GPUs and NVIDIA proprietary software, GeForce NOW transforms the gaming experience with always up-to-date games on always the latest hardware, a streaming experience rivaling that of a local PC, and near-instant launch - just click and play! For more details, see http://www.geforce.com/geforce-now .
Our mission empowers every GeForce NOW engineer to deliver confidently by removing obstacles, catching issues early, and turning deployments into a strength rather than a liability. We handle pipelines using GitLab's CI platform and Jenkins, along with Flux CD and Argo CD rollouts, StackStorm event-driven automation, HashiCorp Vault for credential storage, and automation across multi-cloud environments.
What You'll Own
This is the Kubernetes-native deployment platform paired with its automation system and the backend services and APIs that support it. We deliver zero-downtime global rollouts with automatic drift detection. Our reliable staging environments replicate production accurately. We provide cloud automation for bare-metal and multi-cloud environments. If a GFN service ships, runs, or self-heals, you build the infrastructure behind it.