> Markdown version of [/jobs/ext/2551585-senior-software-engineer-core-infrastructure-services-dgx-cloud](https://www.wearedevelopers.com/jobs/ext/2551585-senior-software-engineer-core-infrastructure-services-dgx-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, Core Infrastructure Services - DGX Cloud - **Company:** NVIDIA Ltd. - **Location:** United States - **Experience:** Expert - **Salary:** $168,000.0 - $322,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Border Gateway Protocol, Cloud Computing, Databases, Linux, Distributed Systems, Domain Name System (DNS), InfiniBand, Virtual Private Networks (VPN), Python (Programming Language), Enterprise Messaging Systems, Network Architecture, Routing, OAuth, Remote Direct Memory Access, Redis, Ansible, Prometheus, Workflow Management Systems, Network Routers, Load Balancing, Grafana, Firewalls (Computer Science), Kubernetes, Infrastructure Automation Frameworks, Iptables, Apache Kafka, Free and Open-Source Software, Amazon Simple Queue Service (SQS), Terraform, Microservices - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=eb38f376e80c61a8 ## About the Role * BS or equivalent experience with 8+ years of relevant industry experience. * Strong proficiency in Python and Go, with experience building production-quality software. * Experience building cloud-native microservices and APIs on Kubernetes using frameworks such as FastAPI, gRPC, or REST. * Experience with infrastructure automation (Terraform, Ansible), workflow orchestration (Temporal), and distributed systems using databases, Redis, and messaging platforms (Kafka, NATS, SQS). * Experience designing, building, and operating production infrastructure services such as DNS, NTP, AAA (RADIUS/OAuth), and observability platforms. * Strong Linux fundamentals with experience in observability (Prometheus, Grafana, OpenTelemetry, gNMI), networking (BGP, switching, routing, load balancing), and security (VPNs, firewalls, iptables/nftables). * Excellent problem-solving, communication, and collaboration skills. Ways to stand out from the crowd: * Hands-on experience with network infrastructure including switches, routers, and firewalls. * Familiarity with InfiniBand, RDMA, and AI/HPC networking.Experience with NetBox, Nautobot, or similar network source of truth platforms. * Contributions to open-source software. Experience with public cloud platforms. ## Description * Develop software that enables infrastructure orchestration, self-service workflows, and platform automation. * Own integrations with internal and external platforms to automate infrastructure provisioning and lifecycle management. * Build observability and security capabilities that improve the reliability and resilience of our infrastructure. * Partner with infrastructure and networking teams to deliver production services at scale. * Drive operational excellence through automation, monitoring, incident response, and continuous improvement. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Keeping applications secure by evolving OAuth 2.0 and OpenID Connect](https://www.wearedevelopers.com/videos/100152-keeping-applications-secure-by-evolving-oauth-2-0-and-openid-connect) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [The Fastest-Growing Tech Sectors to Look Out for in 2025](https://www.wearedevelopers.com/magazine/373-the-fastest-growing-tech-sectors-to-look-out-for-in-2025)