Senior DevOps Engineer - E-commerce

NVIDIA Ltd.
Las Cruces, NM, United States
26 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$100,000.0 - $130,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Amazon Cloudfront Amazon S3 Continuous Integration DevOps Domain Name System (DNS) Python (Programming Language) Reliability Engineering Akamai Software Engineering Datadog
+14 more
Scripting Software Modules System Availability Delivery Pipeline Reliability of Systems Amazon Relational Database Service Kubernetes Information Technology Deployment Automation Api Gateway Amazon Simple Queue Service (SQS) Terraform Docker Jenkins

Job description

We are looking for an outstanding DevOps and Site Reliability Engineer to join the NVIDIA e-commerce team. You will be a key architect of our e-commerce platform, ensuring that our systems are scalable, resilient, and automated. The ideal candidate is a Terraform expert who views infrastructure as code (IaC) not just as a tool, but as a philosophy. You will bridge the gap between development and operations, focusing on system reliability, high availability, and the performance of our global e-commerce platform.

What you’ll be doing:

  • Architect and refine automated deployment Jenkins pipelines to ensure seamless, zero-downtime releases.
  • Design, build, and maintain enterprise-scale infrastructure using Terraform. Establish modular, reusable patterns for AWS resources.
  • Optimize and manage sophisticated AWS environments with a focus on cost-efficiency and security.
  • Transition our monitoring from reactive to proactive using AI-powered observability tools (e.g., Datadog Watchdog) for automated root cause analysis (RCA) and anomaly detection.
  • Define and monitor SLOs and SLAs. Lead incident response and conduct thorough post-mortems to improve system resilience.

Requirements

  • 8+ years or equivalent industry experience
  • Bachelor’s/Master’s Degree in Computer Science, Software Engineering, or equivalent experience.
  • Exceptionally strong background in developing CI/CD processes and deployment pipelines using Jenkins.
  • Extensive experience architecting on AWS Cloud and running services such as API Gateway, Lambda, EKS/ECS, RDS, S3, and SQS.
  • Expert-level knowledge of Terraform (including state management, workspaces, and complex module development).
  • Advanced experience with Kubernetes (EKS) and Docker, including orchestration, service meshes, and Helm.
  • Strong proficiency in a scripting language, such as Python, for automation and custom tooling.
  • Strong communication skills.

Ways to stand out from the crowd:

  • Deep understanding of DNS and CDNs (e.g., Akamai, CloudFront).
  • Demonstrated use of AI tools to improve productivity and the quality of releases.
  • Applies secure-by-design principles across infrastructure, deployment automation, and operational processes.
  • AWS certifications are preferred.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000 USD - 276,000 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5., $100,000.00 per year

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jofdav.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:34 min

Leveraging Akamai edge workers for broad geographic scale

Austin Gil · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · WWC Europe 2026

1:58 min

Application performance and its direct business impact

Jérôme Vieilledent · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all