Principal Cloud Platform Engineer

SambaNova Systems, Inc.
San Jose, CA, United States
13 days ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$144,000.0 - $189,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Amazon Web Services Microsoft Azure Unix Cloud Computing Computer Programming Databases Computer Engineering Continuous Delivery Continuous Integration Data Centers
+35 more
DevOps Github Python (Programming Language) Linux System Administration Memcached NoSQL Octopus Deploy Open Source Technology Redis Reliability Engineering Ansible Prometheus Software Engineering SQL Databases AI Infrastructure Rust (Programming Language) Datadog Scripting Graphics Processing Unit (GPU) Google Cloud Computer Network Operations Cloud Platform System Autoscaling Grafana Caching HybridCloud Cloudformation Information Technology Machine Learning Operations Hardware Infrastructure Terraform Docker Elk Stack Jenkins Golang

Job description

As a Principal Cloud Platform Engineer, you will be specializing in our AI Inferencing Service and will be the guardian of its reliability, performance, and scalability. You will bridge the gap between software development and operations, applying an engineering mindset to solve operational challenges. Your primary focus will be ensuring our inference endpoints have exceptional uptime, low-latency response times, and efficient resource utilization, directly impacting the experience of our customers and the success of our AI products. This role includes participating in a shared on-call rotation to maintain 24/7 service reliability., Some of your responsibilities will include:

  • Shared ownership of the production inferencing service across regions, covering availability, latency, performance, change management, and capacity planning
  • Standing-up and automating AI infrastructure in new regions
  • Participating in a shared primary/secondary on-call rotation, and leading incident response
  • Building monitoring, alerting, and dashboards in Prometheus, Grafana, and Datadog for service health, model latency and throughput, and accelerator utilization
  • Finding and eliminating performance bottlenecks
  • Designing auto-scaling policies that handle variable inference loads
  • Managing cloud and on-prem infrastructure as code in Terraform and Ansible
  • Building CI/CD pipelines that safely deploy new model versions and service updates
  • Forecasting infrastructure needs against the product roadmap and usage trends, and working with finance to manage cloud spend
  • Defining and reporting on SLOs and SLIs for the inferencing platform, using that data to prioritize reliability work

Requirements

  • B.S. in Computer Science, Computer Engineering, or related field
  • 3+ years of experience in a Site Reliability Engineering, DevOps
  • Experience supporting a large-scale, customer-facing service in a public cloud environment (AWS, GCP, Azure)
  • Strong programming and scripting skills in languages like Python, Go, Rust, or Java
  • Proven experience with containerization and orchestration technologies (Docker and Kubernetes)
  • Deep understanding of monitoring and observability principles and tools (e.g., Prometheus, Grafana, ELK Stack, Datadog)
  • Experience with Infrastructure as Code (e.g., Terraform, CloudFormation)
  • Experience with CI/CD principles and tools (e.g., Jenkins, GitHub Actions, ArgoCD)
  • Strong Linux/Unix system administration fundamentals

Preferred Qualifications

  • Experience in a hybrid environment bridging cloud and on-premise/data center infrastructure.
  • Direct experience supporting ML/AI inferencing services in production.
  • Familiarity with GPU-accelerated computing and optimizing workloads for NVIDIA GPUs for purposes of mapping to RDUs.
  • Knowledge of model serving frameworks like vLLM, SGLang or Ray.
  • Understanding of MLOps principles and practices.
  • Experience with managing and tuning databases (SQL or NoSQL) and caching systems (Redis, Memcached)., Accidental Death and Dismemberment (AD&D), Amazon Web Services (AWS), Ansible, Artificial Intelligence (AI), Autoscaling, Business Operations, Caching, Capacity Management, Capacity and Performance Management, Change Management, Cloud Computing, Computer Engineering, Computer Programming, Computer Science, Continuous Deployment/Delivery, Continuous Integration, Cost Control, Customer Experience, Customer Relations, Customer Retention/Renewal, Customer/Client Research, Database Administration, DevOps, Docker, Establish Priorities, Expense Management, Finance, Financial Trend Analysis, Flexible Spending Accounts, Forecasting, GCP (Good Clinical Practices), GPU (Graphics Processing Unit), GitHub, Go Programming Language (Golang), Government Organizations, Healthcare, Hybrid Cloud, Incident Response, Insurance, Java, Jenkins, Linux Administration, Microsoft Windows Azure, Network Operations Center, NoSQL, On Call, Open Source, Product Planning, Public Cloud, Python Programming/Scripting Language, Redis, Reliability Engineering, Reporting Dashboards, Resource Utilization, Rust Programming Language, SQL (Structured Query Language), Scripting (Scripting Languages), Software Development, Software Engineering, Unix System Administration, memcached

Benefits & conditions

Base Salary Range:

Base Pay Range

$144,000-$189,000 USD

Submission Guidelines

Please note that in order to be considered an applicant for any position at SambaNova Systems, you must submit an application form for each position for which you believe you are qualified.

EEO Policy

SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.

Benefits Summary for US-Based, Full-Time Employment Positions

SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution. We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care. Our library of well-being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.

About the company

Title: Principal Cloud Platform Engineer Organization: SambaNova Systems, Inc. Location: San Jose Description:

The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale.

SambaNova Suite is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy. Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets.

About the team

The Cloud Platform team owns the production inferencing service that serves SambaNova’s models to customers on RDU accelerators, including capacity planning, deployment, monitoring, and incident response across regions in the United States, Asia, Europe, and Latin America.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:03 min

Microsoft integrating native Unix coreutils into Windows environments

Chris Heilmann +2 · LIVE

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:04 min

Defining timestamps and the international standard format

Denny Biasiolli Denny Biasiolli · Europe 2026 Virtual

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all