> Markdown version of [/jobs/ext/2568287-principal-cloud-platform-engineer](https://www.wearedevelopers.com/jobs/ext/2568287-principal-cloud-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Cloud Platform Engineer - **Company:** SambaNova Systems, Inc. - **Location:** San Jose, CA, United States - **Experience:** Experienced - **Salary:** $144,000.0 - $189,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Unix, Cloud Computing, Computer Programming, Databases, Computer Engineering, Continuous Delivery, Continuous Integration, Data Centers, DevOps, Github, Python (Programming Language), Linux System Administration, Memcached, NoSQL, Octopus Deploy, Open Source Technology, Redis, Reliability Engineering, Ansible, Prometheus, Software Engineering, SQL Databases, AI Infrastructure, Rust (Programming Language), Datadog, Scripting, Graphics Processing Unit (GPU), Google Cloud, Computer Network Operations, Cloud Platform System, Autoscaling, Grafana, Caching, HybridCloud, Cloudformation, Information Technology, Machine Learning Operations, Hardware Infrastructure, Terraform, Docker, Elk Stack, Jenkins, Golang - **Published:** August 23, 2026 - **Apply:** https://www.careerbuilder.com/job-details/principal-cloud-platform-engineer-san-jose-ca--b34b1baf-798c-4f55-9ae3-2410204b556b ## About the Role * B.S. in Computer Science, Computer Engineering, or related field * 3+ years of experience in a Site Reliability Engineering, DevOps * Experience supporting a large-scale, customer-facing service in a public cloud environment (AWS, GCP, Azure) * Strong programming and scripting skills in languages like Python, Go, Rust, or Java * Proven experience with containerization and orchestration technologies (Docker and Kubernetes) * Deep understanding of monitoring and observability principles and tools (e.g., Prometheus, Grafana, ELK Stack, Datadog) * Experience with Infrastructure as Code (e.g., Terraform, CloudFormation) * Experience with CI/CD principles and tools (e.g., Jenkins, GitHub Actions, ArgoCD) * Strong Linux/Unix system administration fundamentals Preferred Qualifications * Experience in a hybrid environment bridging cloud and on-premise/data center infrastructure. * Direct experience supporting ML/AI inferencing services in production. * Familiarity with GPU-accelerated computing and optimizing workloads for NVIDIA GPUs for purposes of mapping to RDUs. * Knowledge of model serving frameworks like vLLM, SGLang or Ray. * Understanding of MLOps principles and practices. * Experience with managing and tuning databases (SQL or NoSQL) and caching systems (Redis, Memcached)., Accidental Death and Dismemberment (AD&D), Amazon Web Services (AWS), Ansible, Artificial Intelligence (AI), Autoscaling, Business Operations, Caching, Capacity Management, Capacity and Performance Management, Change Management, Cloud Computing, Computer Engineering, Computer Programming, Computer Science, Continuous Deployment/Delivery, Continuous Integration, Cost Control, Customer Experience, Customer Relations, Customer Retention/Renewal, Customer/Client Research, Database Administration, DevOps, Docker, Establish Priorities, Expense Management, Finance, Financial Trend Analysis, Flexible Spending Accounts, Forecasting, GCP (Good Clinical Practices), GPU (Graphics Processing Unit), GitHub, Go Programming Language (Golang), Government Organizations, Healthcare, Hybrid Cloud, Incident Response, Insurance, Java, Jenkins, Linux Administration, Microsoft Windows Azure, Network Operations Center, NoSQL, On Call, Open Source, Product Planning, Public Cloud, Python Programming/Scripting Language, Redis, Reliability Engineering, Reporting Dashboards, Resource Utilization, Rust Programming Language, SQL (Structured Query Language), Scripting (Scripting Languages), Software Development, Software Engineering, Unix System Administration, memcached ## Description As a Principal Cloud Platform Engineer, you will be specializing in our AI Inferencing Service and will be the guardian of its reliability, performance, and scalability. You will bridge the gap between software development and operations, applying an engineering mindset to solve operational challenges. Your primary focus will be ensuring our inference endpoints have exceptional uptime, low-latency response times, and efficient resource utilization, directly impacting the experience of our customers and the success of our AI products. This role includes participating in a shared on-call rotation to maintain 24/7 service reliability., Some of your responsibilities will include: * Shared ownership of the production inferencing service across regions, covering availability, latency, performance, change management, and capacity planning * Standing-up and automating AI infrastructure in new regions * Participating in a shared primary/secondary on-call rotation, and leading incident response * Building monitoring, alerting, and dashboards in Prometheus, Grafana, and Datadog for service health, model latency and throughput, and accelerator utilization * Finding and eliminating performance bottlenecks * Designing auto-scaling policies that handle variable inference loads * Managing cloud and on-prem infrastructure as code in Terraform and Ansible * Building CI/CD pipelines that safely deploy new model versions and service updates * Forecasting infrastructure needs against the product roadmap and usage trends, and working with finance to manage cloud spend * Defining and reporting on SLOs and SLIs for the inferencing platform, using that data to prioritize reliability work ## Related Videos - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [WeAreDevelopers LIVE - Node and Package Security](https://www.wearedevelopers.com/videos/2138-wearedevelopers-live-node-and-package-security) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [The Time Paradox: Building Timezone-Safe Python/Django Applications](https://www.wearedevelopers.com/videos/1915-the-time-paradox-building-timezone-safe-python-django-applications) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)