> Markdown version of [/jobs/ext/244636-principal-cloud-platform-engineer](https://www.wearedevelopers.com/jobs/ext/244636-principal-cloud-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Cloud Platform Engineer - **Company:** Rch Solutions - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Query Performance, Artificial Intelligence, Computing Platforms, Microsoft Azure, BigQuery, Cloud Engineering, Computer Networks, Databases, Distributed Systems, Elasticsearch, Identity and Access Management, Node.Js, Platform as a Service (PAAS), Role-Based Access Control, Prometheus, Search Technologies, Management of Software Versions, AI Infrastructure, Google Cloud, Load Balancing, Cloud Platform System, Autoscaling, Large Language Models, Grafana, Generative AI, Backend, AI Platforms, Kubernetes, Low Latency, Terraform - **Published:** May 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=71eb33896a7cd32a ## About the Role Do you have experience in Terraform?, * 5+ years hands-on background in high-scale platform engineering (internal platforms, PaaS, or shared infra) * Deep Kubernetes Platform Expertise + Hands-on experience with GKE: o Cluster upgrades, node pool management, autoscaling o Managing failures, disruptions, and complex maintenance scenarios o RBAC, namespaces, network policies * GCP IAM, Workload Identity, Secret Manager * GCP Storage: BigQuery, GCS, Firestore * Terraform and IaaC experience with GitOps workflows (ArgoCD, Flux or equivalent) * Strong observability practices using: + Google Cloud Operations Suite (Stackdriver) + Prometheus / Grafana * Hands-on experience operating vector databases in production, ideally Weaviate: + Query performance tuning + Cluster stability and scaling behavior * Distributed Systems & GCP Architecture + Solid understanding of distributed systems design and failure modes + Multi-zone / regional architectures + Google Cloud Load Balancing, * Experience with Elasticsearch, OpenSearch, Azure AI Search or similar distributed search systems. * Experience with Vector DBs other than Weviate: Milvus, Pinecone, Qdrant or pgvector * Experience in designing, building and maintaining end-to-end observability for LLM-based systems using Grafana, LangFuse, and LangSmith: performance, latency, token usage, and alerting. * Exposure to GenAI platforms and LLM-based applications * Experience in Life Science domain. ## Description RCH Solutions is seeking a Principal Cloud Platform Engineer with deep expertise in Kubernetes-based infrastructure to join our Cloud Engineering team. This role is ideal for individuals who take pride in designing, operating and evolving large-scale multi-tenant AI Platforms enabling real-world data and AI applications in the life sciences domain. This role is focused on platform-level engineering. You will own the reliability, scalability, and operational excellence of shared infrastructure supporting RAG-based AI workloads, with a strong emphasis on Kubernetes cluster operations and vector database systems. You'll collaborate closely with Data Engineers and AI Engineers and support them by providing a cloud-hosted scalable multi-tenant infrastructure platform., * Platform & Infrastructure Engineering * + Design, operate, and continuously improve production-grade K8s clusters at the platform level. + Lead complex cluster lifecycle management, including: o Version upgrades and dependency coordination o Failure recovery and incident resolution o Non-trivial maintenance and system evolution o Build and maintain highly reliable, scalable, multi-tenant infrastructure. + Build and maintain end-to-end observability for LLM-based systems using Grafana, LangFuse, and LangSmith - covering performance, latency, token usage, and alerting. * Multi-Tenant Platform Architecture * Architect and operate shared infrastructure across multiple teams and use cases. * Implement and enforce: + RBAC and access control models + Tenant isolation and security boundaries + Resource management and fairness at scale * Ensure platform stability under diverse and competing workloads * Vector Retrieval & AI Infrastructure * + Operate and optimize vector database systems (Weaviate preferred) in production environments. + Support and scale Retrieval-Augmented Generation (RAG) systems + Drive improvements in: o Query performance and latency o Cluster tuning and resource efficiency o Operational stability of retrieval pipelines * Production Ownership & Reliability * + Take technical ownership of production systems over time + Build and maintain strong practices in: o Observability (metrics, logs, tracing) o Incident response and root cause analysis o Long-term system health and resilience + Proactively identify and resolve reliability risks * Cross-Functional Collaboration * + Work closely with backend and GenAI engineers to ensure seamless integration with the platform. + Contribute to a balanced team structure, with a strong infrastructure core and targeted application-layer support. ## Related Videos - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)