> Markdown version of [/jobs/ext/1495314-senior-cloud-infrastructure-engineer-ai-platform](https://www.wearedevelopers.com/jobs/ext/1495314-senior-cloud-infrastructure-engineer-ai-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Cloud Infrastructure Engineer, AI Platform - **Company:** Procore - **Location:** Austin, TX, United States - **Experience:** Expert - **Salary:** $140,960.0 - $193,820.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Software as a Service, Cloud Computing, Data Retrieval, Machine Learning, Performance Tuning, Prometheus, Software Engineering, Datadog, Data Logging, Google Cloud, Load Balancing, Data Ingestion, Large Language Models, Grafana, Parallel Computation, AI Platforms, Low Latency, Optimization Algorithms, Terraform - **Published:** July 30, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/87869463/1 ## About the Role * 5+ years of hands-on experience with cloud platforms * Hands-on experience with cloud platforms, with strong expertise in Google Cloud Platform (GCP) and Infrastructure as Code (e.g., Terraform) * A demonstrated history of performance analysis, tuning, and infrastructure cost optimization, with the ability to speak about trade-offs and quantified impact * Experience building or working on multi-tenant SaaS platforms * Experience setting up end-to-end observability, including logging, metrics, and alerting using tools like Prometheus, Grafana, Datadog, or GCP Operations suite Preferred Qualifications * A fundamental understanding of the challenges in training and serving large machine learning models (e.g., memory constraints, computational complexity) * Strong understanding of VectorDBs, LLMs, and Agentic Observability tools (Datagrid uses Milvus for our VectorDB and Arize for agent tracing) * Experience with the Gemini API, specifically managing LLM quotas and load balancing * Hands-on experience with LLM serving frameworks and optimization techniques (quantization, tensor parallelism, FlashAttention) * Experience designing multi-tenant SaaS architectures and implementing resource quotas and cost allocation * High-level knowledge of agentic systems and best practices ## Description We are hiring a Senior Cloud Infrastructure Engineer to build the core infrastructure that powers the next generation of AI agents. Our agents must ingest, enrich, and vectorize massive, multi-modal datasets from over 100 customer sources. The core challenge is twofold: how do we do this in a way that is radically cost-efficient, while still allowing agents to deliver thorough responses in seconds?, This position reports to a Senior Manager, Software Engineering and will be 2 days per week hybrid role in our Austin office. We're looking for someone to join us immediately. What You'll Do * Build & Optimize for Scale & Cost: Implement and optimize our highly scalable, multi-tenant data ingestion and vectorization pipeline, with a relentless focus on improving cost and performance. * Implement Robust Monitoring: Create and maintain dashboards, alerts, and logging to ensure system health, identify performance bottlenecks, and provide immediate visibility into production issues. * Contribute to Millisecond Latency: Be a key contributor in performance tuning across the stack. You will help identify and eliminate bottlenecks in data retrieval, model inference, and agent response times to ensure a snappy, real-time user experience. * Engineer Multi-Tenant Architectures: Design and implement scalable, secure, and cost-effective multi-tenant infrastructures for our SaaS-based AI products, ensuring strict tenant isolation and fair resource allocation. ## Related Videos - [The OpenTelemetry mistakes I keep seeing (and how to stop making them)](https://www.wearedevelopers.com/videos/100158-the-opentelemetry-mistakes-i-keep-seeing-and-how-to-stop-making-them) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)