Senior Cloud Infrastructure Engineer, AI Platform

Procore
West, United States of America
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 194K

Job location

West, United States of America

Tech stack

API
Artificial Intelligence
Software as a Service
Cloud Computing
Data Retrieval
Machine Learning
Performance Tuning
Prometheus
Software Engineering
Datadog
Data Logging
Google Cloud Platform
Load Balancing
Data Ingestion
Large Language Models
Grafana
Parallel Computation
AI Platforms
Low Latency
Optimization Algorithms
Terraform

Job description

This position reports to a Senior Manager, Software Engineering and will be 2 days per week hybrid role in our Austin office. We're looking for someone to join us immediately. What You'll Do

  • Build & Optimize for Scale & Cost: Implement and optimize our highly scalable, multi-tenant data ingestion and vectorization pipeline, with a relentless focus on improving cost and performance.
  • Implement Robust Monitoring: Create and maintain dashboards, alerts, and logging to ensure system health, identify performance bottlenecks, and provide immediate visibility into production issues.
  • Contribute to Millisecond Latency: Be a key contributor in performance tuning across the stack. You will help identify and eliminate bottlenecks in data retrieval, model inference, and agent response times to ensure a snappy, real-time user experience.
  • Engineer Multi-Tenant Architectures: Design and implement scalable, secure, and cost-effective multi-tenant infrastructures for our SaaS-based AI products, ensuring strict tenant isolation and fair resource allocation.

Requirements

  • 5+ years of hands-on experience with cloud platforms
  • Hands-on experience with cloud platforms, with strong expertise in Google Cloud Platform (GCP) and Infrastructure as Code (e.g., Terraform)
  • A demonstrated history of performance analysis, tuning, and infrastructure cost optimization, with the ability to speak about trade-offs and quantified impact
  • Experience building or working on multi-tenant SaaS platforms
  • Experience setting up end-to-end observability, including logging, metrics, and alerting using tools like Prometheus, Grafana, Datadog, or GCP Operations suite

Preferred Qualifications

  • A fundamental understanding of the challenges in training and serving large machine learning models (e.g., memory constraints, computational complexity)
  • Strong understanding of VectorDBs, LLMs, and Agentic Observability tools (Datagrid uses Milvus for our VectorDB and Arize for agent tracing)
  • Experience with the Gemini API, specifically managing LLM quotas and load balancing
  • Hands-on experience with LLM serving frameworks and optimization techniques (quantization, tensor parallelism, FlashAttention)
  • Experience designing multi-tenant SaaS architectures and implementing resource quotas and cost allocation
  • High-level knowledge of agentic systems and best practices

Benefits & conditions

140,960.00 - 193,820.00 USD Annual

This role may also be eligible for Equity Compensation and/or Bonus Incentive Compensation. Procore is committed to offering competitive, fair, and commensurate compensation. Actual compensation will be based on a candidate's job-related skills, experience, education or training, and location. For Los Angeles County (unincorporated) Candidates:

Procore will consider for employment all qualified applicants, including those with arrest or conviction records, in accordance with the requirements of applicable federal, state, and local laws, including the City of Los Angeles' Fair Chance Initiative for Hiring Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act.

Apply for this position