LLMOps / AgentOps Engineer

Sierra Business Solution LLC
Santa Clara, CA, United States
27 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Cloud Computing DevOps Python (Programming Language) Performance Tuning Reliability Engineering Management of Software Versions Datadog Google Cloud Multi-Agent Systems Mttr Machine Learning Operations

Job description

Senior subject-matter expert for the operations, observability, and lifecycle management of AI agents in production ( AgentOps LLMOps). Owns the frameworks and practices to safely deploy, monitor, evaluate, and continuously improve live agents ensuring reliability, safety, cost-efficiency, and business-KPI performance across the Intel Agent Factory., Define and operate the AgenticOps framework: agent registry, versioning, guarded rollout, and rollback for production agents.

Establish continuous evaluation and monitoring: quality, autonomy, safety (guardrails, Model Armor), latency, cost, and reuse metrics.

Implement observability and tracing for multi-agent systems (Agent Engine Observability, Cloud MonitoringLoggingTrace).

Own the 5-gate validation-to-production process and post-release escape management for delivered agents.

Design human-in-the-loop (HITL) supervision, feedback loops, and automated pre-production simulations for safe rollout.

Track and report agent business KPIs (CSAT, TAT, MTTR, cost savings) via AgentScore Agent 360 dashboards.

Drive cost governance for agent runtimes: model tiering, context caching, batchflex inference, budget caps and alerts.

Collaborate with DevOps SME (deploy) and AI & Data SME (grounding) to close the build-deploy-operate-improve loop advise Intel on AgenticOps ownership transfer.

Requirements

Strong LLMOps MLOps AgentOps experience operating GenAI or agentic systems in production.

Hands-on with Google Cloud agent runtimes: Vertex AI, Agent Engine, and observability tooling.

Agent evaluation and safety: eval frameworks, guardrails, Model Armor, HITL, promptrobustness testing.

Monitoring, tracing, and reliability engineering (SRE) for AI workloads.

Cost governance and performance tuning for LLMagent workloads.

Proficiency in Python strong grasp of agent lifecycle and governance.

Preferred (Good-to-Have) Skills

Experience with ADK, A2A, MCP, and multi-agent orchestration in production.

BigQueryLooker for agent analytics and KPI dashboards.

Responsible-AI, model governance, and auditcompliance frameworks.

Prior enterprise-scale AI platform operations experience.

Experience & Certifications

9 12+ years in MLAI platform operations, SRE, or LLMOps with production agenticGenAI exposure (Tier 5 6)., Google Cloud Professional (MLDevOps) certification preferred.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Loading talks and stories from around this role…