> Markdown version of [/jobs/ext/2718450-ai-platform-agentic-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2718450-ai-platform-agentic-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Platform & Agentic Infrastructure Engineer - **Company:** AIG INTERNAL AUDIT - **Location:** San Jose, United States - **Experience:** Expert - **Salary:** $178,000.0 - $321,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, JIRA, Audit Trail, Automation of Tests, Computer-Aided Audit Tools, Cloud Computing, Continuous Integration, Information Engineering, Data Infrastructure, Data Structures, Disaster Recovery, Apache Hive, Identity and Access Management, Python (Programming Language), Key Management, Machine Learning, Node.Js, OAuth, Blockchain, Standard Sql, Data Streaming, TypeScript, Web Applications, Data Logging, Data Processing, System Availability, Large Language Models, Multi-Agent Systems, Backend, Real Time Data, Build Tools, Data Management, Gsuite, Terraform - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-ai-platform-agentic-infrastructure-engineer-okx-8792066 ## About the Role * 7+ years building and operating resilient backend or platform systems in production, including on-call ownership over time. * Proven brownfield migrations: you have taken a founder-built or prototype system to production grade while it stayed in daily use. * Strong engineering fundamentals: data structures and algorithms, fluent Python, strong SQL and data modeling, plus one additional systems language (TypeScript/Node or Go), with the full-stack range to connect the infrastructure yourself. * Agentic runtime and harness engineering, proven by building: the field is too young to demand years of it, so we weigh real systems shipped over tenure. You have genuinely built model-agnostic agent orchestration, model routing and evaluation across providers, agent SDKs (Claude Agent SDK or equivalent), MCP servers, skill- and hook-based agent tooling, and the evaluation and red-team harnesses that grade agent behavior, with responsible-AI controls (hallucination, bias, drift). * Data engineering with provenance and lineage, and security and data-protection engineering by default (encryption, identity and access management, secrets, retention). You build systems that could withstand external audit, by design. * Cloud (AWS or GCP) plus resilience engineering: infrastructure-as-code (Terraform), CI/CD, HA, DR, SLOs, and observability, backed by automated testing and documentation. * You ship inside locked-down enterprise environments (TLS-intercepting proxies, endpoint detection and response tooling, restricted installs, security guardrails, OAuth admin consent) without treating security as someone else's problem. * Partnership: you translate audit needs into systems, explain technical risk to non-engineers, and drive adoption. Nice to Haves * Multi-agent orchestration frameworks, MCP servers, and tooling and plugin development across the Anthropic, OpenAI, and Google model ecosystems. * Model-risk or AI-governance program experience; frameworks such as the NIST AI Risk Management Framework. * Continuous-auditing or continuous-controls-monitoring platforms; streaming and real-time data at scale; statistical anomaly detection and applied machine learning beyond LLMs; vector stores and retrieval; LLM cost engineering. * Experience passing external audit, SOC 2, or SOX (helpful, not required). * Crypto and blockchain literacy; regulated financial-services, fintech, or crypto experience. ## Description Cloud on AWS or GCP. Claude as the starting point in a deliberately multi-model architecture (Anthropic, OpenAI, Google), using the right model for each job, including the plugins and integrations you build. Google Workspace for reporting and evidence. JIRA for the audit team's workflow. Lark for team communications, alert bots, and the corporate wiki. You inherit a working Google-native prototype estate (Apps Script web apps, Drive-synced automation, locally scheduled jobs, Claude Code agent tooling) and evolve it without breaking daily use. Infrastructure-as-code, CI/CD, and observability throughout. Treat this stack as the starting point, not a constraint: you build production-grade systems end to end on what exists today, and you are expected to propose, prove, and adopt better components as demand and capabilities evolve. Production-grade here means service levels sized for an internal assurance platform: board-cycle windows are sacred, recovery is measured in hours, and this is not a 24/7 pager culture., * Hive Mind runs in the cloud with HA, DR, SLOs, and audit logging that passes an internal controls review. * A governed data foundation with provenance and lineage is live across multiple audit domains, integrated with OKX's group data infrastructure where it exists. * An agentic runtime and harness with evaluation and red-teaming gates what reaches production. * The hiring manager is out of the operational loop: no scheduled job runs on a personal machine, every system has a runbook and a non-founder owner. Decommissioning the founder's laptop as infrastructure is a literal milestone. When these goals compete, the priority order is: keep the estate alive, then the cloud migration with observability, then the harness gating production, then the data foundation., * Inherit, operate, and progressively migrate the working prototype estate (Google Workspace-native automation across Apps Script, Drive, and the Docs/Sheets/Slides APIs; locally scheduled jobs; and Claude Code agent tooling) to the target platform without interrupting daily and board-cycle workflows. Working software wins arguments; migrate by strangling, not rewriting. Reuse and integrate with OKX's prevailing and evolving AI capabilities and data infrastructure, including enterprise-approved models and gateways, Model Context Protocol (MCP) servers and connectors, security tooling, and group data platforms, before building parallel capability. * Re-architect Hive Mind into a resilient AWS or GCP platform with high availability (HA), disaster recovery (DR), defined service-level objectives (SLOs), and full observability, and own it in production. * Build the agentic runtime and harness: orchestration, multi-model routing that sends each task to the right model, tool, MCP, and plugin integration across model providers, and the evaluation, red-team, and regression harnesses that grade agents before and after production, with evaluation gates, judge calibration, and cost controls that hold for each model, including prompt-injection and data-exfiltration threat modeling for agents that read untrusted content. * Build the Responsible-AI and model-governance layer: hallucination, bias, and drift controls, output validation, guardrails, and complete logging. Extend the existing governance design (human-calibrated judge gates, golden-set regression, provenance registry) rather than replacing it. New model providers enter through the same governance and approved-tooling review, not around it. * Build data infrastructure with provenance and lineage: immutable audit trails, versioned evidence, reproducible pipelines, and traceability from source to report. * Engineer data protection: encryption, key and secrets management, least-privilege access, sensitive-data handling, residency, and defensible retention. * Stand up the computer-assisted audit technique (CAAT) and continuous-monitoring data foundation: analytics over full populations, with exceptions streamed in real time. Auditors and the hiring manager define the audit logic; you make it run at production grade. * Own the cloud foundation: infrastructure-as-code, CI/CD, identity, networking, observability, and cost controls, and make the platform examinable. ## Related Videos - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Keeping applications secure by evolving OAuth 2.0 and OpenID Connect](https://www.wearedevelopers.com/videos/100152-keeping-applications-secure-by-evolving-oauth-2-0-and-openid-connect) - [The Private AI Platform: Why Agentic Apps Need a Private Application Platform](https://www.wearedevelopers.com/videos/100162-the-private-ai-platform-why-agentic-apps-need-a-private-application-platform) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) - [AI Won't Fix Your Engineering Culture](https://www.wearedevelopers.com/videos/100266-ai-won-t-fix-your-engineering-culture) - [Delay the AI Overlords: How OAuth and OpenFGA Can Keep Your AI Agents from Going Rogue](https://www.wearedevelopers.com/videos/1637-delay-the-ai-overlords-how-oauth-and-openfga-can-keep-your-ai-agents-from-going-rogue) ## Related Articles - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)