TELECOMMUTE AI Architect - Remote

Lightning Minds Inc.
United States
5 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

A/B Testing Artificial Intelligence Amazon Web Services Automated Storage and Retrieval Systems Microsoft Azure Software as a Service Cloud Computing Security Computer Programming Databases Data Centers Disaster Recovery Distributed Systems
+37 more
Failover Identity and Access Management Python (Programming Language) Lua (Scripting Language) Routing Open Source Technology Ping (Networking Utility) Release Management Systems Integration AI Infrastructure Private Cloud Environment Datadog Data Logging Pulumi Flexi (Photoshop Plugin) Google Cloud Load Balancing System Availability Large Language Models Multi-Agent Systems IT Architecture Multi-Cloud Caching HybridCloud Rate Limiting Event Driven Architecture AI Platforms Infrastructure Automation Frameworks Low Latency Deployment Automation Machine Learning Operations Api Design Terraform Splunk Software Version Control Dynatrace Microservices

Job description

Architecture & Design

  • Design end-to-end architecture for agentic AI systems (multi-agent orchestration, tool-use frameworks, memory/state management, planning and reasoning loops) deployed across hybrid and multi-cloud environments.
  • Architect and implement an AI Gateway layer to unify access to multiple LLM/model providers (OpenAI, Anthropic, Google, open-source models, self-hosted models) with centralized routing, rate limiting, load balancing, caching, and failover.
  • Define reference architectures for hybrid cloud AI deployments, balancing workloads across on-prem, private cloud, and public cloud (AWS/Azure/Google Cloud Platform) based on data residency, latency, cost, and compliance requirements.
  • Establish patterns for model orchestration, agent-to-agent communication, and tool/function calling across distributed systems.
  • Design for interoperability across cloud-native AI services (Bedrock, Azure AI Foundry, Microsoft Co-pilot)
  • Integrating MCP Registry, governance of Agent Registry
  • Microsoft EntraID as identity provider for MCP authentication
  • Integrating AI gateway with AWS Valkey cache and other vector / RAG databases
  • Specific implementation experience of Kong AI gateway with plug-in based architecture, ability create plug-ins using Lua script or Python
  • Observability integration with Dynatrace for agents and AI gateway telemetry; Splunk integration for audit logs

AI Gateway & Governance

  • Own the strategy and implementation of the AI Gateway as the control plane for all AI/LLM traffic including authentication, authorization, usage metering, cost governance, PII/data redaction, prompt/response logging, and audit trails.
  • Implement guardrails for model governance: version control, A/B testing, fallback routing, semantic caching, and observability across multiple model providers.
  • Define policies for responsible AI, data privacy, and security across the gateway and agentic workflows, ensuring compliance with regulatory frameworks (GDPR, SOC 2, HIPAA, etc. as applicable).

Multi-Cloud & Infrastructure

  • Architect resilient, scalable, and cost-optimized infrastructure spanning multiple cloud providers and on-prem data centers.
  • Drive infrastructure-as-code practices (Terraform, Pulumi) for consistent multi-cloud provisioning of AI/ML infrastructure.
  • Evaluate and integrate vector databases, feature stores, and RAG pipelines across hybrid environments.
  • Ensure high availability, disaster recovery, and business continuity for mission-critical agentic AI applications.

Leadership & Collaboration

  • Partner with engineering, product, security, and data teams to align AI architecture with business objectives.
  • Provide technical leadership and mentorship to AI/ML engineers and platform teams.
  • Evaluate emerging tools, frameworks, and vendors in the agentic AI and AI infrastructure ecosystem (e.g., MCP Registry, Kong AI Gateway, Agent365).
  • Create architecture documentation, design standards, and best practices to guide organization-wide AI adoption.
  • Act as a trusted advisor to leadership on AI platform strategy, build-vs-buy decisions, and technology roadmaps.
  • Setup Operating model to operate and govern AI gateway usage enterprise-wide
  • Scalability, hosting topology, define centralized vs federated model
  • Deployment architecture, change management, release management and incident management
  • Alignment with United AI governance Framework for agentic systems
  • Integrate skills, tools, MCP servers - internal or SaaS vendor source

Requirements

  • 8+ years of experience in software/solutions architecture, with 3+ years focused on AI/ML systems.
  • Proven hands-on experience architecting agentic AI systems multi-agent frameworks, autonomous workflows, tool/function calling, and orchestration.
  • Strong expertise in hybrid and multi-cloud architecture (AWS, Azure, Google Cloud Platform, and on-premises/private cloud integration).
  • Demonstrated experience designing or implementing an AI Gateway / LLM Gateway(e.g., Kong, custom-built gateways) for managing multi-model access, routing, rate limiting, and observability.
  • Deep understanding of LLM ecosystems (OpenAI, Anthropic, open-source models) and model serving infrastructure.
  • Experience with RAG architectures, vector databases (e.g.Milvus, pgvector), and knowledge retrieval systems.
  • Solid grounding in API design, microservices, event-driven architecture, and distributed systems.
  • Strong knowledge of cloud security, IAM, networking, and compliance frameworks across multi-cloud environments (IDPs like EntraID and Ping.
  • Proficiency with infrastructure-as-code (Terraform/Harness, DeclarativeKong), CI/CD pipelines, and observability tooling (Splunk, DynaTrace).
  • Programming proficiency in Python (required).

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

1:51 min

Overview of the three Google Maps routing applications

Germán Álvarez · LIVE

1:34 min

Transitioning from traditional software development to artificial intelligence consulting

Patrick Schnell Patrick Schnell · Coffee With Developers

Videos

See all

Related articles

See all