> Markdown version of [/jobs/ext/1964716-ai-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/1964716-ai-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Infrastructure Engineer - **Company:** Kirkland and Ellis - **Location:** Houston, TX, United States - **Experience:** Expert - **Salary:** $133,000.0 - $166,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Application Performance Management, Microsoft Azure, Cloud Computing, Cloud Computing Security, Cloud Engineering, Information Systems, DevOps, Identity and Access Management, Python (Programming Language), Key Management, Network Security, Log Analysis, Performance Tuning, Windows PowerShell, Reliability Engineering, Azure Active Directory, Search Technologies, AI Infrastructure, Policy as Code, Cloud Monitoring, Grafana, AI Platforms, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Bicep, Azure AKS, Terraform - **Published:** August 7, 2026 - **Apply:** https://staffjobsus.kirkland.com/jobs/17823936-ai-infrastructure-senior-engineer-i ## About the Role Are you passionate about building and operating secure, scalable AI platforms that power real-world innovation?, * Education & Certifications: Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, along with relevant Microsoft Azure certifications (e.g., Azure Administrator Associate or similar). * Experience: 4-6 years in cloud infrastructure, platform engineering, or site reliability roles, including at least 3 years working with Microsoft Azure in production environments. * Cloud & Platform Expertise: Hands-on experience managing Azure services, infrastructure-as-code, Kubernetes (AKS), and enterprise cloud environments with a focus on reliability and scalability. * Automation & Scripting: Strong scripting capabilities in Python and PowerShell to support automation, tooling, and operational efficiency. * Networking & Identity: Solid understanding of enterprise networking, access management, and secure cloud architecture. * Monitoring & DevOps Practices: Experience with observability tools, CI/CD pipelines, and modern deployment practices, including GitOps and policy-as-code. * AI Platform Exposure: Familiarity with Azure-based AI services and platform-level considerations such as capacity management, content controls, and governance. * Collaboration & Communication: Ability to work cross-functionally with engineering, infrastructure, and security teams while translating complex technical concepts into practical outcomes. * Operational Excellence Mindset: Experience in incident response, on-call support, and continuous improvement within production environments. If you're excited to help build and operate cutting-edge AI infrastructure, collaborate with high-performing teams, and drive meaningful platform innovation in this AI Infrastructure Engineer role, we'd love to hear from you! ## Description As an AI Infrastructure Engineer, you'll play a key role in shaping and supporting a modern AI platform within the Information Technology team, specifically the AI Infrastructure group. You'll help ensure the reliability, performance, and scalability of shared AI services, enabling engineering and automation teams to deliver impactful solutions efficiently. Working closely with Cloud Engineering and other technology partners, you'll contribute to a high-performing, enterprise-grade environment supporting advanced AI capabilities across the firm. * Platform Build & Configuration: Implement Infrastructure-as-Code (IaC) using tools like Terraform or Bicep, maintain standardized Azure environments, and manage core services such as Azure OpenAI, Azure AI Foundry, Azure Kubernetes Service (AKS), and Azure AI Search. * Cloud Environment Management: Configure and maintain networking (private endpoints, hub-and-spoke architecture, network security groups), identity and access (Microsoft Entra ID, Managed Identity), and secrets (Azure Key Vault) aligned to security best practices. * Operational Monitoring & Reliability: Monitor platform health using Azure Monitor, Application Insights, and Log Analytics; respond to alerts, troubleshoot incidents, and support on-call rotations to maintain service continuity. * Capacity & Performance Optimization: Manage quotas, scaling, and performance for AI services while supporting capacity planning aligned to business growth. * Change & Release Enablement: Execute platform updates, maintenance, and upgrades using established change management processes while supporting onboarding of new AI workloads. * Security & Compliance: Enforce security controls, governance policies, and Responsible AI practices; remediate vulnerabilities and support audit and compliance reporting. * Cost Management & Optimization: Drive visibility into platform usage and costs, ensuring proper tagging, rightsizing, and efficient resource allocation. * Documentation & Continuous Improvement: Create and maintain clear documentation, runbooks, and operational procedures while contributing to platform enhancements and roadmap evolution. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Azure AI Foundry for Developers: Open Tools, Scalable Agents, Real Impact](https://www.wearedevelopers.com/videos/1541-azure-ai-foundry-for-developers-open-tools-scalable-agents-real-impact) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts)