AI Infrastructure Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+26 more
Job description
The AI Infrastructure Engineer is a hands-on platform and infrastructure engineer responsible for building, deploying, and operating the cloud foundations that enable enterprise AI and agentic solutions across the Truist Agentic Enterprise (TAE).
This role focuses on infrastructure, automation, deployment, reliability, observability, and operational excellence for AI-enabled workloads. The engineer partners closely with software engineers, data teams, architects, security teams, and platform teams to deliver scalable, secure, and production-ready environments for agentic systems and AI applications.
The ideal candidate combines strong cloud engineering expertise with modern software engineering skills, including Python, Infrastructure as Code (IaC), CI/CD automation, containerization, and cloud-native operations.
The role requires a strong understanding of platform architecture, deployment patterns, monitoring, resiliency, and operational support for enterprise-scale AI solutions.
Working as an individual contributor, the engineer executes independently on assigned initiatives while contributing to engineering excellence, platform stability, and delivery efficiency across the organization., Following is a summary of the essential functions for this job. Other duties may be performed, both major and minor, which are not mentioned below. Specific activities may change from time to time. Build, deploy, and support infrastructure platforms that enable the development and operation of AI agents, AI applications, and intelligent automation solutions.
- Develop and maintain Infrastructure as Code (IaC) using Terraform and related automation technologies to provision and manage cloud resources.
- Design and support CI/CD pipelines that enable reliable build, test, deployment, and release processes for AI and agentic workloads.
- Implement containerized deployment solutions using modern container and orchestration technologies.
- Configure and manage cloud-native services, networking, security controls, storage, compute resources, and supporting platform capabilities.
- Partner with development teams to operationalize AI workloads, APIs, services, and agentic solutions in production environments.
- Implement observability capabilities including monitoring, logging, tracing, alerting, telemetry, and performance analytics.
- Support platform reliability through automation, resiliency engineering, incident response, root-cause analysis, and operational improvements.
- Contribute to cloud architecture decisions, deployment standards, environment design, and infrastructure best practices.
- Develop automation and tooling using Python and related technologies to improve operational efficiency, platform consistency, and deployment speed.
- Implement security, compliance, access control, and governance requirements within cloud and deployment platforms.
- Create and maintain technical documentation, deployment procedures, operational runbooks, and platform support materials.
- Collaborate with engineering, architecture, security, product, and operations teams to deliver reliable and scalable AI platform capabilities.
- Contribute reusable infrastructure modules, automation patterns, and platform standards that improve engineering productivity and operational stability.
Requirements
The requirements listed below are representative of the knowledge, skill and/or ability required. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.
- Bachelor’s degree in Computer Science, Engineering, Information Systems, a related field, or equivalent education, training, and work-related experience.
- Minimum of 5 years of professional experience in infrastructure engineering.
- Strong knowledge of enterprise infrastructure technologies including cloud, network, database, storage, platform, computing, and middleware., * 5+ years of professional experience in cloud engineering, platform engineering, infrastructure engineering, DevOps, Site Reliability Engineering (SRE), or related disciplines.
- Strong experience with Infrastructure as Code (IaC), including Terraform.
- Strong programming and automation experience using Python.
- Experience designing and supporting CI/CD pipelines and deployment automation.
- Experience working with container technologies and orchestration platforms.
- Experience supporting cloud-native environments in Microsoft Azure.
- Working knowledge of cloud architecture, networking, security, storage, compute, and operational best practices.
- Experience implementing monitoring, logging, alerting, and observability solutions in production environments.
- Experience supporting highly available and reliable production systems.
- Strong troubleshooting, analytical, and problem-solving skills.
- Strong written and verbal communication skills with the ability to work effectively across technical and business teams.
- Experience supporting AI, machine learning, generative AI, or agentic workloads in cloud environments.
- Experience with Azure AI services, enterprise AI platforms, and AI deployment architectures.
- Experience with Amazon Web Services (AWS) in addition to Microsoft Azure.
- Experience deploying and operating containerized workloads using Kubernetes or related orchestration platforms.
- Experience implementing platform observability, distributed tracing, and operational telemetry for complex applications.
- Experience with DevSecOps practices, infrastructure security, and automated compliance controls.
- Familiarity with large language models (LLMs), AI development workflows, model deployment patterns, and retrieval-based architectures.
- Experience operating in highly regulated environments such as financial services, cybersecurity, healthcare, or other governance-driven industries.
- Experience contributing to infrastructure standards, operational best practices, and platform engineering initiatives.
Benefits & conditions
General Description of Available Benefits for Eligible Employees of Truist Financial Corporation: All regular teammates (not temporary or contingent workers) working 20 hours or more per week are eligible for benefits, though eligibility for specific benefits may be determined by the division of Truist offering the position. Truist offers medical, dental, vision, life insurance, disability, accidental death and dismemberment, tax-preferred savings accounts, and a 401k plan to teammates. Teammates also receive no less than 10 days of vacation (prorated based on date of hire and by full-time or part-time status) during their first year of employment, along with 10 sick days (also prorated), and paid holidays. For more details on Truist’s generous benefit plans, please visit our Benefits site. Depending on the position and division, this job may also be eligible for Truist’s defined benefit pension plan, restricted stock units, and/or a deferred compensation plan. As
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Industries Outside of AI Are Hiring The Most AI Experts?
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
Navigating the AI Shift
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production