Infrastructure Architect

Ark Infotech Spectrum
Dallas, TX, United States
7 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Computing Platforms Microsoft Azure Big Data Cloud Computing Cloud Engineering Cloud Foundry Information Systems Disaster Recovery Distributed Systems
+19 more
OpenShift Performance Tuning Azure Machine Learning Software Engineering Systems Architecture AI Infrastructure Google Cloud Cloud Platform System System Availability Prompt Engineering Multi-Cloud Generative AI AI Platforms Kubernetes Information Technology Data Management Machine Learning Operations Hardware Infrastructure Virtual Agents

Job description

Architecture Leadership Lead architecture and technical direction for enterprise AI platform capabilities, including: Enterprise Generative AI platforms Agentic AI platforms and agent runtime environments Model serving and inference infrastructure Prompt engineering, evaluation, and testing frameworks AI governance, risk management, and guardrails AI observability, monitoring, and operations Multi-cloud AI platform strategy and architecture

Platform Architecture & Design Design, evaluate, and guide architecture across AI and cloud technologies such as: Red Hat OpenShift AI (RHOAI) Google Cloud Vertex AI Gemini models Azure AI Foundry Amazon Bedrock Anthropic Claude OpenAI services and models Cloud-native platform scalability and resiliency solutions

Strategic Initiatives Lead and collaborate on initiatives involving: Capacity planning and performance optimization GPU infrastructure and platform strategy Large-scale NVIDIA-based AI infrastructure architectures Multi-region and multi-cloud resiliency Disaster recovery planning and cloud DR strategies Active-active platform architectures High availability and business continuity solutions

Requirements

The successful candidate will demonstrate deep expertise in cloud-native architectures, AI/ML platforms, distributed systems, and platform engineering, along with the ability to influence technical roadmaps and lead cross-functional initiatives., Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related field; or equivalent combination of education and relevant experience. Minimum of 10 years of experience in software engineering, infrastructure engineering, systems architecture, or related technical disciplines. Minimum of 5 years of experience designing and implementing large-scale cloud-native platforms. Experience designing, deploying, and supporting highly available, mission-critical production systems. Experience with: Kubernetes and/or OpenShift Distributed systems architectures Public cloud platforms and cloud-native architectures AI/ML platforms and infrastructure API and integration platforms Data platforms and data-intensive applications Knowledge of: Generative AI systems and architectures Retrieval-Augmented Generation (RAG) Agentic AI frameworks and platforms Model serving and inference architectures MLOps practices and tooling Demonstrated ability to lead technical initiatives across multiple teams and stakeholders. Ability to work from or relocate to one of the following approved locations, Experience with one or more major cloud providers, including AWS, Azure, or Google Cloud Platform. Experience architecting GPU-accelerated AI infrastructure. Experience implementing AI governance, security, risk management, and compliance controls. Experience building or supporting multi-region, highly resilient enterprise platforms. Relevant industry certifications in cloud, AI/ML, Kubernetes, or architecture disciplines.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

4:47 min

Exploring OpenShift and Red Hat Developer Sandbox resources

Markus Eisele Markus Eisele · World Congress 2024

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

Videos

See all

Related articles

See all