AI Foundation Model Engineer
Role details
Job location
Tech stack
Job description
The AI Foundation Model Engineer will design, develop, deploy, and optimize enterprise-grade Generative AI applications powered by Foundation Models, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), Agentic AI, and modern AI engineering frameworks. This role is responsible for transforming AI concepts into secure, scalable, observable, resilient, and production-ready enterprise solutions on the organization's AI Research & Innovation Platform (AIRP), which is currently hosted on AWS while following a cloud-agnostic architecture.
The organization is building a reusable enterprise AI platform that supports multiple business domains, and this role plays a critical part in engineering production AI capabilities. The engineer will develop AI solutions supporting business use cases such as KYC, credit underwriting, Banker 360, Customer 360, pitch book generation, deal intelligence, financial crime detection, sanctions screening, enterprise knowledge management, and intelligent workflow automation. Strong hands-on experience with AWS AI services, cloud-native architectures, Terraform, Infrastructure-as-Code (IaC), Kubernetes, Docker, DevOps, CI/CD, MLOps, and LLMOps is essential for delivering secure, scalable, and maintainable AI applications.
The successful candidate will own the end-to-end lifecycle of LLM-powered applications, RAG pipelines, AI services, model-serving platforms, APIs, observability, evaluation frameworks, deployment automation, monitoring, rollback strategies, and continuous optimization. Working closely with AI Researchers, Platform Engineers, Cloud Engineering, Product, Security, Risk, Compliance, and DevOps teams, the AI Foundation Model Engineer will build reusable AI services that accelerate enterprise AI adoption while maintaining security, governance, Responsible AI, and operational excellence., * Design, develop, and deploy LLM-powered enterprise applications, including knowledge assistants, document intelligence solutions, conversational AI, workflow agents, AI copilots, summarization platforms, enterprise search, decision-support systems, and intelligent automation solutions.
- Build and optimize Retrieval-Augmented Generation (RAG) pipelines using embeddings, semantic search, vector databases, document chunking strategies, reranking, retrieval optimization, response grounding, citation frameworks, and context management.
- Develop scalable AI services integrating Foundation Models, LLM APIs, model gateways, AI orchestration frameworks, AWS AI services, enterprise authentication, data services, and cloud-native platform components.
- Collaborate with Cloud Engineering and DevOps teams to implement Terraform Infrastructure-as-Code (IaC), reusable cloud modules, Kubernetes deployments, Docker containers, CI/CD pipelines, environment promotion, release management, deployment automation, rollback procedures, and production operations.
- Adapt and optimize foundation models using techniques such as LoRA, PEFT, instruction tuning, fine-tuning, transfer learning, model distillation, quantization, prompt engineering, domain adaptation, and inference optimization.
- Optimize production AI workloads for latency, throughput, scalability, token efficiency, inference cost, GPU utilization, reliability, resiliency, and user experience.
- Implement comprehensive LLMOps and MLOps practices, including model evaluation, prompt evaluation, retrieval quality assessment, monitoring, observability, automated testing, versioning, deployment governance, rollback strategies, and continuous improvement.
- Build enterprise observability capabilities capturing prompt logs, retrieval performance, hallucination indicators, groundedness, latency, token consumption, model drift, feedback loops, operational metrics, service health, and cost telemetry.
- Embed AI security, Responsible AI, privacy, model risk management, secure data handling, access controls, compliance, and governance into AI application architecture, development, deployment, and production operations.
- Develop and maintain production documentation, API specifications, technical design documents, deployment guides, runbooks, release notes, testing evidence, operational procedures, and audit-ready implementation artifacts.
- Collaborate with AI Researchers, Product Managers, Platform Engineers, Cloud Architects, Security, Risk, Compliance, and Business teams to deliver enterprise-ready AI capabilities that align with organizational architecture standards and business objectives.
Requirements
- 7+ years of experience in AI/ML Engineering, Software Engineering, Platform Engineering, Applied Machine Learning, or Generative AI Engineering.
- Strong hands-on experience with Large Language Models (LLMs), Foundation Models, Transformers, Retrieval-Augmented Generation (RAG), embeddings, semantic search, vector databases, Agentic AI, and enterprise Generative AI application development.
- Advanced Python development skills with frameworks such as PyTorch, TensorFlow, Hugging Face Transformers, LangChain, LlamaIndex, Semantic Kernel, or equivalent AI engineering frameworks.
- Experience building and deploying production AI services using REST APIs, microservices, Docker, Kubernetes, cloud-native architectures, CI/CD pipelines, MLOps, LLMOps, monitoring platforms, and enterprise deployment pipelines.
- Strong hands-on experience with AWS AI and Cloud Services, including Amazon Bedrock, SageMaker, OpenSearch, Kendra, Lambda, ECS/EKS, API Gateway, IAM, CloudWatch, or equivalent hyperscaler AI services.
- Practical experience with Terraform, Infrastructure-as-Code (IaC), DevOps, CI/CD, release management, model evaluation, inference optimization, secure data handling, and enterprise deployment practices.
- Strong understanding of AI security, Responsible AI, model governance, observability, monitoring, production operations, scalability, and performance optimization.
Preferred Experience
- Experience within Banking, Financial Services, FinTech, Insurance, Risk Management, Financial Crime, AML, Compliance, Enterprise Knowledge Management, or regulated enterprise environments.
- Experience with AWS Bedrock, Amazon SageMaker, OpenSearch, Kendra, Lambda, ECS/EKS, Azure OpenAI, Vertex AI, Databricks, MLflow, Kubeflow, vLLM, Triton Inference Server, LangGraph, model gateways, vector databases, and enterprise AI platforms.
- Experience implementing cloud-agnostic AI architectures, reusable Terraform modules, platform engineering standards, AI governance controls, model risk management, audit requirements, AI cost governance, and private or open-source LLM deployments.
- Familiarity with Responsible AI, AI Governance, Model Risk Management, Security, Privacy, Compliance, Enterprise AI Architecture, and Production AI Operations.