> Markdown version of [/jobs/ext/2405241-ai-foundational-model-engineer](https://www.wearedevelopers.com/jobs/ext/2405241-ai-foundational-model-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Foundational Model Engineer - **Company:** VDart, Inc. - **Location:** Jersey City, NJ, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Engineering, Continuous Integration, Data as a Services, Data Security, Python (Programming Language), Machine Learning, Open Source Technology, Performance Tuning, Release Management, Cloud Services, Tensorflow, Runbook, Search Technologies, Software Deployment, Software Engineering, Pytorch, Retrieval-Augmented Generation, Transfer Learning, Delivery Pipeline, Large Language Models, Model Validation, Containerization, AI Platforms, Kubernetes, Low Latency, HuggingFace, Machine Learning Operations, Terraform, Serverless Computing, Databricks - **Published:** August 14, 2026 - **Apply:** https://www.careerjet.com/jobad/us18c5dddd44af0627bbba5f830850c957 ## About the Role * 7+ years in AI/ML engineering, platform engineering, software engineering, or applied machine learning. * Hands-on experience with LLMs, transformers, embeddings, RAG, semantic search, and GenAI application patterns. * Strong Python engineering skills with PyTorch, TensorFlow, Hugging Face, LangChain, LlamaIndex, Semantic Kernel, or equivalent frameworks. * Experience deploying production AI services using APIs, containers, Kubernetes, CI/CD, cloud-native services, and monitoring platforms. * Practical exposure to AWS AI/cloud services or comparable cloud-native AI deployment experience, with ability to ramp quickly on AWS-hosted AIRP patterns. * Working knowledge of Terraform/IaC, DevOps pipelines, release management, model evaluation, inference optimization, and secure data handling. ## Description * Design, build, deploy, and optimize enterprise-grade AI systems powered by foundation models, LLMs, retrieval-augmented generation, and agentic workflows. * The role converts AI concepts into secure, scalable, observable, and supportable production systems on the enterprise AI-ready platform (AIRP), which is currently AWS-hosted while following a cloud-agnostic architecture blueprint. * Hands-on AWS AI and cloud engineering is a major asset because AIRP currently runs on AWS. * Candidates should be comfortable working with Terraform/IaC and CI/CD teams to move AI services and infrastructure through controlled deployment pipelines. * Experience should map to business AI use cases such as KYC, credit underwriting, pitch book generation, Banker 360, Customer 360, deal library intelligence, financial crime quality, and sanctions screening. * Primary ownership * Production LLM applications, RAG pipelines, AI services, and model-serving integrations for AIRP. * End-to-end LLMOps/MLOps lifecycle from experimentation to deployment, monitoring, evaluation, rollback, and continuous improvement. * Reusable AI service components, APIs, prompts, retrieval logic, and observability patterns that can be federated across multiple business use cases. * Key responsibilities * Design and implement LLM-powered applications such as knowledge assistants, document intelligence solutions, workflow agents, summarization tools, and decision-support systems. * Build RAG pipelines using embeddings, chunking strategies, vector databases, semantic retrieval, reranking, response grounding, and citation patterns. * Integrate AI capabilities with AWS-hosted platform components, including model APIs, model gateways, data services, container platforms, and enterprise authentication patterns. * Collaborate with cloud engineering teams on Terraform modules, IaC templates, environment promotion, CI/CD pipelines, release controls, and rollback procedures. * Adapt and optimize models using LoRA, PEFT, instruction tuning, distillation, transfer learning, quantization, and domain adaptation techniques where appropriate. * Optimize inference workloads for latency, throughput, token efficiency, cost, reliability, and user experience. * Implement model and application observability, including prompt logs, retrieval quality, hallucination indicators, drift signals, feedback loops, cost telemetry, and service health. * Embed security, privacy, Responsible AI, and model risk controls into AI application design and delivery. * Create production documentation, runbooks, release notes, test evidence, and audit-ready implementation records., * Banking, risk, compliance, financial crime, operations, or enterprise technology background. * Experience with AWS Bedrock, SageMaker, OpenSearch, Kendra, Lambda, EKS/ECS, Azure OpenAI, Vertex AI, Databricks, vLLM, Triton, MLflow, Kubeflow, or model gateways. * Exposure to cloud-agnostic application patterns, reusable IaC modules, model risk, AI governance, audit controls, AI cost governance, and private or open-source LLM deployments. * Initial screening questions * Describe a production LLM or RAG system you built. What was your role and what changed after launch? * Which AWS AI or cloud services have you used for production AI delivery, and what design trade-offs did you make? * How have you worked with Terraform, IaC modules, or CI/CD pipelines to deploy AI services? * How did you evaluate groundedness, hallucination rate, retrieval quality, latency, and cost? * How did you secure sensitive data and prevent leakage in the AI pipeline? * What observability and rollback mechanisms did you implement?, Lead Machine Learning Engineer (MLOps, KServe + building Kubernetes Clusters, PyTorch, TensorFlow on AWS) As a Capital One Machine Learning Engineer (MLE), you'll be part of an Agi… + 1 day ago + ## Related Videos - [Developer Experience, Platform Engineering and AI powered Apps](https://www.wearedevelopers.com/videos/990-developer-experience-platform-engineering-and-ai-powered-apps) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Agentic AI - From Theory to Practice: Developing Multi-Agent AI Systems on Azure](https://www.wearedevelopers.com/videos/1532-agentic-ai-from-theory-to-practice-developing-multi-agent-ai-systems-on-azure) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)