AI Foundation Model Engineer

Niche IT Software Solutions LLC
Jersey City, United States
12 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
9 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Microsoft Azure Continuous Integration Data Security Monitoring of Systems Python (Programming Language) Machine Learning Open Source Technology Tensorflow Runbook
+15 more
Search Technologies Software Deployment Enterprise Software Applications Cloud Platform System Pytorch Large Language Models Model Validation Build Management AI Platforms Kubernetes HuggingFace Machine Learning Operations Terraform Docker Databricks

Job description

We are seeking a Senior AI Foundation Model Engineer to build and deploy secure, scalable, enterprise-grade AI solutions using LLMs, RAG and agentic workflows. The role involves developing production AI applications and reusable services for an AWS-hosted, cloud-agnostic AI platform., · Build LLM applications, RAG pipelines, knowledge assistants, document intelligence solutions and workflow agents.

· Develop embeddings, semantic search, reranking, grounding and citation capabilities.

· Deploy and manage AI services using APIs, Docker, Kubernetes, CI/CD and cloud-native infrastructure.

· Collaborate on Terraform/IaC, environment promotion, release controls and rollback procedures.

· Optimize models and inference for accuracy, latency, throughput, token usage, reliability and cost.

· Implement LLMOps/MLOps covering evaluation, monitoring, observability, feedback loops and continuous improvement.

· Ensure security, privacy, Responsible AI, governance and audit readiness.

· Maintain production documentation, runbooks and release records.

Requirements

· Strong hands-on experience with LLMs, transformers, GenAI, RAG, embeddings and vector databases.

· Advanced Python skills and experience with PyTorch, TensorFlow, Hugging Face, LangChain, LlamaIndex, Semantic Kernel or similar frameworks.

· Production deployment experience using APIs, containers, Kubernetes, CI/CD and monitoring tools.

· Practical AWS AI/cloud experience, preferably with Bedrock, SageMaker, OpenSearch, Lambda and EKS/ECS.

· Working knowledge of Terraform/IaC, MLOps/LLMOps, model evaluation, inference optimization and secure data handling.

Preferred Experience

· Banking, risk, compliance, financial crime or enterprise technology experience.

· Experience with Kendra, Azure OpenAI, Vertex AI, Databricks, vLLM, Triton, MLflow, Kubeflow or model gateways.

· Knowledge of LoRA, PEFT, instruction tuning, quantization, model governance and private/open-source LLM deployments.

Alternate Titles: LLM Engineer, GenAI Engineer, AI Platform Engineer, RAG Engineer, Applied ML Engineer or NLP Engineer.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

2:50 min

Introduction and the value of runbooks

Hila Fish · WWC 2023

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa · LIVE

Videos

See all

Related articles

See all