AI Foundation Model Engineer

Usg Inc.
Jersey City, NJ, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Engineering Continuous Integration Data Security Python (Programming Language) Machine Learning Open Source Technology Performance Tuning Tensorflow Search Technologies
+20 more
Software Deployment Software Engineering Data Logging Enterprise Software Applications Pytorch Transfer Learning Large Language Models Model Validation Software Application Programming Generative AI Containerization AI Platforms Kubernetes HuggingFace Machine Learning Operations Virtual Agents Restful APIs Serverless Computing Databricks Microservices

Job description

Seeking an experienced AI Foundation Model Engineer to design, build, deploy, and optimize enterprise-grade AI solutions powered by Large Language Models (LLMs), Generative AI, Retrieval-Augmented Generation (RAG), and agentic AI workflows. This role is responsible for developing scalable, secure, and production-ready AI applications while ensuring operational excellence, observability, governance, and compliance within enterprise environments., * Design and develop LLM-powered applications including knowledge assistants, document intelligence platforms, workflow agents, summarization tools, and decision-support systems.

  • Build Retrieval-Augmented Generation (RAG) pipelines using embeddings, semantic search, vector databases, chunking strategies, reranking, response grounding, and citation mechanisms.
  • Fine-tune and optimize foundation models using techniques such as LoRA, PEFT, instruction tuning, transfer learning, knowledge distillation, quantization, and domain adaptation.
  • Develop scalable APIs, microservices, model-serving infrastructure, and integration services across cloud, hybrid, and containerized environments.
  • Optimize inference workloads for latency, throughput, token efficiency, scalability, reliability, cost optimization, and user experience.
  • Implement observability solutions for AI applications including prompt logging, retrieval quality metrics, hallucination detection, model drift monitoring, service health, user feedback, and cost telemetry.
  • Embed security, privacy, Responsible AI, model governance, and enterprise risk controls throughout the AI application lifecycle.
  • Create production documentation, deployment guides, runbooks, release documentation, testing evidence, and audit-ready implementation artifacts.
  • Collaborate with AI Researchers, Platform Engineers, Security, Product, Architecture, and Business teams to deliver enterprise AI capabilities.

Requirements

The ideal candidate combines strong AI/ML engineering expertise with cloud-native software development and production deployment experience., * 7+ years of experience in AI/ML Engineering, Applied Machine Learning, Platform Engineering, Software Engineering, or related disciplines.

  • Hands-on experience developing applications using Large Language Models (LLMs), Transformers, embeddings, Retrieval-Augmented Generation (RAG), semantic search, and Generative AI architectures.
  • Strong Python development experience with frameworks such as PyTorch, TensorFlow, Hugging Face, LangChain, LlamaIndex, Semantic Kernel, or equivalent AI frameworks.
  • Experience deploying production AI services using REST APIs, microservices, containers, Kubernetes, CI/CD pipelines, cloud-native services, and monitoring platforms.
  • Strong understanding of model evaluation, fine-tuning, inference optimization, secure data handling, and AI application performance tuning.
  • Experience working with cloud platforms and distributed AI workloads.
  • Excellent problem-solving, software engineering, and collaboration skills., * Experience within Banking, Financial Services, FinTech, Risk Management, Compliance, Financial Crime, Operations, or Enterprise Technology.
  • Experience with Azure OpenAI, AWS Bedrock, Google Vertex AI, Databricks, vLLM, Triton Inference Server, MLflow, Kubeflow, AI model gateways, or similar enterprise AI platforms.
  • Familiarity with Responsible AI, AI Governance, Model Risk Management, Audit Controls, AI Cost Governance, and private or open-source LLM deployments.
  • Experience deploying enterprise-scale AI platforms in regulated environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

BalÔzs Kiss · WWC 2023

2:27 min

Managing traffic and tracking costs with Databricks Unity Catalog

Viktoria Semaan Viktoria Semaan Ā· WWC Europe 2026

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter Ā· WWC 2022

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou Ā· Coffee With Developers

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa Ā· LIVE

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski Ā· LIVE

Videos

See all

Related articles

See all