Generative AI Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+7 more
Job description
Philadelphia Gas Works (PGW) is seeking an experienced Generative AI Engineer / On-Prem LLM & Vector Database Consultant to design, deploy, and optimize an enterprise-grade on-premises Large Language Model (LLM) and Vector Database solution in a secure environment. The ideal candidate will possess strong expertise in open-source LLM deployment, Retrieval-Augmented Generation (RAG), vector databases, semantic search, and enterprise AI infrastructure while ensuring data privacy, security, and high-performance inference in private or air-gapped environments., Design and implement an on-premises LLM architecture for secure enterprise AI applications. Deploy and optimize open-source LLMs including Meta Llama 3, Mistral, and Mixtral. Build and implement Retrieval-Augmented Generation (RAG) pipelines. Design, configure, and manage Vector Database solutions for semantic search. Develop Python-based LLM inference, orchestration, prompt engineering, and integrations. Optimize model performance using CPU inference, model quantization, and inference tuning. Generate and manage embeddings, metadata filtering, and semantic search workflows. Implement enterprise-grade authentication, authorization, access controls, and audit logging. Ensure compliance with secure, private, and air-gapped deployment requirements. Deliver deployment architecture, technical documentation, and knowledge transfer sessions. Build a working prototype integrating LLM + Vector Database + RAG Pipeline. Collaborate with engineering teams to deploy scalable AI infrastructure.
Requirements
Strong experience deploying Open-Source LLMs (Meta Llama 3, Mistral, Mixtral). Proficient in Python, Prompt Engineering, LLM Inference, Model Orchestration, and AI Integration. Experience with CPU-based inference, Model Quantization, and performance optimization. Hands-on experience with Vector Databases (Qdrant, Chroma, Milvus, pgvector). Proven expertise in building Retrieval-Augmented Generation (RAG) pipelines. Experience with Embeddings, Metadata Filtering, and Semantic Search. Strong knowledge of Enterprise Security, Data Privacy, and Air-Gapped Deployments. Experience implementing Authentication, Authorization, Access Controls, and Audit Logging. Preferred Qualifications: Experience with LangChain and/or LlamaIndex. Knowledge of Rust, Go, or C++. Experience with Docker and Kubernetes for on-prem deployments. Familiarity with inference frameworks: o vLLM o llama.cpp o Hugging Face Transformers Experience working in regulated or enterprise environments. Experience designing enterprise AI reference architectures.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
How to Become an AI Engineer
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?