> Markdown version of [/jobs/ext/1936780-generative-ai-engineer](https://www.wearedevelopers.com/jobs/ext/1936780-generative-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Generative AI Engineer - **Company:** Philadelphia Gas Works - **Location:** Philadelphia, PA, United States (Remote available) - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Application Integration Architecture, Audit Trail, User Authentication, C++ (Programming Language), Databases, Database Development, Python (Programming Language), Metadata, Open Source Technology, Performance Tuning, Search Technologies, Systems Integration, AI Infrastructure, Large Language Models, Prompt Engineering, Generative AI, HuggingFace, Docker - **Published:** August 5, 2026 - **Apply:** https://www.dice.com/job-detail/e96452c7-b8d6-490f-858c-e9e465cdba3f ## About the Role Strong experience deploying Open-Source LLMs (Meta Llama 3, Mistral, Mixtral). Proficient in Python, Prompt Engineering, LLM Inference, Model Orchestration, and AI Integration. Experience with CPU-based inference, Model Quantization, and performance optimization. Hands-on experience with Vector Databases (Qdrant, Chroma, Milvus, pgvector). Proven expertise in building Retrieval-Augmented Generation (RAG) pipelines. Experience with Embeddings, Metadata Filtering, and Semantic Search. Strong knowledge of Enterprise Security, Data Privacy, and Air-Gapped Deployments. Experience implementing Authentication, Authorization, Access Controls, and Audit Logging. Preferred Qualifications: Experience with LangChain and/or LlamaIndex. Knowledge of Rust, Go, or C++. Experience with Docker and Kubernetes for on-prem deployments. Familiarity with inference frameworks: o vLLM o llama.cpp o Hugging Face Transformers Experience working in regulated or enterprise environments. Experience designing enterprise AI reference architectures. ## Description Philadelphia Gas Works (PGW) is seeking an experienced Generative AI Engineer / On-Prem LLM & Vector Database Consultant to design, deploy, and optimize an enterprise-grade on-premises Large Language Model (LLM) and Vector Database solution in a secure environment. The ideal candidate will possess strong expertise in open-source LLM deployment, Retrieval-Augmented Generation (RAG), vector databases, semantic search, and enterprise AI infrastructure while ensuring data privacy, security, and high-performance inference in private or air-gapped environments., Design and implement an on-premises LLM architecture for secure enterprise AI applications. Deploy and optimize open-source LLMs including Meta Llama 3, Mistral, and Mixtral. Build and implement Retrieval-Augmented Generation (RAG) pipelines. Design, configure, and manage Vector Database solutions for semantic search. Develop Python-based LLM inference, orchestration, prompt engineering, and integrations. Optimize model performance using CPU inference, model quantization, and inference tuning. Generate and manage embeddings, metadata filtering, and semantic search workflows. Implement enterprise-grade authentication, authorization, access controls, and audit logging. Ensure compliance with secure, private, and air-gapped deployment requirements. Deliver deployment architecture, technical documentation, and knowledge transfer sessions. Build a working prototype integrating LLM + Vector Database + RAG Pipeline. Collaborate with engineering teams to deploy scalable AI infrastructure. ## Related Videos - [How E.On productionizes its AI model & Implementation of Secure Generative AI.](https://www.wearedevelopers.com/videos/623-how-e-on-productionizes-its-ai-model-implementation-of-secure-generative-ai) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [A Data Mesh needs Open Metadata](https://www.wearedevelopers.com/videos/505-a-data-mesh-needs-open-metadata) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Make it simple, using generative AI to accelerate learning](https://www.wearedevelopers.com/videos/969-make-it-simple-using-generative-ai-to-accelerate-learning) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)