> Markdown version of [/jobs/ext/2850292-ai-engineer](https://www.wearedevelopers.com/jobs/ext/2850292-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Engineer - **Company:** Seneca Resources - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Data Systems, Information Extraction, Python (Programming Language), Search Technologies, Data Ingestion, Large Language Models, Machine Learning Operations, Virtual Agents, Databricks, Web Api, Data Generation - **Published:** September 11, 2026 - **Apply:** https://senecahq.com/wp-content/plugins/bullhorn-oscp/#/jobs/47514 ## About the Role U.S. citizenship and active T5/SSBI federally adjudicated clearance required. [8]+ years building applied ML/AI or data systems, with demonstrated delivery of LLM and RAG systems you personally built - not notebook demos. Hands-on Databricks. Document processing at scale: OCR, layout-aware parsing, chunking tradeoffs, poor-quality source handling. Local/self-hosted LLM serving - vLLM, TGI, Ollama, llama.cpp, or equivalent - including running open-weight models in an isolated or air-gapped environment without reliance on external API endpoints. Structured extraction and grounded generation with source attribution. LLM evaluation methodology - you can explain how you measured correctness and what the evaluation missed. Privacy-preserving synthetic data generation from CUI, PII, or comparably restricted source data, with an understanding of re-identification risk. Strong Python. Government or defense contracting experience. Preferred Qualifications RAG built inside a government or FedRAMP-authorized environment (Azure OpenAI in GCC High, AWS GovCloud, Bedrock within an authorized boundary). Direct experience with FedRAMP Moderate, NIST 800-171, CMMC L2, or CUI handling. Databricks Vector Search, Mosaic AI Agent Framework and Agent Evaluation, Asset Bundles, MLflow. Unity Catalog governance. H2O (h2oGPTe, Driverless AI). Soft Skills Self-directed execution against a fixed milestone with minimal oversight. Honest reporting of model behavior - comfortable saying what an evaluation does and does not establish. Collaboration across technical and non-technical teams. Clear documentation and active knowledge transfer. ## Description This role builds LLM and document-intelligence capabilities on a greenfield data and AI platform in a high-trust federal environment. Work is hands-on, scoped to a near-term product demonstration, and built to migrate - documented, portable, and transferred to the internal team. Engagement is 1 year with option to extend. Active T5/SSBI clearance required., * Develop RAG and document-processing pipelines on a cloud-based lakehouse platform (Databricks), including OCR, data ingestion, chunking, and retrieval. * Build workflows for LLM-based summarization, information extraction, and evidence-grounded text generation with source attribution. * Generate synthetic document corpora that capture fidelity and quality variation for meaningful analysis. * Establish evaluation frameworks for retrieval accuracy, groundedness, hallucination checks, and structured output validation. * Document processes, package deliverables, and track models and artifacts using tools like MLflow. * Collaborate with internal teams to transfer knowledge and facilitate operational deployment. ## Related Videos - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Building Sovereign AI: Lessons from Deploying Secure RAG Systems using Confidential Computing](https://www.wearedevelopers.com/videos/100108-building-sovereign-ai-lessons-from-deploying-secure-rag-systems-using-confidential-computing) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Web APIs you might not know about](https://www.wearedevelopers.com/videos/281-web-apis-you-might-not-know-about) - [How to Avoid LLM Pitfalls - Mete Atamel and Guillaume Laforge](https://www.wearedevelopers.com/videos/1328-how-to-avoid-llm-pitfalls-mete-atamel-and-guillaume-laforge) - [The Data Mesh as the end of the Datalake as we know it](https://www.wearedevelopers.com/videos/156-the-data-mesh-as-the-end-of-the-datalake-as-we-know-it) ## Related Articles - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [A 5-Step Open-Source Setup for Agentic Engineering](https://www.wearedevelopers.com/magazine/738-a-5-step-open-source-setup-for-agentic-engineering) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)