> Markdown version of [/jobs/ext/3543890-ai-data-platform-engineer](https://www.wearedevelopers.com/jobs/ext/3543890-ai-data-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Data Platform Engineer - **Company:** Banner Quality Management Inc. - **Location:** Brook Park, OH, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** LangGraph Framework, .NET Framework, Application Programming Interfaces (APIs), Agile Methodology, Artificial Intelligence, Airflow, Amazon Web Services, Audit Trail, User Authentication, C Sharp (Programming Language), Cloud Database, Code Review, Encodings, Content Analysis, Information Engineering, Data Governance, Data Infrastructure, Data Security, Relational Databases, Digital Architecture, Identity and Access Management, Information Extraction, Python (Programming Language), Microsoft SQL Server, Natural Language Processing, Node.Js, OAuth, Open Source Technology, OpenID, Azure Active Directory, Standard Sql, Search Technologies, Microsoft SharePoint, SQL Databases, Systems Integration, TypeScript, Unstructured Data, Management of Software Versions, Web Applications, Data Processing, Enterprise Software Applications, Data Classification, Okta, Pytorch, Retrieval-Augmented Generation, Large Language Models, Apache Spark, Hallucination Detection, Llamaindex, Generative AI, Backend, Agentic-AI, Build Management, Containerization, Data Lakes, AI Platforms, Git Flow, Scikit Learn, Information Technology, Amazon Bedrock, HuggingFace, Data Management, LiteLLM, Front End Software Development, Api Design, Restful APIs, Data Pipelines, Serverless Computing, Docker - **Published:** October 2, 2026 - **Apply:** https://www.juju.com/job/16_31096255 ## About the Role * Bachelor's or Master's degree in Computer Science, Data Science, Artificial Intelligence, Engineering, or a related field (PhD preferred for senior levels). * 3+ years of hands-on experience building and deploying AI and data-intensive systems in production. * Proficiency in Python and strong SQL; experience with modern data engineering libraries and frameworks. * Hands-on AWS experience across storage, compute, data, and AI services (object storage, managed relational and vector databases, serverless and containerized compute, managed foundation model and knowledge base services), including IAM and security best practices. * Demonstrated experience designing data lake or lakehouse architectures and pipelines for heterogeneous data at scale, including document processing and embedding workflows. * Experience building and securing REST APIs and/or MCP servers: API design, authentication, versioning, containerization (Docker), and deployment. * Working knowledge of OIDC/OAuth 2.0 flows, JWTs, and enterprise identity providers (Microsoft Entra ID required; Keycloak a plus). * Strong understanding of data handling in regulated/high-stakes domains: CUI/PII-aware design, access controls, auditability, and NLP over technical documents., * Experience implementing GenAI solutions (not just experimentation) using APIs, SDKs, and frameworks (e.g., Amazon Bedrock, LangChain/LangGraph, LlamaIndex, LiteLLM or other model gateways, Hugging Face). * Work with large language models (RAG, embeddings and vector databases, agent frameworks, NL-to-SQL) for technical document analysis or knowledge extraction. * Experience with data workflow orchestration and processing tools (e.g., AWS Step Functions, Glue, Lambda, Airflow, dbt, Spark); specific tooling is not prescribed and candidates are expected to help select it. * Experience evaluating and monitoring LLM applications (e.g., RAG evaluation frameworks, LLM-as-judge, tracing/observability platforms, OpenTelemetry). * Experience implementing robust guardrails, content safety, and groundedness controls. * ML literacy sufficient to operationalize, monitor, and retrain existing models and embedding pipelines; familiarity with scikit-learn or PyTorch a plus. * C#/.NET a plus, for integration with existing web applications and backend services. * TypeScript/Node.js a plus, for MCP servers and frontend integration. * Experience integrating with M365/SharePoint (e.g., Microsoft Graph API). * Experience with Git/GitHub workflows, branching strategies, and code review in an Agile team. * Prior Federal government, defense, or contracting experience Personality or self-management skills: * Strong communication, presentation, and storytelling skills * Experience collaborating with multiple cross-functional partners * Experience working in a fast-paced, iterative environment where multi-tasking and time-management skills are critical * Thrive under pressure and enjoy working in an environment with competing priorities * Experience supporting developers in the implementation of design deliverables * Ability to perform work within specific timeframes and adhere to deadlines * Willingness to keep skills current and quickly adapt to new technologies, as needed. ## Description * Architect, build, and operate a cloud data lake on AWS that ingests and organizes structured, semi-structured, and unstructured data from relational databases (e.g., SQL Server, managed RDS), object storage, documents, transcripts, RSS/news feeds, and M365/SharePoint sources, and makes that data AI-ready. * Design and build end-to-end data pipelines: ingestion, transformation, document parsing and chunking, metadata enrichment, embedding generation, and loading into vector/semantic and structured stores, with orchestration, scheduling, data quality checks, and monitoring. * Develop and operationalize retrieval-augmented generation (RAG) and agentic GenAI capabilities using AWS-native foundation model services and open frameworks, including knowledge bases, embeddings, vector search, natural-language-to-SQL, and tool use over the various datasets. * Design, develop, secure, and document APIs and MCP servers that expose lake datasets and enterprise systems to LLM orchestrators, chat interfaces, and web applications as governed, discoverable tools. * Implement authentication and authorization for AI services, APIs, and data access using OIDC/OAuth 2.0 and Microsoft Entra ID; configure and administer identity providers (including Keycloak) for scopes, roles, service-to-service authentication, and federation. * Build evaluation and observability for GenAI systems: retrieval quality, answer groundedness, hallucination detection, validation of NL-to-SQL outputs against source data, tracing, and cost/latency monitoring. * Embed data governance into pipelines and retrieval: data classification and marking, document-level and row/column-level access controls, PII handling, lineage, and audit logging. * Design and optimize prompts and agent workflows for large language models (LLMs) to ensure accurate, grounded, context-aware outputs with source citations. * Evaluate and recommend tools: help the team select the right AWS services, open-source, and third-party components for each job based on capability, cost, security accreditation, and maintainability, avoiding over-engineering. * Integrate AI capabilities into existing web applications, tools, training platforms, and M365/SharePoint solutions. * Collaborate with subject-matter experts, engineers, data stewards, and mission programs to define use cases, gather requirements, and measure impact (e.g., reduced analysis time, improved risk foresight, higher detection accuracy). * Ensure AI systems follow the Customer's "AI advises, humans decide" posture and responsible AI principles: human-in-the-loop review, source traceability, clear distinction between retrieved data and AI-generated content, explainability, robustness, safety, and compliance with federal guidelines (e.g., NIST AI Risk Management Framework, Customer AI governance). * Stay current on frontier AI advances (e.g., agentic systems, the MCP ecosystem, trustworthy AI) and adapt them to constrained, high-assurance aerospace domains. * Document architecture, pipelines, code, and processes to support reproducibility, audits, and knowledge transfer across the Customer organization. * Must live within Greater Cleveland Area. * No relocation funding available ## Related Videos - [Beyond GPT: Building Unified GenAI Platforms for the Enterprise of Tomorrow](https://www.wearedevelopers.com/videos/1525-beyond-gpt-building-unified-genai-platforms-for-the-enterprise-of-tomorrow) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Keeping applications secure by evolving OAuth 2.0 and OpenID Connect](https://www.wearedevelopers.com/videos/100152-keeping-applications-secure-by-evolving-oauth-2-0-and-openid-connect) - [Inside the AI Revolution: How Microsoft is Empowering the World to Achieve More](https://www.wearedevelopers.com/videos/869-inside-the-ai-revolution-how-microsoft-is-empowering-the-world-to-achieve-more) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Delay the AI Overlords: How OAuth and OpenFGA Can Keep Your AI Agents from Going Rogue](https://www.wearedevelopers.com/videos/1637-delay-the-ai-overlords-how-oauth-and-openfga-can-keep-your-ai-agents-from-going-rogue) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)