AI Engineer

MarineTraffic
United States
1 day ago
Apply on www2.jobdiva.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$26,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Data Systems Information Extraction Python (Programming Language) OpenAI Search Technologies Retrieval-Augmented Generation Large Language Models Ollama Machine Learning Operations
+4 more
Virtual Agents Evaluation of Large Language Models Databricks Web Api

Job description

Build RAG and document-processing pipelines on the Databricks lakehouse - ingestion, OCR of mixed-quality sources, chunking, embedding, and retrieval. Build LLM workflows for summarization, structured extraction, and evidence-grounded generation with source attribution. Generate synthetic document corpora with the fidelity and quality variation needed for meaningful results. Stand up the evaluation harness - retrieval quality, groundedness and hallucination checks, structured-output validity, human-in-the-loop review. Report results in numbers. Package deliverables as jobs and Asset Bundles, tracked in MLflow, and document everything the internal team needs to own the work.

Requirements

U.S. citizenship and active T5/SSBI federally adjudicated clearance required. [8]+ years building applied ML/AI or data systems, with demonstrated delivery of LLM and RAG systems you personally built - not notebook demos. Hands-on Databricks. Document processing at scale: OCR, layout-aware parsing, chunking tradeoffs, poor-quality source handling. Local/self-hosted LLM serving - vLLM, TGI, Ollama, llama.cpp, or equivalent - including running open-weight models in an isolated or air-gapped environment without reliance on external API endpoints. Structured extraction and grounded generation with source attribution. LLM evaluation methodology - you can explain how you measured correctness and what the evaluation missed. Privacy-preserving synthetic data generation from CUI, PII, or comparably restricted source data, with an understanding of re-identification risk. Strong Python. Government or defense contracting experience., RAG built inside a government or FedRAMP-authorized environment (Azure OpenAI in GCC High, AWS GovCloud, Bedrock within an authorized boundary). Direct experience with FedRAMP Moderate, NIST 800-171, CMMC L2, or CUI handling. Databricks Vector Search, Mosaic AI Agent Framework and Agent Evaluation, Asset Bundles, MLflow. Unity Catalog governance. H2O (h2oGPTe, Driverless AI).

Soft Skills Self-directed execution against a fixed milestone with minimal oversight. Honest reporting of model behavior - comfortable saying what an evaluation does and does not establish. Collaboration across technical and non-technical teams. Clear documentation and active knowledge transfer.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www2.jobdiva.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:46 min

Integrating OpenAI REST APIs within Unity environments

Zaid Zaim Zaid Zaim · World Congress 2024

1:56 min

Demonstrating a local recipe finder powered by Ollama

Sandra Ahlgrimm Sandra Ahlgrimm +1 · World Congress 2024

9:56 min

Expanding browser capabilities with modern web APIs

Ire Aderinokun · JS Congress

2:12 min

Navigating technical clarity as a global black belt

Chris Heilmann Chris Heilmann +2 · LIVE

3:25 min

Introduction to OpenAI and SingleStore for financial bots

Akmal Chaudhri Akmal Chaudhri · LIVE

1:31 min

Essential AI and human skills for future teams

Alexander Weißhaupt Alexander Weißhaupt +1 · World Congress 2025

Videos

See all

Related articles

See all