> Markdown version of [/jobs/ext/2609005-ai-engineer-document-intelligence-financial-data-pipelines](https://www.wearedevelopers.com/jobs/ext/2609005-ai-engineer-document-intelligence-financial-data-pipelines). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Engineer - Document Intelligence & Financial Data Pipelines - **Company:** Everforth CyberCoders - **Location:** United States (Remote available) - **Salary:** $145,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Microsoft Azure, Databases, Content Analysis, Relational Databases, JSON, Python (Programming Language), PostgreSQL, RabbitMQ, Redis, SQL Databases, SQLAlchemy, Data Processing, Large Language Models, Fastapi, Data Management, Restful APIs, Automation Anywhere, Microservices - **Published:** August 24, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=3f90057ed5c92cbe ## About the Role * Python expertise. Deep proficiency with async Python, type annotations, and data manipulation; comfortable in a modern Python 3.12+ fully-typed codebase. * LLMs and multimodal vision. Hands-on production experience with frontier LLM APIs (OpenAI, Anthropic, or equivalent) using vision capabilities to read complex multi-page financial tables. * Structured AI outputs. Advanced use of Pydantic, Instructor, or native structured-output APIs to force LLMs into validated, system-ready JSON - not hoping the model behaves, enforcing it. * Document intelligence and OCR. Proven experience with enterprise document parsing pipelines; prior Azure Document Intelligence or Azure Content Understanding exposure strongly preferred. * Relational data discipline. Solid PostgreSQL data modeling and migration practices (SQLAlchemy/Alembic or equivalent); reconciliation logic lives in the database as much as in the pipeline. * Security and data-handling mindset. Familiarity with financial data obligations, zero-data-retention LLM agreements, regional endpoints, and why redaction must precede the API boundary. Nice to have * Domain experience in accounting, fintech, legal, or audit-compliance contexts * PDF internals: manipulating documents at the object or glyph level (pikepdf, pdfminer.six, pypdfium2, or similar) * Orchestration with distributed task queues - our pipelines run as taskiq workers over RabbitMQ and Redis * PII pipeline experience with privacy frameworks or custom NLP/regex hybrid engines * Observability instrumentation with OpenTelemetry shipped to Azure Monitor ## Description * Own the AI extraction layer. Design and maintain high-accuracy pipelines that parse tabular, structured, and unstructured financial data from scanned and digital PDFs - trust statements, payroll documents, vendor reports - using Azure Document Intelligence and frontier LLM vision APIs. * Build and harden PII redaction. Extend our in-house PDF redaction toolkit so that SSNs, EINs, account numbers, and names are masked programmatically before data crosses any external API boundary. * Design reconciliation workflows. Implement the cross-referencing and matching logic - deterministic where Python or SQL wins, agentic only where justified - that reconciles data extracted from disparate financial sources. * Deliver clean, documented REST APIs. Wrap AI workflows into secure FastAPI services with strict Pydantic-validated inputs and outputs that our core development team consumes with confidence. * Enforce guardrails and run evals. Detect and reject low-confidence extractions, build automated regression suites against real documents, and maintain measurable accuracy benchmarks over time. * Operate reliably on Azure. Deploy and monitor microservices on AKS, instrument with OpenTelemetry traces and metrics to Azure Monitor, and keep services auditable and production-grade. ## Related Videos - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Intro to FastAPI](https://www.wearedevelopers.com/videos/462-intro-to-fastapi) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [Garbage In, Garbage Out: Engineering Reliable AI Document Extraction Pipelines](https://www.wearedevelopers.com/videos/100301-garbage-in-garbage-out-engineering-reliable-ai-document-extraction-pipelines) - [Building and Deploying Multi-Agent Systems with ADK and Vertex AI](https://www.wearedevelopers.com/videos/1918-building-and-deploying-multi-agent-systems-with-adk-and-vertex-ai) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [13 AI Tools You Have to Try](https://www.wearedevelopers.com/magazine/219-13-ai-tools-you-have-to-try) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)