> Markdown version of [/jobs/ext/2987773-staff-senior-software-engineer-document-intelligence-data-infrastructure](https://www.wearedevelopers.com/jobs/ext/2987773-staff-senior-software-engineer-document-intelligence-data-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff/Senior Software Engineer, Document Intelligence & Data Infrastructure - **Company:** Lever, Inc. - **Location:** New York, NY, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Data Conversion, Data Infrastructure, Fault Tolerance, Open Source Intelligence, Simple Data Format, Management of Software Versions, Web Crawlers, Real Time Systems, Helios, Data Pipelines - **Published:** September 18, 2026 - **Apply:** https://jobs.lever.co/unusual/543f45a2-8858-4f1e-a940-fb9e42e7df8e/apply ## About the Role * Experience building international public-sector data pipelines across languages and jurisdictions. * Production experience with real-time transcription, speaker diarization, and tone or prosody analysis. * Experience operating large-scale web acquisition against complex and frequently changing sources. * Familiarity with legal, legislative, regulatory, archival, or similarly difficult corpora. * Experience deploying secure processing systems in GovCloud, air-gapped, or other restricted environments. ## Description Build and operate the acquisition and processing infrastructure that turns global public-sector and open-source data into reliable, searchable intelligence. This role owns the path from source discovery and collection through normalization, extraction, enrichment, and publication. You will expand our platform to include new countries, languages, institutions, and source formats covering many additional areas of open source intelligence to advance Proxi's natural capabilities. We are looking for a senior engineer who has built and operated large-scale acquisition or document-processing systems in production. Expected areas of expertise: * Web crawlers and connectors for continuously changing international sources. * Fault-tolerant pipelines with reliable scheduling, recovery, replay, and backfill capabilities. * Processing complex documents, structured files, images, audio, and video across languages. * Stable schemas, source lineage, correction propagation. * Operating secure, observable, scalable, and cost-efficient processing systems across cloud and restricted environments. * Document conversion, OCR, layout analysis, and structural extraction across complex file formats. * Multilingual audio, video, and document processing with reliable alignment and attribution. * Scalable inference infrastructure for batch and real-time processing across CPU and GPU workloads. * End-to-end provenance, versioning, correction propagation, and reproducible reprocessing. * Secure handling of untrusted content with rigorous quality evaluation, observability, and cost controls., * Expand Proxi's ingestion and document-processing platform from source acquisition through downstream publication. * Construct crawling and connector infrastructure required to discover, acquire, and continuously maintain international data sources. * Scale the distributed processing environment that supports document conversion, extraction, transcription, enrichment, replay, and backfills. * Establish common data contracts that normalize heterogeneous and multilingual sources without discarding their jurisdictional context or provenance. * Maintain pipeline reliability, security, observability, and cost efficiency while ensuring corrections propagate through every downstream system. * Build and operate real-time audio intelligence pipelines for low-latency multilingual transcription, speaker attribution, timestamp alignment, and confidence-calibrated tone analysis across live and recorded media. ## Related Videos - [Building an AI-Ready Content Lake: Scaling RAG and Document AI Beyond Demos](https://www.wearedevelopers.com/videos/1977-building-an-ai-ready-content-lake-scaling-rag-and-document-ai-beyond-demos) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [WeAreDevelopers Live: Browser Extensions, Honey Scam, Jailbreaking LLMs and more](https://www.wearedevelopers.com/videos/1286-wearedevelopers-live-browser-extensions-honey-scam-jailbreaking-llms-and-more) - [Implementing continuous delivery in a data processing pipeline](https://www.wearedevelopers.com/videos/73-implementing-continuous-delivery-in-a-data-processing-pipeline) - [Python-Based Data Streaming Pipelines Within Minutes](https://www.wearedevelopers.com/videos/1233-python-based-data-streaming-pipelines-within-minutes) - [ChatGPT vs Google: SEO in the Age of AI Search - Eric Enge](https://www.wearedevelopers.com/videos/1287-chatgpt-vs-google-seo-in-the-age-of-ai-search-eric-enge) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [Dev Digest 139 - Soft and hard queries](https://www.wearedevelopers.com/magazine/487-dev-digest-139-soft-and-hard-queries)