> Markdown version of [/jobs/ext/3028357-research-informatics-software-engineer](https://www.wearedevelopers.com/jobs/ext/3028357-research-informatics-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Research Informatics Software Engineer - **Company:** Excelsior Sciences Inc. - **Location:** New York, NY, United States - **Salary:** $120,000.0 - $160,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon Web Services, Business Analytics Applications, Data Analysis, C++ (Programming Language), Cloud Computing, Continuous Integration, Data Validation, Information Engineering, Data Infrastructure, Cursor (Graphical User Interface Elements), Digital Technology, Graph Database, Python (Programming Language), Laboratory Information Management Systems, PostgreSQL, Neo4j, Performance Tuning, Scientific Computating, Software Construction, Software Engineering, SQL Databases, Enterprise Software Applications, Feature Engineering, GitHub Copilot, Pytorch, Large Language Models, Multi-Agent Systems, Apache Spark, Backend, Kubernetes, Information Technology, Virtual Agents, Data Pipelines, Docker, Databricks - **Published:** September 22, 2026 - **Apply:** https://www.thejobnetwork.com/job/30eafd6f-d2e8-42c8-9dc4-ee767b3a7262/research-informatics-software-engineer ## About the Role * Master's degree in Computer Science (or a closely related field) with relevant coursework in cloud computing and the fundamentals of AI and ML. * Demonstrated experience building data pipelines, feature engineering, or scientific data workflows (e.g., Spark/Databricks-style pipelines, data quality checks, performance tuning). * Hands-on experience with cloud platforms (AWS), containers (Docker/Kubernetes), and modern data/backend tools (SQL, PostgreSQL, orchestration frameworks). * Strong proficiency with AI coding assistants and coding agents (e.g., Cursor, Claude Code, GitHub Copilot, or similar tools). * Familiarity with LLM concepts, RAG, retrieval, or multi-agent systems (coursework, projects, or professional exposure). * Willingness and aptitude to rapidly learn commercial LIMS/ELN or analytical platforms (e.g., Genedata, CDD Vault, Virscidian Analytical Studio); prior exposure is a plus. * Proficiency in Python and SQL; additional experience with C++/C, high-performance ML tooling, or scientific computing libraries is a plus. * Strong collaboration skills and ability to work at the intersection of software engineering, data, and scientific applications. Preferred Qualifications * Practical, hands-on experience with LLM / agentic AI systems, including RAG, GraphRAG, multi-agent architectures, or production retrieval-augmented pipelines. * Experience optimizing high-performance ML or scientific models (e.g., protein structure prediction, surrogate modeling, Bayesian optimization). * Hands-on work with multi-agent systems, knowledge graphs (Neo4j), or agent frameworks/SDKs. * Familiarity with Airflow or Prefect, PyTorch, and related ML/LLM tooling. * Direct experience with commercial LIMS/ELN or analytical platforms such as Genedata, CDD Vault, Virscidian Analytical Studio, or similar-especially their data models, APIs, and integration points. * Experience with CI/CD, testing, schema design, and production-grade software practices. * Interest in applying agentic AI and robust data engineering to laboratory and drug-discovery workflows. ## Description * Integrate, extend, and support vendor Laboratory Information Management Systems (LIMS), Electronic Lab Notebooks (ELN), and analytical informatics platforms. Scope includes platforms such as Genedata, CDD Vault, Virscidian Analytical Studio, and similar systems-focusing on data models, workflows, APIs, sample/analytical data flows, and connections to instruments and enterprise systems. * Design, implement, and maintain scalable data pipelines and APIs that make scientific data (samples, assays, analytical results, automation streams) FAIR, high-quality, and machine-actionable for both human scientists and AI agents. Leverage modern data platforms, warehouses/lakes, and orchestration tools. * Build and operate cloud-native components (primarily AWS) using containers (Docker/Kubernetes), infrastructure patterns, CI/CD, and workflow orchestration to support lab informatics and AI workloads. * Prototype and productionize agentic AI / GenAI solutions-LLM agents, RAG and GraphRAG systems, multi-agent workflows, and prompt-engineered / retrieval-augmented pipelines-that automate or augment laboratory informatics processes, data interpretation, and closed-loop experimentation. * Collaborate with research scientists and cross-functional engineering teams to translate scientific needs into reliable software, data products, and AI capabilities; contribute to documentation, testing, and knowledge transfer. * Apply software engineering best practices (agile / AI-agile delivery, testing, schema design, performance tuning) in a scientific computing context. * Support continuous improvement of lab digital systems, including data quality, observability, and readiness for AI agents. ## Related Videos - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Carl Lapierre - Exploring Advanced Patterns in Retrieval-Augmented Generation](https://www.wearedevelopers.com/videos/1235-carl-lapierre-exploring-advanced-patterns-in-retrieval-augmented-generation) ## Related Articles - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)