> Markdown version of [/jobs/ext/2711585-software-engineer](https://www.wearedevelopers.com/jobs/ext/2711585-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # software engineer - **Company:** Iambic Therapeutics, Inc - **Location:** San Diego, CA, United States (Remote available) - **Experience:** Starter - **Salary:** $110,000.0 - $162,000.0 - **Contract:** Permanent contract - **Skills:** Test Suite, Access Network, Artificial Intelligence, Amazon Web Services, Amazon S3, Applicant Tracking Systems, Audit Trail, Big Data, Cloud Computing, Data Validation, Data Cleansing, Information Engineering, Extract Transform Load (ETL), Distributed Data Store, Document-Oriented Databases, Github, Python (Programming Language), Knowledge Management, Machine Learning, Performance Tuning, Software Engineering, Unstructured Data, Web Pages, Workflow Management Systems, Chatbots, Large Language Models, Multi-Agent Systems, Pytest, Kubernetes, Software Coding, GPT, Docker, Servicenow - **Published:** September 4, 2026 - **Apply:** https://www.builtincolorado.com/job/software-engineer-agentic-data-pipelines/11014282?handler=ApplyRedirect ## About the Role This role is ideal for candidates who combine strong software engineering instincts with scientific understanding of biomedical data, and who are excited about using LLMs as tools to solve practical data problems., * Master's degree in a computational STEM field, or a Bachelor's with 2+ years of industry experience * Strong Python engineering skills, including experience building and maintaining production-quality software * Hands-on experience with LLM APIs (e.g., Claude, GPT) and agentic patterns such as tool use, orchestration, and multi-step reasoning * Familiarity with biomedical or chemical data sources and formats (e.g., PDB, UniProt, ChEMBL, SDF/MOL, FASTA, or similar) * Comfort with data engineering fundamentals: ETL design, data validation, and working with structured and unstructured data at scale * Hands-on experience with Python testing frameworks (e.g., pytest fixtures, parametrization), * Experience with agent orchestration frameworks, and with evaluation harnesses for LLM-generated code * Familiarity with cloud infrastructure and workflow orchestration (e.g., AWS, Docker, Kubernetes) * Knowledge of multimodal biomedical data-spanning small molecules, proteins, assays, images, 'omics, and/or clinical records * Experience with large-scale dataset construction or curation for ML model training * Knowledge of agent security practices: sandboxing, scoped credentials, prompt injection * Interest in a longer-term project: a natural language orchestrator that lets drug prosecution team members request inference, fine-tuning, virtual screens, and dataset analysis without writing code using our internal tools ## Description Build agentic LLM systems that acquire, clean, validate, curate, and document large-scale biomedical datasets for multimodal model training. Develop reliable agent architectures, tooling, evaluation harnesses, quality-control workflows, and reproducible data pipelines. Collaborate with ML scientists, monitor distributed production pipelines, diagnose failures, and implement secure sandboxed execution with restricted access and audit logging., We are seeking a software engineer to join our team at Iambic Therapeutics, working on data acquisition and curation for Enchant, our multimodal transformer model trained at scale on a wide variety of biomedical data. In this role, you will design and build agentic systems that generate code to acquire, clean, format, quality-control, and generate auditable data reports for the large-scale datasets that power Enchant training. The model writes the code, the code runs the pipeline. You will work at the intersection of LLM-based automation and biomedical data engineering-developing AI agents that can navigate heterogeneous data sources, enforce quality standards, and operate reliably at scale., * Design, build, and maintain agentic systems that turn a pointer to a biomedical data source (web page, S3 path, GitHub repository, a table in a paper) into a reviewed, versioned dataset. The agent's output is committed pipeline code and its test suite, so re-running it later reproduces the same dataset. Develop LLM-based pipelines for data cleaning, normalization, and formatting across diverse data modalities (e.g., molecular, genomic, clinical, literature) * Implement automated quality-control workflows that detect anomalies, flag inconsistencies, and enforce data standards * Evaluate and iterate on agent architectures, prompting strategies, tool definitions, validation loops, and evaluation harnesses that make agent-generated code trustworthy, improving reliability and throughput over time * Collaborate with ML scientists on the Enchant team to understand data requirements and translate them into scalable acquisition and processing systems * Monitor and maintain distributed data pipelines in production, diagnosing failures and improving robustness over time * Document data provenance, processing decisions, and quality metrics to support reproducibility and auditing * Operate the agents safely with sandboxed execution, least-privilege credentials, restricted network access, audit logs, and raising potential security risks to the team, Leads architecture and hands-on engineering across a futures commission merchant platform, including market data, order routing, settlement, reporting, risk, and liquidation. Drives replatforming away from legacy vendors, improves reliability and technical quality, partners with product, risk, operations, finance, and compliance, and mentors engineers. Requires extensive backend distributed-systems experience, direct FCM or derivatives expertise, regulated-finance experience, and proficiency in Go, Java, or similar technologies., Artificial Intelligence * Machine Learning * Natural Language Processing * Software * Conversational AI Lead and coach a recruiting team while owning day-to-day talent acquisition execution across engineering, research, and go-to-market functions. Run full-cycle recruiting, support complex searches, build AI-assisted workflows and scalable recruiting systems, use funnel data to diagnose performance, develop recruiters, and improve hiring quality through structured interviewing and stakeholder coaching. Top Skills: AIAshbyJuiceboxMetaview PNC Bank ## Related Videos - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [pytest: Simple, rapid and fun testing with Python](https://www.wearedevelopers.com/videos/213-pytest-simple-rapid-and-fun-testing-with-python) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Automagic Configuration in Python](https://www.wearedevelopers.com/videos/363-automagic-configuration-in-python) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [13 AI Tools You Have to Try](https://www.wearedevelopers.com/magazine/219-13-ai-tools-you-have-to-try) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [The Prompt Engineer ✍️](https://www.wearedevelopers.com/magazine/216-the-prompt-engineer)