> Markdown version of [/jobs/ext/2709172-ml-engineer](https://www.wearedevelopers.com/jobs/ext/2709172-ml-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Engineer - **Company:** Albert Invent Corp - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Microsoft Azure, Big Data, Unix, Cloud Computing, Continuous Integration, Information Engineering, Data Integration, Data Systems, Distributed Systems, Fault Tolerance, Monitoring of Systems, Information Retrieval, Machine Learning, Open Source Technology, Performance Tuning, Systems Development Life Cycle, Software Tools, Prometheus, Software Engineering, Data Storage Management, Google Cloud, Delivery Pipeline, Large Language Models, Grafana, Reliability of Systems, Backend, Git, Fastapi, Containerization, Data Lakes, Kubernetes, Information Technology, Dask, Machine Learning Operations, Virtual Agents, Restful APIs, Software Version Control, Docker, Microservices - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-staff-ai-engineer-job-2-8293667 ## About the Role * A strong passion for AI/ML engineering and scalable data systems. * An ability to prioritize system scalability and fault tolerance while focusing on innovative AI/ML solutions. * A collaborative mindset and excellent communication skills. * A commitment to quality and innovation in AI and data engineering. Key competencies * A degree in Computer Science, AI, or a related field with 7+ years of industry experience (Bachelor's) or 5+ years (Master's or PhD) in software engineering, emphasizing expertise in building scalable, fault-tolerant AI systems. * Advanced knowledge of modern AI frameworks (e.g., LangChain, LangGraph, AutoGen, Crew.ai). * Experience with vector databases (e.g., Pinecone, Milvus, Pgvector, ChromaDB). * Strong understanding of distributed systems and microservices architecture. * Proficiency in REST API development using FastAPI REST Framework or similar tools. * Familiarity with cloud platforms (e.g., AWS, GCP, Azure) and containerization technologies (e.g., Docker, Kubernetes). * Proven track record of deploying production-grade AI systems. * Experience leading technical teams and fostering collaborative environments. Preferred/Bonus Points * Advanced degree in Computer Science, AI, or related fields. * Background in chemical and materials science applications. * Contributions to open-source AI projects. * Experience with fine-tuning and optimizing large language models (LLMs). * Expertise in system monitoring and observability tools (e.g., Prometheus, Grafana). * Experience with data engineering tools (e.g., Airflow, Delta Lake, Dask). * Familiarity with Unix scripting and version control systems (e.g., Git). * Prior experience working in AI/ML-focused environments using tools such as Ray, Torch, Kubeflow. * Experience mentoring junior developers/engineers in best practices. ## Description We are seeking a highly motivated and talented individual with a passion for AI/ML engineering and agent technologies. In this role, you will unleash your creativity, intelligence, and curiosity to build scalable AI systems that empower researchers and chemists at leading chemical and materials organizations to push the boundaries of innovation. As an ML Engineer specializing in LLMs and agent technologies, you will play a critical role in our mission to streamline workflows and provide robust, scalable solutions that support AI/ML capabilities for thousands of researchers worldwide. This role is central to the development of autonomous systems and tools tailored to chemical and materials science applications. What you'll do We are seeking an exceptional ML Engineer with a focus on LLMs and RAG systems. This role prioritizes designing and developing scalable, fault-tolerant AI systems while maintaining a strong focus on domain-specific AI solutions. You will play a critical role in building robust infrastructure to support high-performance applications and tools, enabling seamless data integration and transformation to power AI/ML capabilities in chemical and materials science. Scalable AI System Development: * Design, build, and maintain scalable, fault-tolerant AI systems leveraging OpenAI and Anthropic models. * Develop RAG architectures to ensure efficient, high-performance information retrieval tailored to chemical and materials science. * Optimize system performance to handle large-scale data and application demands. AI Agent Development: * Build and maintain intelligent AI agents using modern frameworks. * Collaborate with domain experts to refine agent capabilities for specific scientific workflows. Data Engineering and Integration: * Architect and maintain vector database solutions for efficient data storage and retrieval. * Develop pipelines for ingestion, transformation, and storage to enable AI/ML workflows. * Collaborate with platform and ML engineers to integrate AI/ML models with backend systems. System Reliability and Fault Tolerance: * Implement robust error-handling, monitoring, and alerting mechanisms to ensure system resilience. * Troubleshoot and resolve system bottlenecks and failures. CI/CD and Deployment Pipelines: * Design, implement, and maintain CI/CD pipelines for AI systems and data workflows. * Promote automation and best practices to enhance the development lifecycle. Adoption of Emerging Technologies: * Stay informed on the latest trends and tools in AI/ML engineering and agent technologies. * Introduce and implement new technologies to improve system scalability, data integration, and developer productivity. Collaboration and Cross-Functional Teamwork: * Work closely with AI/ML, data engineering, and platform teams to understand and deliver on technical requirements. * Contribute to architectural decisions that impact the overall platform ecosystem. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [WeAreDevelopers LIVE - Node and Package Security](https://www.wearedevelopers.com/videos/2138-wearedevelopers-live-node-and-package-security) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [The Time Paradox: Building Timezone-Safe Python/Django Applications](https://www.wearedevelopers.com/videos/1915-the-time-paradox-building-timezone-safe-python-django-applications) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)