> Markdown version of [/jobs/ext/2718644-backend-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2718644-backend-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Backend & Infrastructure Engineer - **Company:** Albert Invent Corp - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon Web Services, Microsoft Azure, Cloud Computing, Computer Programming, Data Infrastructure, Data Security, Data Stores, Distributed Systems, Fault Tolerance, Python (Programming Language), Open Source Technology, Performance Tuning, Queueing Systems, RabbitMQ, Redis, Software Tools, Prometheus, Scientific Computating, Software Engineering, Datadog, Data Logging, Pulumi, Autoscaling, Flask (Web Framework), Grafana, Caching, Reliability of Systems, Backend, Fastapi, Event Driven Architecture, Build Management, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Low Latency, Apache Kafka, Machine Learning Operations, Restful APIs, Terraform, Microservices - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/staff-ml-ops-engineer-job-2-8293668 ## About the Role * Deep expertise in Python backend development and distributed systems * Strong Kubernetes and cloud infrastructure experience * A builder's mindset-you want to create foundational systems that others build on * Genuine interest in science and technology; curiosity about how your work enables scientific discovery * A commitment to building systems that are reliable, maintainable, and scalable Key competencies * A degree in Computer Science or a related field with 7+ years of industry experience (Bachelor's) or 5+ years (Master's or PhD) in software engineering * Experience supporting AI/ML teams or deploying ML systems in production * Experience with GPU workloads and scheduling * Advanced proficiency in Python including async programming and performance optimization * Deep experience with Kubernetes-cluster management, networking, autoscaling, and troubleshooting * Strong background in distributed systems and microservices architecture * Experience with cloud platforms (AWS, GCP, or Azure) and infrastructure-as-code * Proficiency in REST API development using FastAPI, Flask, or similar * Experience with containerization and CI/CD pipelines * Track record of operating production systems at scale Preferred/Bonus Points * Familiarity with scientific computing or research environments * Background in or curiosity about chemistry, materials science, or related fields * Familiarity with data engineering tools (Airflow, Dagster, or similar) * Experience with vector databases or search infrastructure * Expertise in observability tools (Prometheus, Grafana, Datadog) * Experience with message queues and event-driven architectures (Kafka, Redis, RabbitMQ) * Contributions to open-source projects * Experience mentoring engineers ## Description As our Backend & Infrastructure Engineer, you will architect and build the core systems that power everything our AI/ML team delivers-the APIs, infrastructure, and distributed systems that make intelligent capabilities possible at scale. This is a foundational role: you'll shape how AI gets built and shipped here., We are seeking a highly motivated and talented individual with deep expertise in Python backend development, Kubernetes, and distributed systems. You'll be embedded with ML engineers and researchers, building robust systems that turn ambitious AI ideas into production realities-whether that's powering agent-based workflows, scaling inference, or enabling scientific computing pipelines. The infrastructure you build will directly enable researchers at the world's largest chemical and materials companies to leverage AI in ways that weren't possible before-accelerating discovery, enabling inverse design of novel materials, and transforming how science gets done. What you'll do Infrastructure & Kubernetes: * Design, deploy, and maintain Kubernetes infrastructure supporting AI/ML workloads * Manage containerized services, autoscaling, networking, and resource optimization Backend Development: * Design and build high-performance Python APIs and services using FastAPI or similar frameworks * Architect backend systems for scalability, reliability, and low latency * Build integrations between AI/ML systems and the broader Albert platform Distributed Systems: * Build and operate distributed systems that handle compute-intensive and high-throughput workloads * Design for fault tolerance, graceful degradation, and horizontal scalability * Implement async workflows, job queues, and task orchestration as needed Data Infrastructure: * Architect and maintain data pipelines and storage systems supporting AI/ML workflows * Work with vector databases, caches, and other data stores as required by ML systems * Ensure efficient data access patterns for training and inference workloads Reliability & Operations: * Implement observability including logging, metrics, tracing, and alerting * Own system reliability-troubleshoot issues, conduct post-mortems, and continuously improve * Design CI/CD pipelines and promote automation best practices * Implement infrastructure-as-code practices using Terraform, Helm, ArgoCd, Pulumi, or similar tools Collaboration: * Partner closely with ML engineers to understand requirements and deliver production-ready infrastructure * Translate ML prototypes and research code into scalable, maintainable systems * Contribute to technical decisions that shape the team's architecture ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)