> Markdown version of [/jobs/ext/2722632-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2722632-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Infrastructure Engineer - **Company:** Reducto, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Software Debugging, Python (Programming Language), Open Source Technology, Large Language Models, Kubernetes, Infrastructure Automation Frameworks, Storage Technologies - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/infrastructure-engineer-reducto-8103157 ## About the Role * Have 5+ years of hands-on experience in building or supporting production-grade infrastructure and reliability processes for high-throughput systems. * Are comfortable with Python or similar languages, and exceptional at working across cloud platforms, container orchestration (e.g., Kubernetes), networking, and storage technologies. * Build your own tools on the fly to diagnose, experiment, and address reliability problems-whether it's an internal dashboard or an automated remediation workflow. * Bring a quantitative, hands-on approach to system operations, automation, and continuous improvement., * Have prior experience founding a company or building products/infrastructure in early-stage, high-growth environments. * Are excited about automating incident management processes with LLMs/AI. * Are driven, ambitious, and deeply care about both technical excellence and collaborative problem-solving. * Keep up with the latest trends in cloud, observability, and SRE best practices. * Are passionate about open-source and have contributed tools or automation to reliability communities. * Have built or optimized monitoring, incident response, or high-performance computing systems for demanding AI/ML, fintech, or enterprise clients. ## Description * Designing, building, and maintaining highly available, scalable infrastructure to support intensive AI/ML workloads and real-time model deployments. * Implementing robust monitoring, alerting, and observability systems to ensure system health, performance, and uptime across cloud and on-prem environments. * Debugging, optimizing, and automating infrastructure for fast iteration and rapid deployment cycles, focusing on both reliability and developer velocity. * Proactively identifying, investigating, and resolving incidents to minimize downtime and maintain world-class service levels for enterprise customers. * Collaborating closely with engineers, ML specialists, and founders to shape product, infrastructure, and security strategies. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Using AI Without Losing Your Skills](https://www.wearedevelopers.com/videos/2045-using-ai-without-losing-your-skills) - [Startup Presentation: StorX - Future of Cloud Storage](https://www.wearedevelopers.com/videos/1176-startup-presentation-storx-future-of-cloud-storage) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)