Infrastructure Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
- Designing, building, and maintaining highly available, scalable infrastructure to support intensive AI/ML workloads and real-time model deployments.
- Implementing robust monitoring, alerting, and observability systems to ensure system health, performance, and uptime across cloud and on-prem environments.
- Debugging, optimizing, and automating infrastructure for fast iteration and rapid deployment cycles, focusing on both reliability and developer velocity.
- Proactively identifying, investigating, and resolving incidents to minimize downtime and maintain world-class service levels for enterprise customers.
- Collaborating closely with engineers, ML specialists, and founders to shape product, infrastructure, and security strategies.
Requirements
- Have 5+ years of hands-on experience in building or supporting production-grade infrastructure and reliability processes for high-throughput systems.
- Are comfortable with Python or similar languages, and exceptional at working across cloud platforms, container orchestration (e.g., Kubernetes), networking, and storage technologies.
- Build your own tools on the fly to diagnose, experiment, and address reliability problems-whether it’s an internal dashboard or an automated remediation workflow.
- Bring a quantitative, hands-on approach to system operations, automation, and continuous improvement., * Have prior experience founding a company or building products/infrastructure in early-stage, high-growth environments.
- Are excited about automating incident management processes with LLMs/AI.
- Are driven, ambitious, and deeply care about both technical excellence and collaborative problem-solving.
- Keep up with the latest trends in cloud, observability, and SRE best practices.
- Are passionate about open-source and have contributed tools or automation to reliability communities.
- Have built or optimized monitoring, incident response, or high-performance computing systems for demanding AI/ML, fintech, or enterprise clients.
Benefits & conditions
- Insurance: Generous health insurance covering medical, dental, and vision.
- Health and Wellness Budget: We provide up to $150/mo reimbursement for health and wellness spending, such as gym memberships, fitness classes, or similar.
- Parental Leave: Work with us to build a leave schedule that works for you and your family
About the company
Reducto is the agentic document platform for leading AI teams who demand enterprise performance at scale. We provide a comprehensive toolkit for working with documents the way a human would, combining custom in-house and leading frontier models to power efficient and accurate document workflows.
We’ve grown rapidly, increasing revenue 8x year over year and partnering with hundreds of companies, from leading AI teams like Harvey, Vanta, and Scale, to enterprise customers across FAANG and top trading firms.
Reducto has raised over $100M from world-class investors including a16z, Benchmark, and First Round Capital.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
Dev Digest 120 - Apple and peers
Navigating the AI Shift
Is Software Engineering Over-Saturated?