> Markdown version of [/jobs/ext/1914192-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/1914192-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Infrastructure engineer - **Company:** Ai, Inc - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Encodings, Cyber Security, Software Debugging, DevOps, Fault Tolerance, Python (Programming Language), Reliability Engineering, Prometheus, Software Engineering, Data Logging, Pulumi, Scripting, Google Cloud, Feature Engineering, System Availability, Grafana, Reliability of Systems, Kubernetes, Production Code, Terraform, Golang - **Published:** August 4, 2026 - **Apply:** https://www.careerbuilder.com/job-details/infrastructure-engineer-san-francisco-ca--5e6389ba-571a-487b-9431-6cfd5a11f221 ## About the Role * Track record. 5+ years of experience in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product company. * Breadth. Experience running containerisation in production (a real cluster, not a lab), with experience in Helm and Terraform or Pulumi on at least one major cloud (AWS preferred), plus good proficiency in Python or Go for automation and tooling. * AI in workflow. AI is part of how you ship, not a thing youve read about - agentic tooling (Claude Code, Droid, Codex, internal skills) is in your daily loop, youve built or adopted AI-assisted workflows others now use, and you have strong opinions on where its unreliable. This is a hard requirement, not a bonus. Candidates whose actual daily workflow does not already include AI tooling will not be advanced. * First-principles + decision-making. Demonstrated ability to Challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions to complex reliability problems - reason from constraints and failure modes (not analogy or vendor defaults), name the tradeoff in business terms (reliability vs. velocity, cost vs. blast radius, standardisation vs. one-off), and reject the "best practices" answer when it doesnt fit the problem. * Reversibility & blast-radius. Make reversible calls by default - write the rollback before you touch production, work fluently with monitoring and logging stacks (Prometheus, Grafana, ELK or equivalent), and stress the system in safe places so it comes back stronger. Non-technical * Cross-functional collaboration. Excellent communication, collaboration, and problem-solving skills, with a talent for building strong relationships and Connecting with cross-functional teams - surface non-goals before anyone asks, and partner with product, security, and platform peers as one delivery surface. * Autonomy & end-to-end ownership. A strong sense of ownership and accountability, eager to Own mission-critical systems and drive them toward peak performance and unparalleled reliability. At least one 0-to-1 infrastructure build you owned end-to-end, with the outcome metric attached. Bonus if you have * Software-engineering depth. A software-engineering background, not only config and scripting - youve designed, built, and shipped non-trivial production code (services, libraries, internal frameworks) in Python, Go, or a comparable language, you can read and modify the codebases your infrastructure runs, and you move between infra automation and feature engineering without changing brains., Amazon Web Services (AWS), Artificial Intelligence (AI), Artificial Intelligence (AI) Agents, Automation, Best Practices, Budgeting, Cloud Computing, Communication Skills, Cross-Functional, DevOps, Failure Analysis, Family Planning, GCP (Good Clinical Practices), Go Programming Language (Golang), High Availability, Hubs, Incident Response, Machine Tool, Metrics, Microsoft Windows Azure, Mine Budget/Costs, Motorola Droid, On Call, Pager, Problem Solving Skills, Product Planning, Production Systems, Python Programming/Scripting Language, Root Cause Analysis, Scaffolding, Scripting (Scripting Languages), Security Infrastructure, Software Engineering, Systems Reliability, Writing Skills ## Description At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, youll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isnt just about keeping the lights on; its about pushing the boundaries of whats possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. Youll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need. This is a hybrid position, based out of our New York City, San Francisco, Seattle, or London hubs. Youll report to our director of engineering. ️ What youll do Technical * Breadth across disciplines. Bring deep focus to one problem at a time, with the breadth to move between SRE, DevOps, Infrastructure, and Platform work over a quarter or two as the leverage shifts. This is not a thrash-every-week role - most of the time youre heads-down on one substantial initiative (the on-call posture, the release pipeline, the multi-region Terraform layout, the internal platform surface). Cross-layer fluency is what lets you pick the right next initiative; it isnt a weekly context-switch. * Simplicity / via negativa. Challenge the status quo and remove toil before adding features - automate operational tasks and infrastructure management with Python or Go, reject tools that dont fit the problem, and treat manual on-call work as a defect to be designed out, not a status quo to be staffed up. * Breadth across the stack. Design scalable, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure, working fluently across Kubernetes, Helm, Terraform, and the supporting cloud and AI tooling that backs WRITERs high-traffic platform. * AI in workflow. Run agents in your daily loop - Claude Code, Droid, Codex, internal skills - to investigate incidents, draft Terraform / Helm changes, write runbooks, scaffold tooling, and review PRs. Build the agentic setup as a collective surface: humans and digital teammates working as one team, with shared skills, shared context, and shared on-call workflows. Encode recurring infra tasks as internal skills any teammate (human or agent) can pick up and run, so the teams throughput compounds - not just your own. * Debugging fluency. Lead incident response, post-mortems, and root-cause analyses - trace failures to the underlying problem (never the symptom), apply the learning back into the architecture, and prevent the same incident from happening twice. Non-technical * End-to-end ownership. Own the reliability, performance, and efficiency of WRITERs core services end-to-end - define and uphold the SLOs and error budgets, carry the on-call pager, and stand behind the outcome metric, not just the system you shipped. * Strategic vs. tactical balance. Balance this weeks critical work with the 6-12-month platform direction - ship the on-call-driving fix today while shaping the multi-year observability, cost, and reliability investments that move WRITERs enterprise customers. * Cross-functional collaboration. Operate at the seams with product, security, and engineering peers - provide expert guidance on system design for reliability, performance, and scalability from conception through launch, Connect the infra agenda to product and revenue context, and disagree with evidence, not volume. ## Related Videos - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai)