Infrastructure Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Build resilient, scalable, fault-tolerant infrastructure for WRITER’s high-traffic enterprise generative AI platformMove between SRE, DevOps, Infrastructure, and Platform initiatives as priorities shiftAutomate operational tasks and infrastructure management with Python or GoDesign and operate infrastructure across AWS, GCP, and AzureWork with Kubernetes, Helm, Terraform, and cloud and AI toolingUse AI agents to investigate incidents, draft Terraform and Helm changes, write runbooks, scaffold tooling, and review pull requestsEncode recurring infrastructure tasks as reusable internal skills for human and agent teammatesLead incident response, post-mortems, and root-cause analysesOwn reliability, performance, and efficiency of core services end-to-endDefine and uphold SLOs and error budgets and carry the on-call pagerBalance immediate reliability work with long-term platform, observability, cost, and reliability investmentsCollaborate with product, security, and engineering peers on system
Requirements
design from conception through launchReport to the director of engineeringRequirements5+ years of experience in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product companyExperience running containerisation in productionExperience with Helm and Terraform or Pulumi on at least one major cloud (AWS preferred)Good proficiency in Python or Go for automation and toolingDaily workflow already includes agentic tooling such as Claude Code, Droid, Codex, or internal skills; this is a hard requirementDemonstrated ability to challenge the status quo, identify systemic weaknesses, and propose innovative solutions to complex reliability problemsAbility to reason from constraints and failure modes and articulate tradeoffs in business termsAbility to make reversible decisions, write rollback plans, work with monitoring and logging stacks, and stress systems safelyExcellent communication, collaboration, and problem-solving skillsStrong ownership and accountability for mission-critical systemsAt least one end-to-end 0-to-1 infrastructure build with an attached outcome metricWillingness and ability to work in person in the office 3 days per weekLegally authorized to work in the country where the job is locatedCore CompetenciesDemonstrates expertise in building and operating resilient, scalable infrastructure for high-traffic enterprise platforms, with a strong focus on automation using Python or Go. Proven ability to lead incident response and uphold reliability standards while collaborating effectively with cross-functional teams.Highest-signal resume keywordsInfrastructure EngineeringDevOps PracticesAWS, GCP, AzurePython or Go AutomationHelm and TerraformHard SkillsInfrastructure EngineeringDevOpsAutomationContainerizationIncident ResponseRoot-Cause AnalysisSLO DefinitionMonitoring and LoggingSystem DesignHigh-Availability SystemsSoft SkillsExcellent CommunicationCollaborationProblem-SolvingOwnershipAccountabilityIndustry KeywordsGenerative AIHigh-Traffic PlatformsProduction SystemsReliability EngineeringHigh-Growth Product CompanyTools & TechnologiesKubernetesHelmTerraformPulumiCloud ToolingAgentic Tooling #J-18808-Ljbffr
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Dev Digest 120 - Apple and peers
Highest Paying Tech Companies for Developers
Dev Digest 121 - AI goes offline