SENIOR SOFTWARE ENGINEER, INFRASTRUCTURE
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+16 more
Job description
- Own production uptime, reliability, latency, and system performance
- Improve infrastructure provisioning speed and developer productivity
- Manage infrastructure costs and identify cloud-cost optimization opportunities
- Lead incident response, root-cause analysis, and corrective actions
- Build and maintain production infrastructure on AWS
- Develop infrastructure using Terraform, Kubernetes/EKS, Helm, and Docker
- Work with EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management
- Design systems for capacity planning, autoscaling, queues, and backpressure
- Improve cleanup processes, retry behavior, fault tolerance, and safe rollbacks
- Design reliable services, APIs, databases, and asynchronous workflows
- Define dashboards, alerts, logs, traces, and service-level objectives
- Create runbooks, escalation procedures, and sustainable on-call workflows
- Build and maintain CI/CD pipelines and release-automation systems
- Improve environment creation, configuration, and management
- Write maintainable code for automation, backend services, and internal tools
- Partner with software, research, and operations teams to improve platform reliability
- Diagnose complex production failures across application and infrastructure layers, * Engineers can provision, test, and release changes quickly and safely
- Monitoring identifies problems before they significantly affect users
- Incidents are resolved quickly and followed by meaningful preventive improvements
- Infrastructure expenses remain visible, controlled, and appropriately optimized
- Backend and infrastructure systems are maintainable, well-documented, and resilient
Requirements
This is not a traditional infrastructure-only position. The successful candidate will combine deep production-infrastructure ownership with strong backend and systems-engineering judgment across services, APIs, databases, queues, asynchronous workloads, and production failure modes.
You will own the reliability, scalability, performance, cost, and developer experience of the company’s core infrastructure and backend systems., * Experience owning production cloud infrastructure for a high-availability, user-facing platform
- Deep experience with AWS and containerized environments
- Strong hands-on experience with Kubernetes or Amazon EKS
- Experience using Terraform or similar infrastructure-as-code tools
- Strong understanding of Docker and container orchestration
- Backend-engineering experience involving services, APIs, and databases
- Experience designing or operating queues and asynchronous systems
- Strong understanding of scaling limits, bottlenecks, and production failure modes
- Experience with CI/CD, release automation, and environment management
- Experience implementing observability through logs, metrics, traces, dashboards, and alerts
- Experience with incident response, root-cause analysis, and operational runbooks
- Ability to write maintainable automation or backend code
- Strong systems-engineering judgment and debugging ability
- Ability to maintain approximately 70-80% working-hour overlap with San Francisco or Singapore
PREFERRED QUALIFICATIONS
- Experience with data-heavy or artificial intelligence platforms
- Experience supporting bursty or unpredictable workloads
- Experience operating long-running jobs and distributed workers
- Familiarity with sandboxed or isolated code-execution environments
- Experience with high-throughput data-processing pipelines
- Experience optimizing cloud infrastructure costs
- Experience designing systems for backpressure, retries, cleanup, and safe rollback
- Familiarity with Helm, CodeBuild, ECR, S3, IAM, and AWS networking
- Experience working at an early-stage or rapidly scaling technology company
- Experience collaborating with distributed teams across multiple time zones
- Strong backend development experience in Python, Go, TypeScript, or a comparable language, * Bachelor’s degree in Computer Science, Software Engineering, Computer Engineering, or a related technical discipline is preferred
- Equivalent professional experience will be considered
- Candidates should demonstrate substantial ownership of production infrastructure and backend systems
- Technical judgment, system ownership, and experience resolving real production failures are more important than a specific number of years
Benefits & conditions
- Ownership of business-critical infrastructure and backend systems
- Opportunity to solve complex scaling, reliability, and systems-engineering problems
- Close collaboration with software engineers and AI research teams
- Significant influence over architecture, tooling, and developer experience
- Remote work with substantial San Francisco or Singapore time-zone overlap
- Competitive annual compensation of $100,000-$200,000
- Potential visa and relocation support for strong candidates
HIRING PROCESS
- Application and résumé review
- Initial screening discussion
- Two technical interviews
- Two-to-three-day structured work trial
- Final decision and offer
Kadir.ai is an equal opportunity recruiting partner. Qualified applicants will be considered without regard to race, color, religion, sex, gender identity, sexual orientation, national origin, age, disability, veteran status, or any other status protected by applicable law.
Pay: $100,000.00 - $200,000.00 per year
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Fully Remote Software Engineer Jobs
Is Software Engineering Over-Saturated?
What is Software Engineering in the Age of AI?