> Markdown version of [/jobs/ext/3395150-senior-software-engineer-infrastructure](https://www.wearedevelopers.com/jobs/ext/3395150-senior-software-engineer-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SENIOR SOFTWARE ENGINEER, INFRASTRUCTURE - **Company:** KADIR.AI LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $100,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Application Programming Interfaces (APIs), Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Application Release Automation, Cloud Computing, Databases, Computer Engineering, Continuous Integration, Software Debugging, Fault Tolerance, Identity and Access Management, Python (Programming Language), Key Management, Software Engineering, TypeScript, Data Processing, Autoscaling, System Availability, Backend, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Terraform, Docker, Golang - **Published:** September 22, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=59f4cb03e462dfed ## About the Role This is not a traditional infrastructure-only position. The successful candidate will combine deep production-infrastructure ownership with strong backend and systems-engineering judgment across services, APIs, databases, queues, asynchronous workloads, and production failure modes. You will own the reliability, scalability, performance, cost, and developer experience of the company's core infrastructure and backend systems., * Experience owning production cloud infrastructure for a high-availability, user-facing platform * Deep experience with AWS and containerized environments * Strong hands-on experience with Kubernetes or Amazon EKS * Experience using Terraform or similar infrastructure-as-code tools * Strong understanding of Docker and container orchestration * Backend-engineering experience involving services, APIs, and databases * Experience designing or operating queues and asynchronous systems * Strong understanding of scaling limits, bottlenecks, and production failure modes * Experience with CI/CD, release automation, and environment management * Experience implementing observability through logs, metrics, traces, dashboards, and alerts * Experience with incident response, root-cause analysis, and operational runbooks * Ability to write maintainable automation or backend code * Strong systems-engineering judgment and debugging ability * Ability to maintain approximately 70-80% working-hour overlap with San Francisco or Singapore PREFERRED QUALIFICATIONS * Experience with data-heavy or artificial intelligence platforms * Experience supporting bursty or unpredictable workloads * Experience operating long-running jobs and distributed workers * Familiarity with sandboxed or isolated code-execution environments * Experience with high-throughput data-processing pipelines * Experience optimizing cloud infrastructure costs * Experience designing systems for backpressure, retries, cleanup, and safe rollback * Familiarity with Helm, CodeBuild, ECR, S3, IAM, and AWS networking * Experience working at an early-stage or rapidly scaling technology company * Experience collaborating with distributed teams across multiple time zones * Strong backend development experience in Python, Go, TypeScript, or a comparable language, * Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, or a related technical discipline is preferred * Equivalent professional experience will be considered * Candidates should demonstrate substantial ownership of production infrastructure and backend systems * Technical judgment, system ownership, and experience resolving real production failures are more important than a specific number of years ## Description * Own production uptime, reliability, latency, and system performance * Improve infrastructure provisioning speed and developer productivity * Manage infrastructure costs and identify cloud-cost optimization opportunities * Lead incident response, root-cause analysis, and corrective actions * Build and maintain production infrastructure on AWS * Develop infrastructure using Terraform, Kubernetes/EKS, Helm, and Docker * Work with EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management * Design systems for capacity planning, autoscaling, queues, and backpressure * Improve cleanup processes, retry behavior, fault tolerance, and safe rollbacks * Design reliable services, APIs, databases, and asynchronous workflows * Define dashboards, alerts, logs, traces, and service-level objectives * Create runbooks, escalation procedures, and sustainable on-call workflows * Build and maintain CI/CD pipelines and release-automation systems * Improve environment creation, configuration, and management * Write maintainable code for automation, backend services, and internal tools * Partner with software, research, and operations teams to improve platform reliability * Diagnose complex production failures across application and infrastructure layers, * Engineers can provision, test, and release changes quickly and safely * Monitoring identifies problems before they significantly affect users * Incidents are resolved quickly and followed by meaningful preventive improvements * Infrastructure expenses remain visible, controlled, and appropriately optimized * Backend and infrastructure systems are maintainable, well-documented, and resilient ## Related Videos - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)