> Markdown version of [/jobs/ext/3037849-software-engineer-devops-sre](https://www.wearedevelopers.com/jobs/ext/3037849-software-engineer-devops-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - DevOps/SRE - **Company:** Project44, LLC - **Location:** Chicago, IL, United States - **Experience:** Expert - **Salary:** $130,000.0 - $190,000.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Artificial Intelligence, Amazon Web Services, Bash Shell, Cloud Computing, Code Review, Continuous Integration, Software Debugging, Software Design Documents, DevOps, Distributed Systems, Domain Name System (DNS), Identity and Access Management, Subnetting, Python (Programming Language), Key Management, Role-Based Access Control, Reliability Engineering, Runbook, Data Streaming, Scripting, Google Cloud, Load Balancing, GitHub Copilot, Istio, Mttr, Amazon Virtual Private Cloud (VPC), Kubernetes, Infrastructure Automation Frameworks, Real Time Data, Apache Kafka, Linkerd (Service Mesh), Cloud Optimization, Terraform, Golang - **Published:** September 23, 2026 - **Apply:** https://startup.jobs/senior-software-engineer-devops-sre-project44-10028126 ## About the Role * Strong communication skills, with the ability to clearly explain infrastructure decisions through design docs, postmortems, and cross-functional discussions, and collaborate effectively with both technical and non-technical partners * A strong analytical approach to problem-solving, with comfort navigating ambiguity and making tradeoffs explicit-especially across reliability, velocity, and cost * A sense of ownership, curiosity, and resilience-comfortable voicing ideas, incorporating feedback, and iterating toward better solutions * A platform-minded builder who thinks about the engineers and systems they enable, not just the infrastructure they manage Qualifications * 5+ years of professional experience in platform engineering, infrastructure engineering, SRE, DevOps, or a closely related role Strong experience in: * Kubernetes administration and operations at scale (GKE strongly preferred), including cluster networking, autoscaling, workload isolation, and upgrade management * Kafka operations and architecture, including topic design, consumer group management, broker configuration, and integration with streaming workloads (Strimzi experience a plus) * Infrastructure-as-code using Terraform-module design, state management, CI-driven plan/apply workflows, and drift detection * Google Cloud Platform (GCP) across core services: GKE, GCS, Cloud NAT, VPC networking, IAM, Artifact Registry, and related managed services * Cloud observability tooling and practices-building meaningful dashboards, alerts, SLI/SLO frameworks, and runbooks regardless of the tool stack * Cloud networking fundamentals: VPCs, subnetting, private connectivity, DNS, load balancing, and security group/firewall design * Cloud cost management: identifying and executing on optimization opportunities, building spend visibility, and influencing engineering behavior around resource efficiency Desirable and helpful (but not necessary): * Experience with AWS services alongside a primary GCP environment * Scripting or automation in Python, Go, Bash, or similar for operational tooling * Experience with service mesh technologies (Istio, Linkerd) or advanced Kubernetes networking (Cilium, Calico) * Familiarity with CI/CD systems and GitOps patterns (ArgoCD, Flux, or similar) * Experience building or operating developer platform capabilities (internal developer platforms, golden paths, self-service provisioning) * Using AI-assisted development tools for infrastructure automation, runbook generation, or operational tooling (e.g., Claude, GitHub Copilot, or similar) Work Authorization: Candidates must be authorized to work in the US without current or future employer-sponsored visa support. In-office Commitment: Our office is where ideas spark, connections thrive, and innovation comes alive. We are looking for candidates who are enthusiastic and committed to joining our team on-site, in our beautiful headquarters 3 days a week. Together, we're building something extraordinary-learn, grow, and thrive in our fast-paced, transformative environment. ## Description We expect every project44 team member, regardless of role or function, to actively leverage AI in their day-to-day work. Whether you're building product, serving customers, managing people, or running operations, AI is a tool you're expected to use with intent, curiosity, and judgment. We don't expect everyone to be a data scientist. We do expect everyone to be an intelligent user of AI: able to identify where it adds value, direct it effectively, evaluate outputs critically, and govern it responsibly. We invest in our team's AI fluency because we believe it's a competitive advantage for every person at project44, not just our engineers If you're driven to solve meaningful problems, leverage AI to scale rapidly, drive impact daily, and be part of a high-performance team - we should talk. The Role project44 is looking for a Senior DevOps/SRE/Platform Engineer to join our Foundations Engineering team. You will work in a fast-paced Agile environment designing, building, and operating the foundational systems that power project44's global logistics network-ensuring our platform is reliable, observable, cost-efficient, and built to scale. This is a high-ownership, high-impact role at the intersection of cloud infrastructure, distributed systems, and developer enablement. Key Accountabilities: What you'll work on * Design, build, and operate core infrastructure across our primary GCP environment and secondary AWS footprint, leveraging managed services to deliver scalable, reliable, and cost-effective solutions * Own and evolve our Kubernetes (GKE) platform-cluster lifecycle management, node pool strategy, networking, RBAC, and workload reliability-ensuring developer teams can ship safely and quickly * Build and maintain Terraform-based infrastructure-as-code across cloud environments, enforcing standards for modularity, state management, and repeatable, auditable provisioning * Manage and improve our Kafka and Strimzi streaming infrastructure, supporting high-throughput event pipelines that are foundational to project44's real-time data capabilities * Lead and execute cost optimization initiatives-rightsizing compute, improving resource utilization, reducing waste, and creating visibility and accountability for cloud spend across engineering teams * Champion observability and reliability practices - SLOs/SLIs, incident response, on-call processes, and postmortems - to reduce MTTR and prevent recurrence. * Partner with application engineering teams to define and implement platform standards for deployment, scaling, service mesh, secrets management, and infrastructure security * Participate in a team on-call rotation, debugging and resolving infrastructure and platform incidents with urgency and rigor * Contribute to technical design discussions, architecture reviews, and documentation-helping set a high bar for quality, reliability, and operational excellence across the engineering organization * Mentor and support other engineers through code reviews, pairing, and technical guidance * Evaluate and introduce new tools, platforms, and automation to continuously modernize the infrastructure stack. * Drive cost-optimization initiatives across cloud spend without compromising reliability or performance. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Platform Engineering vs. DevOps Why not both?](https://www.wearedevelopers.com/videos/885-platform-engineering-vs-devops-why-not-both) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026)