> Markdown version of [/jobs/ext/2732727-linux-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2732727-linux-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Linux Infrastructure Engineer - **Company:** IAG NORTH AMERICA, LLC - **Location:** Chicago, IL, United States - **Experience:** Expert - **Salary:** $140,000.0 - $180,000.0 - **Contract:** Permanent contract - **Skills:** Bash Shell, Configuration Management, Code Review, Protocol Stack, Dynamic Host Configuration Protocol, Software Debugging, Linux, File Systems, Domain Name System (DNS), Elasticsearch, HAProxy, Icinga, Python (Programming Language), Kernel-Based Virtual Machine, Nagios, Routing, Nginx, RabbitMQ, Redis, Ansible, Runbook, TCP/IP, Transmission Control Protocol (TCP), Virtual Local Area Networks, Virtualization Technology, Xen Servers, Load Balancing, CheckMK, Git, Containerization, Kubernetes, Bare Metal, Puppet, Vmware - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/senior-linux-infrastructure-engineer-tastylive-9890719 ## About the Role * 6+ years in a Linux systems, infrastructure, or SRE role. Expert-level Linux, specifically: * Performance analysis. You can go from a vague complaint to a named subsystem using the standard tooling - perf, strace, ss, iostat, bpftrace, or equivalents - and explain what the numbers mean. * Troubleshooting methodology. A disciplined, hypothesis-driven approach that works on a system you have never seen before. We care more about how you narrow the problem than about which commands you happen to know. * Systems fundamentals. Processes and signals, systemd, cgroups and namespaces, filesystems, and what actually happens when a host exhausts memory or file descriptors. * Declarative configuration management at production scale - Salt, Ansible, Chef, or Puppet. You have authored and maintained the code, not only run it. * Containerization and orchestration. Production experience with Kubernetes or Nomad, including the operational realities: scheduling, resource pressure, rollouts, and debugging a workload that will not start. * Scripting and automation. Strong Bash and working Python. You write code that other engineers can maintain. * Foundational networking. Comfortable across layers - VLANs and routing, DNS, DHCP, TCP behavior, and TLS - and able to determine whether a problem is the application, the host, or the network. * Proxy and load-balancing experience with Nginx or HAProxy. * Virtualization experience with at least one of VMware, Xen, or KVM. * Git and peer review as a daily habit, including the discipline to keep changes reviewable. * Excellent written communication. Runbooks, design proposals, and incident write-ups are a core part of this job, not an afterthought. * The ability to get productive quickly in areas where you do not yet have depth. Preferred Qualifications * Production ownership of Redis or RabbitMQ as a primary responsibility rather than a dependency. * Vault administration - policies, auth methods, and rotation at scale. * Elastic Stack operations at volume: index lifecycle, mapping decisions, and cluster tuning. * Experience in a regulated environment, or anywhere downtime has a direct revenue cost. * Colocation or bare-metal experience: hardware lifecycle, remote hands, and capacity planning against a finite footprint. ## Description We are looking for a senior Linux engineer who is equal parts operator and builder. You will own the systems layer of our production environment - from the metal to the configuration management code. This is a hands-on role: you will jump into incidents, tackle challenging engineering problems, design systems that scale, and evaluate new software and architectures. This is not a role where the abstractions hide the machine. You are expected to be curious about what is underneath and to be the person other engineers come to when a system is behaving in a way nobody can explain. What You'll Do * Own Linux performance. Diagnose and tune systems under real production load: CPU scheduling and NUMA placement, memory and page cache behavior, disk and filesystem I/O, and network stack tuning. * Lead troubleshooting on hard problems. Work incidents methodically - form a hypothesis, find the cheapest test that falsifies it, and narrow the search rather than changing five things at once. Write postmortems, organize follow-up tasks, and future-proof the environment so the same issue does not recur. * Write and maintain configuration management code. Build infrastructure declaratively with Salt, Ansible, Chef, or Puppet, treating that code with the same standards as application code: reviewed, tested, and version-controlled. * Run containerized workloads. Build, deploy, and operate services on Kubernetes or Nomad, including scheduling behavior, resource limits, health checking, and the failure modes that only appear under contention. * Automate in Bash and Python. Replace manual runbooks with tooling and streamline repeatable work. * Operate core network services. Troubleshoot TCP/IP with confidence and bring solid networking fundamentals. Manage DNS and DHCP as production services - zone management, resolver behavior, TTL strategy, scopes, reservations, and relay configuration. * Operate the traffic and data tier. Configure and troubleshoot Nginx and HAProxy (routing, TLS termination, health checks, connection handling) and support Redis and RabbitMQ in production. * Manage virtualization. Provision and maintain guests across VMware, Xen, or KVM, including capacity planning, host maintenance, and live migration. * Handle secrets properly. Use Vault for secret storage, dynamic credentials, policy, and rotation - and help move the organization off whatever it was doing before. * Own observability. Maintain log aggregation on the Elastic Stack and alerting through Nagios, CheckMK, or Icinga. Tune alerts toward signal; an alert nobody can act on is a bug. * Work through change management and code review. Everything moves through Git and pull requests. You will review other people's changes as seriously as you expect yours to be reviewed. ## Related Videos - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Geometry of Incidents: Connecting User Impact to Architecture](https://www.wearedevelopers.com/magazine/764-the-geometry-of-incidents-connecting-user-impact-to-architecture) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)