Site Reliability Engineer

Nabout Leidos
United States
21 days ago
Apply on jobs.military.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$87,100.0 - $157,450.0
Working hours
Regular working hours

Tech stack

Bash Shell Computer Networks Computer Engineering System Configuration Software Debugging Linux File Systems Elasticsearch Fault Tolerance Monitoring of Systems Python (Programming Language) Ansible
+6 more
Prometheus Scripting Grafana Information Technology Terraform Golang

Requirements

u2022 Enthusiasm for learning new stuff.\n \u2022 Bachelor’s degree in Computer Science, Computer Engineering, a related field, or amazing equivalent skills.\n \u2022 Strong experience administering and troubleshooting Linux systems.\n \u2022 Experience with Linux networking, storage, and file systems.\n \u2022 Experience automating system configuration and deployment.\n \u2022 Experience with Nix and/or NixOS in production, lab, or significant personal environments.\n \u2022 Python, Go, Bash, or similar scripting/programming experience of 2+ years.\n \u2022 Ability to debug complex systems methodically across multiple layers of the stack.\n \n, u2022 Deep experience with NixOS, including custom modules, overlays, flakes, packaging, and reproducible deployments.\n \u2022 Experience operating fleets of Linux systems.\n \u2022 Infrastructure-as-code experience with Nix, Terraform, Ansible, or similar tooling.\n \u2022 Experience designing highly available or fault-tolerant systems.\n \u2022 Monitoring and observability experience with tools such as Prometheus, Grafana, Loki, OpenTelemetry, Elasticsearch, or similar systems.\n

Benefits & conditions

Kudu Dynamics is a 100% employee-owned company, forged out of a decade of experience in computer network operations and staffed with talent who have built, overseen, and enhanced capabilities throughout the entire USG arsenal. Our team of hackers, engineers, makers, and shakers brings deep experience across research, development, deployment, and operations.\n \n Kudu Dynamics is uniquely qualified to anticipate tomorrow’s threats and build the next generation of capabilities.\n \n When you come in for your interview, you’ll see that Kudu is an amazing place to work, where you’ll be surrounded by experts who are ready to teach and learn. Our team has flexible work hours and work-from-home options. When we do work from the offices, we enjoy home-roasted coffee and award-winning workspaces.\n \n \nJob Description:\n \n Hello! We’re a small team in a fun company looking for someone to come in and help us build systems that stay reliable when things get complicated.\n \n We need a Site Reliability Engineer who has experience building, deploying, automating, and operating complex compute platforms. You’ll work across infrastructure, Linux systems, networking, distributed storage, observability, security, and deployment automation, with a particular emphasis on creating systems that are reproducible, maintainable, resilient, and easy to operate.\n \n A major part of this role will involve Nix and NixOS. We want someone who appreciates declarative systems, reproducible environments, infrastructure-as-code, and reducing configuration drift. You may be building NixOS-based servers, improving deployment pipelines, developing reusable Nix modules, debugging distributed systems, or making sure a platform can be rebuilt predictably from scratch.\n \n Critical components of the environment may include high-performance compute, distributed Linux file systems, network design, security-in-depth, ML/AI infrastructure, high-bandwidth data processing, cloud deployment, and fielded systems.\n \n You don’t need to be an expert at everything. Whatever part you bite off, though, will be yours to own. This is a great opportunity to apply deep systems expertise while learning adjacent areas.\n \n Our team needs your help making sure our customers are repeatedly provided with operational insights from a multi-domain data environment-and that the infrastructure producing those insights is dependable, observable, reproducible, and recoverable.\n \n \nThese Are the Things You\n’\nll Get to Do:\n \u2022 Own the reliability and operation of critical compute and data platforms.\n \u2022 Design, deploy, and maintain Linux infrastructure, including NixOS-based systems.\n \u2022 Build reproducible system configurations and deployment workflows using Nix and infrastructure-as-code.\n \u2022 Develop reusable NixOS modules, packages, flakes, and system configurations.\n \u2022 Improve platform resilience, fault tolerance, recoverability, and maintainability.\n \u2022 Build monitoring, logging, alerting, and observability capabilities that help us understand system behavior before customers notice problems.\n \u2022 Diagnose difficult failures across Linux, networking, storage, containers, and distributed systems.\n \u2022 Automate repetitive operational tasks and eliminate configuration drift.\n \u2022 Design and build security and maintainability components for deployed systems.\n \u2022 Help define operational standards, deployment practices, upgrade strategies, and disaster-recovery procedures.\n \u2022 Perform capacity planning and identify performance bottlenecks across compute, storage, and networking.\n \u2022 Own solutions spanning CNO, analytics, security, infrastructure, and field deployments.\n \u2022 Learn new skills and teach the rest of us what you know.\n \n

About the company

If you’re looking for comfort, keep scrolling. At Leidos, we outthink, outbuild, and outpace the status quo - because the mission demands it. We’re not hiring followers. We’re recruiting the ones who disrupt, provoke, and refuse to fail. Step 10 is ancient history. We’re already at step 30 - and moving faster than anyone else dares.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.military.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all