Site Reliability Engineer (SRE) Splunk Enterprise Administrator

Nabout Leidos
United States
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$107,900.0 - $195,050.0
Working hours
Regular working hours

Tech stack

Agile Methodology User Authentication Automation of Tests Bash Shell Cloud Computing Configuration Management Cyber Security Continuous Integration Data Transport Utility Software Debugging DevOps Python (Programming Language)
+31 more
Parsing Windows PowerShell Red Hat Enterprise Linux Reliability Engineering Ansible Runbook Scaled Agile Framework Security Information and Event Management Syslog Systems Integration Software Vulnerability Management Scripting Computer Network Operations Cloud Platform System Performance Testing Data Ingestion Mttr Software Troubleshooting Reliability of Systems HybridCloud Indexer Information Technology Data Management Hardware Infrastructure Cloud Integration Restful APIs Terraform Splunk Network Server Data Pipelines Jenkins

Job description

Leidos is seeking a Site Reliability Engineer (SRE) Splunk Enterprise Administrator focused on Cyber support the largest IT services program for the Navy. Under the Service Management, Integration, and Transport (SMIT) program, the Leidos team delivers the core backbone of the Navy-Marine Corps Intranet, including cybersecurity services, network operations, service desk, and data transport. Leidos supports the Navy in unifying its shore-based networks and data management to improve capability and service while also saving significant dollars by focusing efforts under one enterprise network.\n \n As part of the SRE organization you will develop and execute tests focused on system resilience, performance underload, and failure scenarios. You will also work in tandem with other Site Reliability Engineers (SREs) and development teams to create automated testing frameworks that simulate real-world conditions that validate system behavior under normal and stress conditions, ensuring our services are resilient and meet established service level objectives (SLOs) The SRE will support the operations and maintenance of the enterprise network. Your work will contribute to the development of robust and scalable services that operate reliably in production.\n \n The Splunk Enterprise Administrator supports mission-critical cybersecurity operations by administering and maintaining distributed Splunk Enterprise platforms across hybrid cloud and on-premises environments. The role is responsible for daily platform operations, performance and ingestion monitoring, data onboarding support, incident troubleshooting, security compliance, documentation, and modernization activities in an Agile/DevOps operating culture.\n \n Key Metrics of Success for the Team: \n \u2022 Improved system reliability, as measured by adherence to Service Level Objectives (SLOs) and reduced Mean Time to Recovery (MTTR). \n \u2022 Comprehensive and regularly updated automated test coverage for all critical systems and infrastructure components. \n \u2022 Timely identification and resolution of performance bottlenecks and failure points. \n \u2022 Integration of automated testing into the CI/CD pipeline, ensuring continuous reliability validation. \n \u2022 Increased scalability and performance of systems under high load due to effective performance testing.\n \n \nWhat You’ll Get to Do:\n \u2022 Administer, maintain, and perform daily Operations and Maintenance (O&M) for distributed Splunk Enterprise environments, including Search Heads, Indexers, Heavy Forwarders, Intermediate Forwarders, Universal Forwarders, and Deployment Servers across hybrid cloud and on-premises infrastructure.\n \u2022 Monitor platform health, data ingestion, indexing throughput, search performance, and retention utilization; troubleshoot data onboarding, parsing, field extraction, forwarding, indexing, and ingestion issues.\n \u2022 Support Splunk Cloud integrations and associated hybrid operational activities.\n \u2022 Maintain Splunk applications, dashboards, alerts, saved searches, and knowledge objects.\n \u2022 Support patching, vulnerability remediation, Security Technical Implementation Guide (STIG) compliance, and system hardening activities.\n \u2022 Support platform upgrades, migrations, infrastructure modernization, technology refresh, and automation initiatives using scripting and configuration-management tools.\n \u2022 Participate in incident response, outage troubleshooting, problem resolution, and root cause analysis.\n \u2022 Maintain operational documentation, architecture diagrams, runbooks, and standard operating procedures (SOPs), and coordinate with cybersecurity, network, server, engineering, and customer teams during operational and modernization activities.\n \u2022 Participate in after-hours support and an on-call rotation as required.\n \n

Requirements

u2022 Requires BS degree and 5-10 years of prior relevant experience or Master’s with 4-8 years of prior relevant experience.\n \u2022 U.S. Citizen and posses an active Secret Security Clearance.\n \u2022 Minimum of DoD 8570.01 IAT Level II Certification required.\n \u2022 Experience designing, developing, and maintaining operational dashboards, visualizations, and executive reporting in Splunk Enterprise and Splunk Cloud, including IT Service Intelligence (ITSI), Service Analyzer, Glass Tables, and KPI-driven service health monitoring. \n \u2022 Experience with automated script design, coding, debugging, and maintenance skills (using bash, python, etc.) preferred. \n \u2022 Ability to work onsite at Norfolk Naval Station Monday through Friday day shift. \n \u2022 Must have a vendor certification e.g., Splunk Enterprise Certified Admin, Splunk Cloud Certified Admin, Scaled Agile Framework (SaFe).\n \u2022 Strong communication, analytical, and problem-solving skills, ability to work in a team environment.\n \u2022 Minimum three years of experience administering Splunk Enterprise v9 and supporting distributed Splunk architectures in hybrid cloud and on-premises production environments.\n \u2022 Strong working knowledge of Splunk data ingestion pipelines, indexing, parsing, forwarding, search optimization, and retention management.\n \u2022 Experience administering and troubleshooting Red Hat Enterprise Linux (RHEL) 8 and/or RHEL 9 servers.\n \u2022 Working knowledge of cybersecurity monitoring and Security Information and Event Management (SIEM) operations, with demonstrated experience troubleshooting operational incidents in complex enterprise environments.\n \u2022 Working knowledge of TCP/UDP networking, SSL certificates, Syslog, REST APIs, and authentication integrations.\n \u2022 Experience using at least one scripting or automation language: Python, Bash, or PowerShell.\n \u2022 Strong troubleshooting, analytical, communication, and documentation skills, with the ability to work independently, manage competing priorities, collaborate across technical teams, and maintain a customer-focused approach in a high-tempo production environment.\n, u2022 Splunk Core Certified Power User and/or Splunk Enterprise Certified Admin certification. Experience supporting Splunk Cloud environments or integrations and both Splunk Enterprise v9 and v10.\n \u2022 Experience working in classified government or Department of Defense environments.\n \u2022 Familiarity with STIGs, the Risk Management Framework (RMF), vulnerability management, and compliance frameworks.\n \u2022 Familiarity with RHEL 10, DevOps practices, and Infrastructure-as-Code concepts.\n \u2022 Experience with automation or configuration-management tools such as Ansible, Jenkins, Chef, or Terraform.\n

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.military.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

46 sec

Automating telemetry collection through robust Telegraf deployment

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:30 min

Discovering and instrumenting services using systemd process enumeration

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

Videos

See all

Related articles

See all