Site Reliability Engineer (SRE) Splunk Enterprise Administrator
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+31 more
Job description
Leidos is seeking a Site Reliability Engineer (SRE) Splunk Enterprise Administrator focused on Cyber support the largest IT services program for the Navy. Under the Service Management, Integration, and Transport (SMIT) program, the Leidos team delivers the core backbone of the Navy-Marine Corps Intranet, including cybersecurity services, network operations, service desk, and data transport. Leidos supports the Navy in unifying its shore-based networks and data management to improve capability and service while also saving significant dollars by focusing efforts under one enterprise network.\n \n As part of the SRE organization you will develop and execute tests focused on system resilience, performance underload, and failure scenarios. You will also work in tandem with other Site Reliability Engineers (SREs) and development teams to create automated testing frameworks that simulate real-world conditions that validate system behavior under normal and stress conditions, ensuring our services are resilient and meet established service level objectives (SLOs) The SRE will support the operations and maintenance of the enterprise network. Your work will contribute to the development of robust and scalable services that operate reliably in production.\n \n The Splunk Enterprise Administrator supports mission-critical cybersecurity operations by administering and maintaining distributed Splunk Enterprise platforms across hybrid cloud and on-premises environments. The role is responsible for daily platform operations, performance and ingestion monitoring, data onboarding support, incident troubleshooting, security compliance, documentation, and modernization activities in an Agile/DevOps operating culture.\n \n Key Metrics of Success for the Team: \n \u2022 Improved system reliability, as measured by adherence to Service Level Objectives (SLOs) and reduced Mean Time to Recovery (MTTR). \n \u2022 Comprehensive and regularly updated automated test coverage for all critical systems and infrastructure components. \n \u2022 Timely identification and resolution of performance bottlenecks and failure points. \n \u2022 Integration of automated testing into the CI/CD pipeline, ensuring continuous reliability validation. \n \u2022 Increased scalability and performance of systems under high load due to effective performance testing.\n \n \nWhat You’ll Get to Do:\n \u2022 Administer, maintain, and perform daily Operations and Maintenance (O&M) for distributed Splunk Enterprise environments, including Search Heads, Indexers, Heavy Forwarders, Intermediate Forwarders, Universal Forwarders, and Deployment Servers across hybrid cloud and on-premises infrastructure.\n \u2022 Monitor platform health, data ingestion, indexing throughput, search performance, and retention utilization; troubleshoot data onboarding, parsing, field extraction, forwarding, indexing, and ingestion issues.\n \u2022 Support Splunk Cloud integrations and associated hybrid operational activities.\n \u2022 Maintain Splunk applications, dashboards, alerts, saved searches, and knowledge objects.\n \u2022 Support patching, vulnerability remediation, Security Technical Implementation Guide (STIG) compliance, and system hardening activities.\n \u2022 Support platform upgrades, migrations, infrastructure modernization, technology refresh, and automation initiatives using scripting and configuration-management tools.\n \u2022 Participate in incident response, outage troubleshooting, problem resolution, and root cause analysis.\n \u2022 Maintain operational documentation, architecture diagrams, runbooks, and standard operating procedures (SOPs), and coordinate with cybersecurity, network, server, engineering, and customer teams during operational and modernization activities.\n \u2022 Participate in after-hours support and an on-call rotation as required.\n \n
Requirements
u2022 Requires BS degree and 5-10 years of prior relevant experience or Master’s with 4-8 years of prior relevant experience.\n \u2022 U.S. Citizen and posses an active Secret Security Clearance.\n \u2022 Minimum of DoD 8570.01 IAT Level II Certification required.\n \u2022 Experience designing, developing, and maintaining operational dashboards, visualizations, and executive reporting in Splunk Enterprise and Splunk Cloud, including IT Service Intelligence (ITSI), Service Analyzer, Glass Tables, and KPI-driven service health monitoring. \n \u2022 Experience with automated script design, coding, debugging, and maintenance skills (using bash, python, etc.) preferred. \n \u2022 Ability to work onsite at Norfolk Naval Station Monday through Friday day shift. \n \u2022 Must have a vendor certification e.g., Splunk Enterprise Certified Admin, Splunk Cloud Certified Admin, Scaled Agile Framework (SaFe).\n \u2022 Strong communication, analytical, and problem-solving skills, ability to work in a team environment.\n \u2022 Minimum three years of experience administering Splunk Enterprise v9 and supporting distributed Splunk architectures in hybrid cloud and on-premises production environments.\n \u2022 Strong working knowledge of Splunk data ingestion pipelines, indexing, parsing, forwarding, search optimization, and retention management.\n \u2022 Experience administering and troubleshooting Red Hat Enterprise Linux (RHEL) 8 and/or RHEL 9 servers.\n \u2022 Working knowledge of cybersecurity monitoring and Security Information and Event Management (SIEM) operations, with demonstrated experience troubleshooting operational incidents in complex enterprise environments.\n \u2022 Working knowledge of TCP/UDP networking, SSL certificates, Syslog, REST APIs, and authentication integrations.\n \u2022 Experience using at least one scripting or automation language: Python, Bash, or PowerShell.\n \u2022 Strong troubleshooting, analytical, communication, and documentation skills, with the ability to work independently, manage competing priorities, collaborate across technical teams, and maintain a customer-focused approach in a high-tempo production environment.\n, u2022 Splunk Core Certified Power User and/or Splunk Enterprise Certified Admin certification. Experience supporting Splunk Cloud environments or integrations and both Splunk Enterprise v9 and v10.\n \u2022 Experience working in classified government or Department of Defense environments.\n \u2022 Familiarity with STIGs, the Risk Management Framework (RMF), vulnerability management, and compliance frameworks.\n \u2022 Familiarity with RHEL 10, DevOps practices, and Infrastructure-as-Code concepts.\n \u2022 Experience with automation or configuration-management tools such as Ansible, Jenkins, Chef, or Terraform.\n
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on jobs.military.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
DevOps Engineer Salary [2023]
Highest Paying Tech Companies for Developers
What’s the Difference between a Junior, Mid, and Senior Developer?
Data Engineer Salary UK