Elastic Site Reliability Engineer (SRE)
Zachary Piper
Boston, MA, United States
3 months ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$180,000.0 - $200,000.0
Working hours
Regular working hours
Job source
Tech stack
Amazon Elastic Compute Cloud
Unix
Cloud Computing
Cyber Security
Continuous Integration
DevOps
Distributed Systems
Elasticsearch
Linux System Administration
Uptime
Networking Basics
Performance Tuning
+15 more
Reliability Engineering
Logstash
Ansible
Security Information and Event Management
Data Logging
System Availability
Delivery Pipeline
SC Clearance
Falcon Platform
Kubernetes
Infrastructure Automation Frameworks
Kibana
Terraform
Splunk
Devsecops
Job description
- Operate, maintain, and optimize large-scale Elastic Stack environments supporting logging, search, observability, and telemetry operations.
- Ensure platform reliability, uptime, scalability, and performance across production mission systems.
- Manage Kubernetes-based Elastic deployments, including ECK operator environments.
- Develop and maintain automation for deployment workflows, monitoring, alerting, and incident response processes.
- Integrate Elastic infrastructure with SIEM and security tooling including Splunk, EDR platforms, and telemetry systems.
- Troubleshoot complex issues across distributed systems, infrastructure, and application environments.
- Implement and support observability frameworks including logging, metrics, tracing, and monitoring solutions.
- Support CI/CD pipelines and infrastructure-as-code initiatives within DevOps environments.
- Maintain operational runbooks, escalation procedures, and technical documentation.
- Participate in on-call support rotations and incident response activities., elastic sre, site reliability engineer, elastic stack, elasticsearch, kibana, logstash, beats, observability, telemetry, logging infrastructure, distributed systems, kubernetes, eck, elastic cloud on kubernetes, sre, devops, platform engineering, infrastructure engineering, production support, linux administration, unix systems, networking, monitoring, tracing, metrics, incident response, automation, ci/cd, infrastructure as code, terraform, ansible, cloud infrastructure, distributed logging, telemetry systems, siem, splunk, edr, crowdstrike, trellix, platform reliability, reliability engineering, scalability, uptime, performance tuning, root cause analysis, operational excellence, federal infrastructure, dod, govcloud, classified systems, mission systems, secret clearance, hanscom afb, langley afb, secure environments, production engineering, elastic observability, elastic security, sre engineer, kubernetes engineer, platform sre, enterprise infrastructure, cloud operations, mission critical systems, elastic engineer, telemetry engineer, security operations, devsecops, automation engineer, distributed architecture, operational support
Requirements
- 5+ years of experience supporting Site Reliability Engineering, DevOps, or infrastructure operations environments.
- Strong hands-on experience with Elastic Stack in enterprise production environments.
- Advanced Kubernetes experience, including ECK operator deployments.
- Strong Linux/Unix administration and networking fundamentals.
- Experience supporting observability, telemetry, logging, and monitoring platforms.
- Experience working within secure, classified, federal, or highly regulated environments.
- Ability to work onsite at Hanscom AFB (MA) or Langley AFB (VA).
- U.S. Citizenship with ability to obtain or maintain a Secret clearance.
Nice-to-Haves:
- Elastic certifications including Elastic Engineer, Security, or Observability.
- Experience with Terraform, Ansible, and CI/CD pipeline automation.
- Exposure to SIEM and EDR technologies including Splunk, CrowdStrike, or Trellix.
- Experience supporting GovCloud, DoD, or federal infrastructure environments.
- Prior experience supporting distributed logging or telemetry platforms.
Soft Skills:
- Strong incident response and operational troubleshooting mindset.
- Ability to remain calm and effective during production outages or high-pressure situations.
- Strong collaboration skills across security, infrastructure, DevOps, and operations teams.
- Excellent communication skills for escalation and operational coordination environments.
- Self-sufficient and capable of operating independently within classified environments.
Benefits & conditions
- Target compensation: $180,000 - $200,000 annually.
- Long-term federal engagement supporting mission-critical infrastructure initiatives.
- Opportunity to support advanced observability and security operations within classified environments.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on clearancejobs.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
over 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
about 2 years ago
EM
Eli McGarvie
React Developer Salary [2023]
over 3 years ago
EM
Eli McGarvie
Find a Developer Job: 12 Best Job Sites For Developers
over 3 years ago
CH
Chris Heilmann
Dev Digest 134 - Where pixels sing?
almost 2 years ago