Site Reliability Engineer (SRE) - SecOps
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+14 more
Job description
We are seeking a highly motivated Systems Reliability Engineer (SRE) to lead the design and implementation of operational excellence across our multi-cloud environments. This role is central to ensuring the scalability, reliability, and performance of our products running in AWS, Azure, and GCP infrastructure. As the lead SRE, you will own the uptime, observability, and system resilience for our critical services. This includes driving architecture decisions, automation practices, and incident response strategies-working closely with the product owner(s), developer teams, and security operations teams.
What You’ll Do
- Design, implement, and own the infrastructure reliability strategy across AWS, Azure, and GCP
- Champion observability by developing and maintaining effective logging, monitoring, and alerting systems
- Define and enforce SLOs/SLAs for critical systems and services
- Lead efforts in performance tuning, system hardening, capacity planning, and disaster recovery
- Own the incident management lifecycle: from detection to postmortem and root cause analysis
- Automate deployment, scaling, and recovery workflows to reduce manual toil
- Contribute to infrastructure as code (Terraform, ARM templates, CloudFormation, etc.)
- Act as a mentor and technical leader to junior engineers and cross-functional partners
- Drive a culture of accountability, ownership, and continuous improvement
- Perform any other related duties as required or assigned., + The Security Builder: You don’t just consume security tools - you extend and improve them. You’re energized by the opportunity to make analysts faster and compliance more automated., We are a Defense-focused company supporting sensitive and cleared workforces. The Site Reliability Engineer (SRE) - SecOps will embrace our commitment to operational excellence, compliance rigor, and a world-class employee experience.
Requirements
- 5+ years of experience in SRE, DevOps, or infrastructure engineering roles
- Proven track record of operating large-scale systems in multi-cloud environments, with hands-on expertise in AWS and GCP
- Strong knowledge of cloud-native architecture, container orchestration with Kubernetes, and CI/CD pipelines
- Proficient in scripting (Python, Bash, etc.) and infrastructure automation tools (e.g., Terraform)
- Experience with monitoring and observability platforms (e.g., Prometheus, Grafana, Datadog, ELK)
- Excellent problem-solving skills with the ability to manage incidents and make sound decisions under pressure
- Clear communicator capable of translating technical concepts to mixed audiences and participating in customer discussions, + Quality Over Speed: You write code that lasts. You push for clean interfaces, good documentation, and tests - even in a fast-moving environment.
- Security-Minded Developer: You treat security as a first-class requirement, not an afterthought. You’re comfortable reading CVEs, threat models, and compliance controls.
- Cross-Functional Partner: You can work fluidly with security analysts, engineers, and compliance professionals, translating needs into reliable software., + Prolonged periods of sitting at a desk and working on a computer
- Must be able to lift up to 15 pounds at times
- May require occasional travel to office locations or client sites
- Ability to communicate effectively in written and verbal form
Benefits & conditions
Benefits for working with us! We are committed to supporting our employees both professionally and personally. Our robust benefits package is designed to promote your well-being, growth, and work-life balance
- Competitive Salary: Recognizing your hard work with attractive compensation and rewarding excellence.
- Health and Wellness Programs: Including medical, dental, & vision insurance options, along with mental health support & wellness initiatives.
- Retirement Planning: Secure your future with our flexible 401(k) plan and matching company contributions.
- Paid Time Off & Holidays: Generous PTO, sick leave, and holiday pay to help you recharge and enjoy life outside of work.
- Employee Assistance Program: Confidential resources for personal and professional support.
- Professional Development: Access to training, certifications, and continuing education to foster your career growth.
About the company
At Arkenstone Defense, we empower defense tech startups with the tools, infrastructure, and compliance solutions they need to become successful prime contractors. Our mission is to remove barriers and help innovators grow - from day one to becoming a trusted prime for the U.S. Government. We’re early, we’re lean, and we’re building something that actually matters. The people who do well here aren’t waiting to be told what to do; they see a gap and fill it.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Find a Developer Job: 12 Best Job Sites For Developers
Where To Find Software Engineering Jobs
Dev Digest 120 - Apple and peers