Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+13 more
Job description
- Own and maintain the end-to-end CI/CD pipeline, including build automation, automated testing gates, artifact management, release processes, and production deployments.
- Lead software deployments into secure, access-controlled operational environments, including hands-on work at customer facilities.
- Serve as a primary technical resource during site deployments, coordinating with infrastructure, security, and technical teams to resolve issues and validate successful implementations.
- Design, maintain, and improve infrastructure supporting both development and production environments, with an emphasis on reliability, scalability, security, and operational efficiency.
- Drive infrastructure-as-code and automation practices to create consistent, reproducible, and maintainable environments while reducing manual operational effort.
- Establish deployment standards, operational procedures, and rollback strategies to ensure releases are documented, repeatable, and recoverable.
- Implement and maintain containerization, orchestration, monitoring, alerting, and observability capabilities across production environments.
- Define service reliability objectives and implement the monitoring and operational practices necessary to measure and maintain system performance.
- Implement security controls for infrastructure and deployment processes, including identity and access management, secrets management, audit logging, and access controls.
- Automate security, compliance, vulnerability, and operational checks within CI/CD pipelines and deployed environments.
- Partner with software architecture and engineering teams on infrastructure decisions involving reliability, scalability, cost, security, and deployment strategy.
- Coordinate technical requirements, deployment schedules, access needs, and other implementation activities with external stakeholders and site personnel.
- Develop and maintain infrastructure documentation, deployment procedures, operational runbooks, and technical standards to support continuity and knowledge transfer.
- Contribute to technical planning, estimation, and continuous improvement initiatives for infrastructure and deployment workstreams.
Requirements
Our client is seeking a Senior Reliability Engineer to own CI/CD, infrastructure, deployment, and operational reliability for a mission-critical software platform operating within secure environments. This role is ideal for an experienced DevOps, site reliability, or infrastructure engineer who is equally comfortable building automated cloud infrastructure and coordinating hands-on deployments at customer sites. The ideal candidate will bring strong technical ownership, operational discipline, and communication skills to ensure reliable, secure, and repeatable deployments., * Active Secret or TS/SCI security clearance.
- Bachelor’s degree or equivalent professional experience in Computer Science, Systems Engineering, Information Technology, or a related technical discipline.
- 5+ years of experience in DevOps, site reliability engineering, infrastructure engineering, or a related discipline involving production deployments.
- Demonstrated experience building, maintaining, and operating CI/CD pipelines from source code commit through production deployment.
- Strong AWS infrastructure experience, including production cloud environments, containerization with Docker, ECS, or EKS, and infrastructure-as-code technologies such as Terraform or CloudFormation.
- Hands-on experience with production monitoring, observability, alerting, troubleshooting, and reliability practices.
- Experience deploying software into external, customer, partner, or other environments outside of internally controlled infrastructure.
- Strong understanding of deployment automation, infrastructure security, access management, configuration management, and operational best practices.
- Excellent communication and coordination skills with the ability to work effectively with engineering teams, security personnel, customer stakeholders, and program leadership.
- Ability to travel and work on-site at customer facilities as required to support deployment and operational activities.
Preferred:
- Experience deploying or supporting technology within classified, air-gapped, SCIF, or similarly restricted environments.
- Experience supporting Department of Defense, military, federal government, or defense contractor programs.
- Experience with Kubernetes and GitOps technologies or workflows such as Argo CD or Flux.
- Familiarity with high-availability and zero-downtime deployment strategies.
- Experience implementing security automation, including policy-as-code, vulnerability scanning, compliance checks, or audit automation.
- Familiarity with security and compliance requirements associated with FedRAMP, ITAR, classified systems, or comparable controlled environments.
- AWS DevOps Engineer - Professional, Solutions Architect - Associate, or comparable cloud certification.
- Located in or willing to relocate to the Albuquerque, New Mexico area.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Is Software Engineering Over-Saturated?
What Are The Top Skills Required For Azure Developers?
Highest Paying Tech Companies for Developers