Senior Site Reliability Engineer (Operations)

AGILETEK SOLUTION LLC
Reston, VA, United States
15 days ago
Apply on www.clearancejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services C Sharp (Programming Language) Cloud Computing Configuration Management Continuous Integration DevOps Python (Programming Language) Octopus Deploy Reliability Engineering Ruby Software Deployment Rust (Programming Language)
+7 more
System Availability Grafana Kubernetes Infrastructure Automation Frameworks Terraform Workday Programming Languages

Job description

We are looking for a highly motivated Site Reliability Engineer to join our growing Infrastructure and Platform Engineering team. You will play a critical role in operating, monitoring, automating, maintaining, and providing metrics and observability for our Kubernetes-based platform and enabling our engineering organization to deliver capabilities and incremental products quickly and reliably. You and the team will have the opportunity to work with cutting-edge technologies, solve complicated problems, and contribute to the foundation of our infrastructure. This role requires a strong technical background, a passion for automation, and a collaborative mindset.

This role involves collaborating with teams across multiple locations and time zones. You will need to be adept at communicating optimally and building positive relationships with colleagues in diverse locations. About the Role You have a growth mindset and will be part of a team promoting a diverse and inclusive environment where you and your workmates are happy, energized and engaged, and who are excited to come to work every day. Responsibilities:

  • Ensuring the Workday Kubernetes based platform is maintained, healthy, and ensures high availability for our customers through, infrastructure automation, CI/CD pipelines, reporting, incident handling and response, and observability tools.
  • Maintain the overall platform: maintain core platform components, ensuring high availability, scalability, and security.
  • Automate and optimize: Automate infrastructure provisioning, configuration management, and application deployments using tools like Terraform and Argo CD.
  • Troubleshooting and support: Provide support and solve for platform-related issues, working closely with development teams to resolve problems.
  • Security and compliance: Implement and maintain security standard methodologies for the platform, ensuring compliance with industry standards.
  • Documentation and knowledge sharing: Build and maintain comprehensive documentation for platform components and processes. Actively participate in knowledge sharing within the team.
  • Collaborate effectively with other engineers and development teams across multiple locations and time zones.
  • Stay up-to-date with the latest technologies and trends in the platform engineering space. This role will support one or more direct or indirect contracts with the U.S. Federal Government which, due to federal government security requirements, mandates that all Workday personnel working on the contracts be United States citizens (naturalized or native).

Requirements

  • 5+ years of hands-on experience working with large scale cloud infrastructure, automation, and overall DevOps methodologies
  • Bachelor’s degree in a computer related field or equivalent work experience Other Qualifications:

  • Infrastructure as code: Proficiency in infrastructure automation tools like Terraform.
  • CI/CD: Experience with building, maintaining, and consuming CI/CD pipelines and tools like Argo CD.
  • Problem-solving: Strong analytical and problem-solving skills.
  • Communication: Excellent communication and collaboration skills. Preferred:

  • Strong understanding of Kubernetes
  • Amazon Web Services proficiency working in a production environment
  • Proficiency in at least one programming language such as C#, Python, Ruby, Rust, or Go programming language proficiency
  • Experience with security auditing and compliance frameworks. Experience working in air gapped cloud regions

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:22 min

Eliminating HR bureaucracy and trusting employees

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

50 sec

Why developer happiness matters in web frameworks

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

5:06 min

Primary reasons for capability gaps in modern recruitment systems

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all