Site Reliability Engineer

Vital Tech Solutions
United States
3 days ago
Apply on www.thejobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Amazon Web Services Amazon S3 HP Systems Insight Manager Python (Programming Language) Linux Commands Reliability Engineering Ansible Data Streaming Cloud Platform System Delivery Pipeline Apache Spark Reliability of Systems
+6 more
Git Information Technology Terraform Software Version Control Docker Databricks

Job description

This role is focused on the reliable, secure, and efficient operation of critical production infrastructure. The Senior Site Reliability Engineer will be responsible for maintaining production environments, troubleshooting incidents, supporting deployments, improving monitoring and diagnostics, and driving greater uptime and operational stability., * Maintain, monitor, and troubleshoot production environments to support system uptime, reliability, and performance.

  • Manage and operate infrastructure using Terraform, Ansible, and Docker.
  • Support and maintain CI/CD pipelines and automate operational workflows using Git and related tooling.
  • Ensure the ongoing reliability of systems operating within AWS environments, including EKS, S3, and EMR.
  • Support operational use of technologies such as Spark, JupyterHub, and Hue.
  • Diagnose and resolve infrastructure and application issues.
  • Conduct root-cause analysis and drive long-term resolution of recurring incidents.
  • Implement, refine, and maintain infrastructure and application monitoring, alerting, and diagnostics.
  • Support deployment activities, maintenance windows, and production changes.
  • Optimize data flows and storage integrations.
  • Collaborate with engineering, product, and client stakeholders to communicate issues, coordinate maintenance, and support operational priorities.
  • Contribute to continuous improvement of operational processes, platform documentation, and reliability best practices.

Requirements

  • Active Secret security clearance or higher is required
  • Strong experience supporting and maintaining production infrastructure.
  • Hands-on Python experience within operational, infrastructure, or support environments.
  • Professional experience with Terraform, Ansible, and Docker.
  • Experience supporting CI/CD pipelines and deployment workflows.
  • Strong proficiency with Git and version-control practices.
  • Strong Linux command-line and systems operations experience.
  • Experience monitoring, diagnosing, and troubleshooting production systems.
  • Strong incident-response and root-cause analysis capabilities.
  • Ability to anticipate and resolve complex operational issues.
  • Strong communication skills and the ability to work effectively in a collaborative, client-facing environment.
  • Ability to learn and adapt to new technologies quickly.
  • Ability to work East Coast business hours.

Preferred Qualifications

  • Experience operating infrastructure within AWS or another major cloud platform.
  • Hands-on experience with AWS services including EKS, S3, and EMR.
  • Familiarity with Spark, JupyterHub, and Hue in an operational environment.
  • Databricks experience.
  • Experience supporting federal government, regulated, or other security-sensitive environments.
  • Experience working directly with external clients or government stakeholders.
  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline.

Full Time

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.thejobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all