Site Reliability Engineer

Harvey Nash
United States
8 days ago
Apply on careers.harveynashusa.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$104,000.0 - $112,320.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Cloud Database System Configuration Linux Instant Messaging Technology Nagios Reliability Engineering Prometheus Runbook Backup and Restore Cloud Platform System Grafana
+6 more
Amazon Virtual Private Cloud (VPC) Kubernetes Information Technology Terraform Software Version Control User Administration

Job description

We are seeking a highly motivated Site Reliability Engineer (SRE) with a strong operational focus to join our growing team. In this role, you will play a vital role in ensuring the smooth operation and performance of our critical infrastructure and services. You’ll work cross-functionally to create alignment and deliver results alongside builders who have helped to shape the success of companies such as Google, Okta, AWS, Snowflake.

What you will do in this role:

  • Deploy software for Cloud Prem and SAAS customers.
  • Respond to and diagnose system incidents in a timely and efficient manner, minimizing downtime and impact on users.
  • Collaborate with other engineers to establish root causes and implement effective resolutions.
  • Continuously improve incident response processes and documentation for future occurrences.
  • Proactively monitor and maintain the health and performance of our infrastructure and services.
  • Perform routine administrative tasks such as system configuration, user management, and data backups.
  • Identify and implement operational improvements to ensure ongoing system reliability and efficiency.
  • Develop and implement scripts and automated solutions to streamline operational tasks and reduce manual workload.
  • Participate in the on-call rotation to address critical incidents outside of regular business hours.
  • Ensure effective handoff between on-call engineers and document post-incident information for future reference.
  • Document processes for support and create, maintain and execute run-books for identified situations
  • Provide tier 2/3 technical support to customers experiencing platform issues or requiring advanced troubleshooting
  • Work directly with customer technical teams to resolve complex deployment, configuration, and integration challenges
  • Conduct technical onboarding sessions and provide guidance on best practices for customer implementations
  • Collaborate with customer success teams to ensure smooth customer experiences and rapid issue resolution
  • Create and maintain customer-facing technical documentation, troubleshooting guides, and knowledge base articles
  • Escalate customer feedback and feature requests to product and engineering teams
  • Participate in customer calls and technical discussions to provide expert-level platform guidance
  • Track and analyze customer support metrics to identify trends and areas for improvement

Requirements

  • Education:BS degree in Computer Science or related field
  • Experience:3+ years of experience in Site Reliability Engineering
  • 2+ years experience working with cloud platform and cloud automation tools especially in AWS
  • Strong experience with Kubernetes, Helm, Linux, AWS networking(VPC) and Terraform
  • Experience with the GitOps model for deployment
  • Familiarity with distributed version control
  • Experience with monitoring and alerting tools (e.g., Prometheus, Grafana).
  • Excellent communication skills with ability to explain technical concepts to both technical and non-technical audiences
  • Strong customer service orientation with patience and empathy when working with frustrated customers

Benefits & conditions

Medical, dental, and vision coverage 401(k) retirement plan Voluntary benefits and insurance options Referral bonus opportunities Pre-tax commuter benefits

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on careers.harveynashusa.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

3:37 min

Why differing legacy workflows complicate monitoring tool migrations

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

Videos

See all

Related articles

See all