Site Reliability Engineer

Hire IT People
Bellevue, WA, United States
about 2 months ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Web Services Monitoring of Systems Mttr Containerization Kubernetes Docker Servicenow

Job description

  • Certified on one or more observability tools like Splunk. AppDynamics, Grafana, Dynatrace etc., * Production support activities including proactive identification of issues leveraging observability tools with the aim of reducing MTTD and MTTR
  • Coordinate all activities required to lead incident triage in compliance with SLAs and OLAs. Corelating inputs from various dashboards & tools to drive resolution.
  • Flexibility to work in 24 X 7 environment

Requirements

  • SRE Mindset in Production support: Proactive issue identification using observability tools. Skills in using different monitoring & observability tools to track system performance
  • Incident commander: Ability to diagnose complex issues and actively drive incident calls working with technical, product SMEs, and Tier 2 SREs.
  • Communication: Excellent communicator who could interact with Director/Sr. Director and above.

Technical expertise

  • Splunk (including Splunk APM and Splunk O11y), AppDynamics, Grafana, RedMetrics, 1000Eyes
  • Knowledge of VMs, Load balancers, Firewalls, API Gateways, DB, Network, Linux / Unix
  • Knowledge of Containerization, Docker, Kubernetes, AWS, PCF, GCP
  • ServiceNow (including AIOps, tools for Self-Heal and automated playbooks)
  • APM, NMON, Wireshark usage and analysis
  • Experience in UEM and synthetic monitoring tools

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on hireitpeople.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · WWC 2022

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all