Infrastructure Support Engineer / Site Reliability Engineer (SRE) / Cloud Operations Engineer

Wrymark, Inc.
San Jose, CA, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Working hours
Regular working hours
Job source

Tech stack

Cloud Computing Domain Name System (DNS) Monitoring of Systems Routing Prometheus TCP/IP Virtual Machines Datadog Load Balancing Grafana Reliability of Systems Firewalls (Computer Science)
+3 more
Splunk New Relic (SaaS) Dynatrace

Job description

  • Monitor production environments.
  • Handle incidents, alerts, and outages.
  • Perform root cause analysis (RCA).
  • Troubleshoot infrastructure issues.
  • Coordinate with customers and internal teams during critical incidents.
  • Ensure system reliability and availability.

Requirements

  • Overall 8+ years of experience

  • Core platform Engineering

  • Incidents, Alerts

  • SRE Mindset

  • Customer handling skills

  • Infra/Cloud basic concepts

  • Network

  • Storage, 1. SRE Mindset
  • Understanding of reliability, availability, and performance.
  • Focus on automation and reducing manual efforts.
  • Experience with incident management and problem management.
  1. Incident & Alert Management * Handling Sev1, Sev2, and Sev3 incidents. * Experience with monitoring tools such as:
  • Datadog
  • New Relic
  • Dynatrace
  • Splunk
  • Grafana
  • Prometheus

  • Performing RCA and post-incident reviews.
  1. Infrastructure Fundamentals

Strong understanding of:

  • Compute: Virtual Machines, CPU, Memory utilization
  • Storage: SAN, NAS, Disk management, Storage troubleshooting
  • Network: TCP/IP, DNS, Load Balancers, Firewalls, Routing basics
  1. Cloud Basics

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · WWC 2022

Videos

See all

Related articles

See all