Infrastructure Support Engineer / SRE / Cloud Operations Engineer

LTD Global
Sunnyvale, CA, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Cloud Computing Computer Networks Domain Name System (DNS) Monitoring of Systems Identity and Access Management Reliability Engineering Prometheus TCP/IP Virtual Machines Datadog
+8 more
Google Cloud Load Balancing Grafana Firewalls (Computer Science) Information Technology Splunk New Relic (SaaS) Dynatrace

Job description

  • Production Operations & Monitoring: Monitor enterprise production environments using tools such as Datadog, Grafana, Prometheus, New Relic, or Splunk.
  • Incident & Outage Management: Lead triaging, escalation, and resolution for Sev1, Sev2, and Sev3 incidents; perform thorough Root Cause Analysis (RCA).
  • Infrastructure Troubleshooting: Identify and resolve operational issues across virtual machines, OS, storage (SAN/NAS), and networking components (DNS, Load Balancers, TCP/IP).
  • Cloud Operations: Support core cloud resources across AWS, Azure, or Google Cloud Platform (VMs, VPCs/Networking, IAM, Storage).
  • Stakeholder & Customer Communication: Act as the primary technical contact during critical incidents, ensuring clear, timely updates to internal teams and customers.

Requirements

  • 10+ years of progressive IT experience in Cloud Operations, System Administration, or Site Reliability Engineering (SRE).
  • Demonstrated experience handling high-severity incident management (Sev1/Sev2/Sev3) and writing post-incident RCA reports.
  • Hands-on proficiency with enterprise monitoring platforms (Datadog, Grafana, Prometheus, Splunk, Dynatrace, or New Relic).
  • Strong foundational knowledge of Compute (VMs, CPU/RAM utilization), Storage (SAN, NAS, Disk management), and Networking (TCP/IP, DNS, Firewalls, Load Balancers).
  • Solid operational grasp of at least one public cloud platform (AWS, Azure, or Google Cloud Platform).
  • Excellent communication and customer-handling skills during high-pressure outages.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:29 min

Recommended community resources for cloud engineers

Piet Van Dongen · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all