Core Platform Engineer - Infrastructure Reliability & Incident Management

FSTONE Technologies
Sunnyvale, United States
10 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Systems Engineering Kernel-Based Virtual Machine Quick EMUlator (QEMU) Reliability Engineering Pure Storage Kubernetes Hardware Infrastructure Pagerduty

Job description

  • Act as first responder and lead Sev-1/Sev-2 production incidents.
  • Lead incident bridges and coordinate cross-functional SMEs.
  • Perform infrastructure triage using logs, metrics, and telemetry.
  • Drive incident communication, resolution, and post-incident improvements.
  • Implement SRE best practices, SLIs/SLOs, observability, and automation.
  • Support on-call, change management, and operational readiness.

Requirements

  • 5-10+ years in SRE, Platform Engineering, Infrastructure, or Systems Engineering.
  • Strong hands-on infrastructure troubleshooting.
  • Deep expertise in at least one: KVM/QEMU, OVN/OVS/SDN, GPU infrastructure, or Storage (Lightbits/Pure Storage).
  • Production Kubernetes/GKE experience preferred.
  • Strong knowledge of monitoring, observability, SLI/SLO, and incident management.
  • Experience with incident.io, PagerDuty, Opsgenie, or similar tools.
  • Strong communication and stakeholder-management skills., A senior infrastructure/SRE engineer who can quickly triage production incidents, identify the affected infrastructure domain, coordinate SMEs, and drive incidents to resolution.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

4:14 min

Handling external dependencies through pure services

Abdu Taviq Abdu Taviq · Europe 2026 Virtual

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

1:57 min

Streamlining incident response and root cause analysis automatically

Mike Mike · World Congress 2025

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

3:31 min

Storing prefetch links automatically via intersection observer api

Rowdy Rabouw Rowdy Rabouw · World Congress 2023

Videos

See all

Related articles

See all