System Specialist
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+4 more
Job description
- Compute: HPC environments, AMD/NVIDIA GPUs, Linux Kernel/OS, NUMA awareness ,CPU pinning, huge pages, and performance optimization
- Storage: NVMe storage and distributed storage systems, High performance Block, file, and Object Storage Experience
- Software-Defined Networking (SDN): Design, deployment, and operational support of SDN infrastructure
For L1 role (Major skill required)
- GPU infrastructure (is not mandatory)
- Networking
- Block storage
- File storage
- Production support and reliability engineering
Hi ,
Embedded Platform/Infrastructure Engineer (Level 2 Engineer)
Sunnyvale, CA or San Jose, CA (5 days Onsite, Final round F2F), Embed within a foundation engineering team and operate as a domain expert.
Participate in on-call rotations and production incident management.
Design, implement, and improve infrastructure reliability and scalability.
Collaborate with architects and engineers on long-term technical roadmaps.
Analyze complex infrastructure performance bottlenecks and system failures.
Build automation to improve operations, deployment, and observability.
Define and track service reliability objectives (SLIs/SLOs).
Contribute to platform standards and engineering best practices.
Partner with data center and operations teams during service-impacting events.
Drive continuous improvements in service stability, efficiency, and performance.
Requirements
Skill required - GPU (Must), Exposure SRE, Storage & Network, GPU Infrastructure & SDN.
As an Embedded PE, we will function as a member of a foundation engineering team,
owning reliability, performance, scalability, and operational excellence for critical
infrastructure supporting fastest-growing AI cloud platforms.
This role requires deep technical expertise in a specific infrastructure domain and the
ability to operate across the full service lifecycle, including design, implementation,
optimization, incident response, and roadmap planning.
Required Qualifications
10 to 15+ years of infrastructure engineering, platform engineering, SRE, or systems engineering
experience.
Deep expertise in a specific infrastructure domain.
Strong Linux systems knowledge.
Demonstrated experience operating production-critical infrastructure.
Hands-on incident response and troubleshooting experience.
Experience building automation and operational tooling.
Ability to clearly articulate technical contributions to complex projects.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Is Software Engineering Over-Saturated?
7 Cloud Computing Trends Coming in 2025 for Developers
A Guide to Green Tech and Green IT Careers