System Specialist

TEKGENCE INC
San Jose, CA, United States
2 days ago
Apply on www.disabledperson.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours

Tech stack

Systems Engineering Build Automation Cloud Computing Data Centers Linux File Systems Distributed Data Store Linux Kernel Performance Tuning Reliability Engineering Graphics Processing Unit (GPU) High Performance Computing
+4 more
SDN Network Hardware Infrastructure Block Storage Nvme

Job description

  • Compute: HPC environments, AMD/NVIDIA GPUs, Linux Kernel/OS, NUMA awareness ,CPU pinning, huge pages, and performance optimization
  • Storage: NVMe storage and distributed storage systems, High performance Block, file, and Object Storage Experience
  • Software-Defined Networking (SDN): Design, deployment, and operational support of SDN infrastructure

For L1 role (Major skill required)

  • GPU infrastructure (is not mandatory)
  • Networking
  • Block storage
  • File storage
  • Production support and reliability engineering

Hi ,

Embedded Platform/Infrastructure Engineer (Level 2 Engineer)

Sunnyvale, CA or San Jose, CA (5 days Onsite, Final round F2F), Embed within a foundation engineering team and operate as a domain expert.

Participate in on-call rotations and production incident management.

Design, implement, and improve infrastructure reliability and scalability.

Collaborate with architects and engineers on long-term technical roadmaps.

Analyze complex infrastructure performance bottlenecks and system failures.

Build automation to improve operations, deployment, and observability.

Define and track service reliability objectives (SLIs/SLOs).

Contribute to platform standards and engineering best practices.

Partner with data center and operations teams during service-impacting events.

Drive continuous improvements in service stability, efficiency, and performance.

Requirements

Skill required - GPU (Must), Exposure SRE, Storage & Network, GPU Infrastructure & SDN.

As an Embedded PE, we will function as a member of a foundation engineering team,

owning reliability, performance, scalability, and operational excellence for critical

infrastructure supporting fastest-growing AI cloud platforms.

This role requires deep technical expertise in a specific infrastructure domain and the

ability to operate across the full service lifecycle, including design, implementation,

optimization, incident response, and roadmap planning.

Required Qualifications

10 to 15+ years of infrastructure engineering, platform engineering, SRE, or systems engineering

experience.

Deep expertise in a specific infrastructure domain.

Strong Linux systems knowledge.

Demonstrated experience operating production-critical infrastructure.

Hands-on incident response and troubleshooting experience.

Experience building automation and operational tooling.

Ability to clearly articulate technical contributions to complex projects.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.disabledperson.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

8:22 min

Simulating a Linux terminal and running Spring Boot

Jakov Semenski · LIVE

Videos

See all

Related articles

See all