> Markdown version of [/jobs/ext/2303531-system-specialist](https://www.wearedevelopers.com/jobs/ext/2303531-system-specialist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # System Specialist - **Company:** TEKGENCE INC - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Systems Engineering, Build Automation, Cloud Computing, Data Centers, Linux, File Systems, Distributed Data Store, Linux Kernel, Performance Tuning, Reliability Engineering, Graphics Processing Unit (GPU), High Performance Computing, SDN Network, Hardware Infrastructure, Block Storage, Nvme - **Published:** August 30, 2026 - **Apply:** https://www.disabledperson.com/jobs/74629334-system-specialist ## About the Role Skill required - GPU (Must), Exposure SRE, Storage & Network, GPU Infrastructure & SDN. As an Embedded PE, we will function as a member of a foundation engineering team, owning reliability, performance, scalability, and operational excellence for critical infrastructure supporting fastest-growing AI cloud platforms. This role requires deep technical expertise in a specific infrastructure domain and the ability to operate across the full service lifecycle, including design, implementation, optimization, incident response, and roadmap planning. Required Qualifications 10 to 15+ years of infrastructure engineering, platform engineering, SRE, or systems engineering experience. Deep expertise in a specific infrastructure domain. Strong Linux systems knowledge. Demonstrated experience operating production-critical infrastructure. Hands-on incident response and troubleshooting experience. Experience building automation and operational tooling. Ability to clearly articulate technical contributions to complex projects. ## Description * Compute: HPC environments, AMD/NVIDIA GPUs, Linux Kernel/OS, NUMA awareness ,CPU pinning, huge pages, and performance optimization * Storage: NVMe storage and distributed storage systems, High performance Block, file, and Object Storage Experience * Software-Defined Networking (SDN): Design, deployment, and operational support of SDN infrastructure For L1 role (Major skill required) * GPU infrastructure (is not mandatory) * Networking * Block storage * File storage * Production support and reliability engineering Hi , Embedded Platform/Infrastructure Engineer (Level 2 Engineer) Sunnyvale, CA or San Jose, CA (5 days Onsite, Final round F2F), Embed within a foundation engineering team and operate as a domain expert. Participate in on-call rotations and production incident management. Design, implement, and improve infrastructure reliability and scalability. Collaborate with architects and engineers on long-term technical roadmaps. Analyze complex infrastructure performance bottlenecks and system failures. Build automation to improve operations, deployment, and observability. Define and track service reliability objectives (SLIs/SLOs). Contribute to platform standards and engineering best practices. Partner with data center and operations teams during service-impacting events. Drive continuous improvements in service stability, efficiency, and performance. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Discover the open source trio you didn’t expect: .NET and PostgreSQL on Linux](https://www.wearedevelopers.com/videos/2042-discover-the-open-source-trio-you-didn-t-expect-net-and-postgresql-on-linux) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [A Guide to Green Tech and Green IT Careers](https://www.wearedevelopers.com/magazine/374-a-guide-to-green-tech-and-green-it-careers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)