Sr. System Engineer

Super Micro Computer, Inc.
San Jose, CA, United States
20 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$137,000.0 - $156,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Big Data Cloud Computing Software Quality Computer Programming Computer Engineering Data Centers Dynamic Host Configuration Protocol Software Debugging DevOps
+21 more
Domain Name System (DNS) IPv4 IPv6 Python (Programming Language) Network Configuration and Change Management Network Architecture OpenStack Software Reliability Testing Tensorflow Shell Script TCP/IP Graphics Processing Unit (GPU) Cloud Platform System High Performance Computing Pytorch Large Language Models Test Scripts System-level Testing Kubernetes Storage Technologies Docker

Job description

As a global leader in server technologies, Supermicro has been growing extremely fast in many key markets such as Cloud Computing, Big Data, HPC, AI and Storage, etc. To meet the market demand, Supermicro is developing end to end enterprise IT solutions with compute, storage, networking all integrated into full rack or multi-rack level systems. Senior System Engineer plays an important role in designing, implementing, testing and deploying rack system solutions for data center and enterprise customers., Includes the following essential duties and responsibilities (other duties may also be assigned):

  • Deploy Rack/Cluster infrastructure and execute comprehensive system level testing on the latest GPUs, CPU processors, Network and Storage, encompassing functionality, compatibility, performance, stress, and reliability testing, leveraging proprietary in-house tools
  • Conduct proof of concept design and testing. Establish expertise in HPC/AI applications and benchmarks, providing optimized benchmarks for HPC/AI applications by fine-tuning system settings, optimizing OS/network configurations, and demonstrating strong problem-solving skills and building robust processes and procedures for HPC/AI solutions
  • Lead day-to-day operational support for Cluster, Storage, HPC and Cloud infrastructure. Identify and document hardware and software quality issues. Collaborate with product management and other Engineering teams to integrate enhancements into future products
  • Write technical documents for test procedures, test reports and troubleshooting procedures related to servers/networks/clusters software and hardware to facilitate knowledge sharing
  • Deliver on-site deployment services to ensure customer acceptance verification and satisfaction
  • Write automation tools for cluster deployment and test environment

Requirements

  • BS/MS in Electrical Engineering, Computer Engineering or a related field, MS preferred
  • 8+ years of work-related experience in server/network/storage hardware configuration, testing, debugging and troubleshooting
  • 8+ years of work-related experience in DevOps or in cloud environments, including but not limited to Docker/Containers and Kubernetes
  • Experience with leading AI/ML frameworks such as PyTorch, TensorFlow, etc.
  • Familiar with TCP/IP protocol stack, UDP, IPv4-IPv6, DNS, DHCP and other Application protocols
  • Familiar with HPC, AI or Cloud benchmark tests, networking architecture
  • Excellent Programming skills in Python and shell scripting
  • Strong communication skills and strong sense of teamwork and good team player
  • Familiar with MLPerf Training/Inference benchmark, LLM, HPL-AI or RCCL/NCCL is a plus
  • CCNA, OpenStack, Openshit, Azure or AWS is a plus Salary Range

About the company

Supermicro is a Top Tier provider of advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon Valley Top 50 technology firms. Our unprecedented global expansion has provided us with the opportunity to offer a large number of new positions to the technology community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.localjobnetwork.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

1:34 min

The pros and cons of campus-wide IP authentication

Christoph Eicke Christoph Eicke · WWC 2025

2:22 min

Introducing Skupper for application connectivity

Alex Soto Alex Soto · WWC 2024

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:04 min

Insights on transitioning from supercomputing to technical education

Andrew Holway · LIVE

3:05 min

Exploring microcontrollers and communication protocols for amateur hardware

Philipp-Alexander Blum · LIVE

Videos

See all

Related articles

See all