10 - Staff Engineer, Software

Celestica, Inc.
Austin, TX, United States
22 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Testing (Software) Artificial Intelligence Amazon Web Services Data Analysis Automation of Tests Microsoft Azure Bash Shell BIOS Ubuntu (Operating System) CentOS Command-Line Interface Profiling
+36 more
Computer Networks Continuous Integration Data Centers Software Debugging Linux Ethernet Firmware Hardware Design InfiniBand Python (Programming Language) Machine Learning NetApp Applications Open Source Technology PCI Express Red Hat Enterprise Linux Software Reliability Testing Tensorflow Software Engineering Subsystems TCP/IP Strategies of Testing Virtualization Technology Scripting Google Cloud Enterprise Software Applications Cloud Platform System Performance Testing Data Ingestion Pytorch Perf (Linux) Containerization Kubernetes Information Technology Hardware Infrastructure U-Boot Docker

Job description

We’re looking for a skilled and experienced Staff Test Engineer, AI Data Center Infrastructure to help ensure the reliability, performance, and scalability of our AI data center’s networking, storage and server infrastructure. In this role, you’ll be a key technical contributor, working with cross-functional teams to develop and execute comprehensive test strategies for our critical hardware, firmware, and software components. Your deep expertise in storage and server systems will be essential as we deliver robust and high-performing solutions that support demanding AI/ML workloads. The ideal candidate will be a hands-on technical leader, capable of mentoring junior engineers, driving test automation, and collaborating across engineering teams to deliver robust and high-performing solutions., * Define and implement test strategies for all storage and server hardware, firmware, and software components within the AI data center environment.

  • Lead the definition and development of holistic test strategies, test plans and test cases for complex data center solutions, including functional, performance, reliability, stress, and endurance testing.
  • Mentor and provide technical guidance to junior test engineers, fostering a culture of technical excellence and continuous improvement.
  • Design and implement automated test frameworks and scripts using languages like Python, Go, or similar, to improve efficiency and coverage of testing.
  • Conduct in-depth performance analysis and bottleneck identification for server platforms (e.g., CPU, GPU, memory, PCIe, networking), security (e.g., secure boot, Root of Trust, Platform Firmware Resilience) and OpenBMC interfaces/features
  • This includes debugging issues related to BIOS, BMC functionality and its interaction with server hardware.
  • Develop and maintain robust testbeds and infrastructure for continuous integration and validation.
  • Utilize open-source and commercial test tools relevant to server, BIOS, OpenBMC and storage validation.
  • Collaborate closely with hardware design, software development, infrastructure, and AI/ML engineering teams to understand requirements and integrate testing throughout the product lifecycle.
  • Communicate test progress, results, and critical issues effectively to stakeholders, including executive leadership.
  • Develop specialized test methodologies to validate performance and reliability under heavy AI/ML workloads (e.g., large model training, inference at scale, data ingestion).
  • Understand and test the interactions

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related technical field.
  • 10+ years of experience in hardware and/or software testing, with at least 5 years focused on enterprise-level storage and server systems.
  • 5+ years of experience in a lead or senior technical role, mentoring junior engineers or leading test initiatives.
  • Deep expertise in server architectures (x86, ARM, GPU servers), CPU/memory subsystems, PCIe, and power management.
  • Extensive experience in server architectures (x86, ARM, GPU servers), BIOS, CPU/memory subsystems, PCIe, power management, and Baseband Management Controllers (BMC) functionality.
  • Strong understanding of enterprise software security (e.g., secure boot, Root of Trust, Platform Firmware Resilience)
  • Proficiency in scripting languages (e.g., Python, Bash) for test automation and data analysis.
  • Experience with Linux operating systems (e.g., Ubuntu, CentOS, RHEL) and command-line tools.
  • Familiarity with networking concepts (Ethernet, TCP/IP, InfiniBand) and network testing methodologies.
  • Experience with test methodologies such as performance testing, reliability testing, stress testing, and fault injection.
  • Excellent problem-solving, analytical, and debugging skills.
  • Strong communication and interpersonal skills, with the ability to collaborate effectively across diverse teams.

Preferred Qualifications

  • Familiarity with OCP (Open Compute Project)
  • Experience with cloud environments (AWS, Azure, GCP) and virtualization technologies.
  • Knowledge of containerization technologies (Docker, Kubernetes).
  • Familiarity with AI/ML frameworks (e.g., TensorFlow, PyTorch) and their infrastructure requirements.
  • Experience with performance profiling tools (e.g., fio, Iometer, Perf, VTune).
  • Contributions to open-source projects related to storage, servers, or testing.
  • Certifications in relevant technologies (e.g., NetApp, Dell EMC, HPE, NVIDIA).

About the company

Celestica (NYSE, TSX: CLS) enables the world’s best brands. Through our recognized customer-centric approach, we partner with leading companies in Aerospace and Defense, Communications, Enterprise, HealthTech, Industrial, Capital Equipment and Energy to deliver solutions for their most complex challenges. As a leader in design, manufacturing, hardware platform and supply chain solutions, Celestica brings global expertise and insight at every stage of product development - from drawing board to full-scale production and after-market services for products from advanced medical devices, to highly engineered aviation systems, to next-generation hardware platform solutions for the Cloud. Headquartered in Toronto, with talented teams spanning 40+ locations in 13 countries across the Americas, Europe and Asia, we imagine, develop and deliver a better future with our customers.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

41 sec

Massive client data loss and bio-digital storage

Chris Heilmann +1 · LIVE

40 sec

Hardware durability labs and robot testing methods

Chris Heilmann +1 · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all