Storage Benchmarking Engineer

Aziro Technologies Llc
Santa Clara, CA, United States
about 1 month ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Microsoft Excel Artificial Intelligence Amazon S3 Bash Shell Cloud Storage Software Debugging Distributed Systems Ethernet InfiniBand Python (Programming Language) Open Source Technology Performance Tuning
+16 more
Remote Direct Memory Access Ansible Statistical Process Control (SPC) Tableau (Software) Technical Data Management Systems AI Infrastructure Graphics Processing Unit (GPU) Delivery Pipeline Software Troubleshooting Kubernetes Information Technology Data Management Data Pipelines Block Storage Docker Nvme

Job description

The Storage Benchmarking Engineer will design, execute, and analyze performance benchmarks spanning both industry-standard storage benchmarks (fio, vdbench, SPEC SFS 2020, IO500, SPC-1/SPC-2) and the emerging class of AI/ML storage workloads (MLPerf Storage, DLIO, and GPU-driven training and inference data pipelines). As AI has made storage a first-order bottleneck in the GPU data path, this role sits at the intersection of high-performance storage and large-scale AI infrastructure., * Configure and scale the HPC/AI lab environment so all systems including GPU servers, high-speed fabrics, and storage achieve maximum efficiency and scale across a variety of test harnesses. Build robust automation so labs can be rapidly configured and reconfigured to meet the demands of different benchmarks.

  • Design and execute storage performance benchmarks using industry-standard tools and methodologies, including fio, vdbench, SPEC SFS 2020, IO500, and SPC-1/SPC-2 (or similar).
  • Design and execute AI/ML storage benchmarks, including MLPerf Storage, DLIO, and representative AI workloads model training and checkpointing, inference and data ingest, RAG/vector-database access patterns, and GPU-driven I/O paths (e.g., GPUDirect Storage, NFS/RDMA). Characterize storage behavior against reference architectures such as NVIDIA DGX/SuperPOD and BasePOD.
  • Perform end-to-end performance troubleshooting and debugging across compute, GPU, network, and storage components to pinpoint and resolve bottlenecks and achieve best-in-class results.
  • Develop and maintain automated benchmarking workflows using tools like Ansible, Python, or Bash to ensure rapid provisioning and efficient, repeatable, reproducible results.
  • Analyze benchmark results, generate detailed reports, and deliver actionable insights to engineering teams for product optimization.
  • Collaborate with engineering, product management, marketing, and sales to align benchmarking efforts with product goals and customer needs.
  • Engage directly with benchmark standards organizations (e.g., SPEC, SNIA, MLCommons) and communities to influence methodologies, drive submissions, and stay ahead of industry and AI infrastructure trends.
  • Deliver high-impact presentations to internal teams, customers, and external stakeholders, translating complex technical data into clear narratives.
  • Write technical marketing documents, whitepapers, and performance summaries to support product launches and customer engagement.
  • Maintain comprehensive documentation of benchmarking processes, configurations, and results.
  • We are primarily an in-office environment, and you will be expected to work from the Santa Clara office in compliance with Everpure s policies, unless you are on PTO, work travel, or other approved leave.

Requirements

This position demands strong end-to-end performance troubleshooting across the entire stack compute (including GPUs), network (including RDMA/InfiniBand and high-speed Ethernet), and storage together with close collaboration across engineering, product management, marketing, and sales. The ideal candidate has hands-on experience with both classic storage benchmarks and AI data-pipeline benchmarking, a track record engaging benchmark standards organizations and communities, and exceptional communication and writing skills., * Bachelor s or Master s degree in Computer Science, Electrical Engineering, or a related field (or equivalent experience).

  • 5+ years of experience in storage performance benchmarking or a related technical role.
  • Proven expertise with storage benchmarking tools, including:
  • Fio flexible I/O tester for workload simulation
  • vdbench versatile storage benchmarking for enterprise workloads
  • SPEC SFS 2020 filesystem and application-level benchmarking
  • IO500 and SPC-1/SPC-2 (or similar) HPC and block storage benchmarking
  • Experience benchmarking AI/ML storage workloads, such as MLPerf Storage, DLIO, or characterizing storage for GPU-based training and inference pipelines (data ingest, checkpointing, GPUDirect Storage, RDMA-based access). (Strongly preferred)
  • Strong end-to-end performance tuning and troubleshooting skills across compute, network, and storage layers and ideally GPU/accelerator data paths.
  • Hands-on experience with automation tools (e.g., Python, Ansible, Bash) for test orchestration and data collection.
  • Demonstrated track record of interfacing with engineering, product management, marketing, and sales teams.
  • Direct engagement with benchmark standards organizations (e.g., SPEC, SNIA, MLCommons) and storage benchmarking communities.
  • Exceptional communication skills, with a proven ability to deliver engaging technical briefs and presentations to technical and non-technical audiences.
  • Strong writing skills, ideally authoring technical marketing documents, whitepapers, or performance summaries.
  • Proficiency in data analysis and visualization tools (e.g., Python, R, Sheets, Excel, Tableau).
  • Deep understanding of and hands-on proficiency with enterprise storage technologies (e.g., NVMe, NVMe-oF, SSDs, and scale-up and scale-out block/file/object storage and distributed systems).
  • Experience with Everpure Storage products (e.g., FlashArray, FlashBlade) or similar enterprise block, file, and object storage platforms.
  • Familiarity with cloud storage (e.g., AWS S3, Azure Blob, Google Cloud Storage) and HPC or AI training environments.
  • Knowledge of containerized and orchestrated benchmarking workflows using Docker or Kubernetes.
  • Familiarity with high-speed networking and GPU fabrics (e.g., InfiniBand, RoCE, NVLink) as they relate to storage performance. (Preferred)
  • Prior contributions to benchmark standards or open-source performance tools.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · World Congress 2025

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all