> Markdown version of [/jobs/ext/2551134-performance-engineer](https://www.wearedevelopers.com/jobs/ext/2551134-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Performance Engineer - **Company:** Aziro Technologies Llc - **Location:** Santa Clara, CA, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon S3, Computer Clusters, Software Debugging, Distributed Systems, InfiniBand, Network File Systems, Performance Tuning, Remote Direct Memory Access, Weka, Rate Limiting - **Published:** August 4, 2026 - **Apply:** https://www.dice.com/job-detail/d79a7093-c2de-48bb-a93c-f92b9fcd710f ## About the Role * Deep, hands-on background in storage performance engineering across file, block, and object protocols, ideally with direct HPC or hyperscale exposure (parallel filesystems, pNFS/NFS at scale, S3-scale object stores). * Working knowledge of GPU cluster architecture RDMA fabrics, GPUDirect Storage, checkpoint/restore patterns for large model training and how storage bottlenecks manifest in mixed compute/network/storage systems. * Fluency with industry benchmark standards (MLPerf Storage, IO500) and load-generation tooling (elbencho, fio, vdbench), plus the judgment to design workload-representative tests beyond canned benchmarks. * Demonstrated ability to build rigorous, technically credible competitive analysis (not slideware) against systems like VAST, DDN, and WEKA architecture-level understanding, not just spec-sheet comparison. * Experience with multi-tenant resource management concepts (QoS, rate limiting, workload isolation) in a distributed systems context. * Comfortable operating across the stack and across audiences deep enough to debug an RDMA queue-pair stall or a metadata hot-partition, articulate enough to brief a NeoCloud customer's technical evaluation team. ## Description * Set the performance architecture agenda. Bring deep, current expertise across file, block, and object storage protocols and translate it into concrete performance requirements for //EXA's data path (NFSv3 direct-to-DN, pNFS layouts, S3 via MDN) informed by how HPC and hyperscale environments actually push storage systems (checkpointing, small-file metadata storms, GPU-starved read patterns, mixed-tenant burst I/O). * Track and act on the NeoCloud / sovereign-cloud shift. Maintain a living view of where //EXA's highest-value deployments are heading GPU-cloud and sovereign-cloud operators (CoreWeave, Crusoe, Nscale, and similar) and make sure //EXA's performance roadmap, reference architectures, and sizing guidance map to how these operators actually buy and operate infrastructure (multi-tenant GPU clusters, bursty training/inference mixes, strict SLAs to end customers). * Own competitive performance positioning. Build and maintain deep, technically substantiated comparisons against VAST Data, DDN, and WEKA not marketing bullet points, but real architectural analysis (metadata scaling model, erasure coding/durability tradeoffs, protocol support, GPU-direct paths, cost/performance at scale) that engineering and field teams can use to win technical evaluations and POCs. * Drive performance tuning for multi-tenant HPC/AI workloads. Lead tuning and validation work spanning the full stack a GPU cluster touches storage (MDN/DN geometry, pack groups, erasure coding layout), networking (RDMA, RoCE/InfiniBand fabric behavior, NIC/queue tuning), and compute (GPU-side I/O patterns, checkpoint/restore, data loader behavior) with particular focus on how these interact when multiple tenants/workloads share the same //EXA fleet. * Build and run the benchmark suite. Own //EXA's benchmark framework and result credibility: MLPerf Storage (v2/v3), elbencho, IO500, fio/vdbench-class synthetic tests, and workload-representative benchmarks for AI training/inference and traditional HPC. Ensure results are reproducible, defensible in public disclosure, and directly comparable to published competitor numbers. * Define QoS, limits, and workload segmentation. Drive the technical requirements and validation for quality-of-service guarantees, per-tenant/per-workload throughput and IOPS limits, and workload isolation the mechanisms that let //EXA make hard SLA commitments in shared, multi-tenant NeoCloud deployments rather than best-effort performance. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [Building Systems that Last](https://www.wearedevelopers.com/videos/1389-building-systems-that-last) - [Hate organising your photos? Try it with 5 Terabytes](https://www.wearedevelopers.com/videos/79-hate-organising-your-photos-try-it-with-5-terabytes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Why Event-Driven Architecture Isn’t About Speed (and When You Actually Need It)](https://www.wearedevelopers.com/magazine/745-why-event-driven-architecture-isn-t-about-speed-and-when-you-actually-need-it)