Engineer, Cloud HPC Platform

Ayar Labs
San Jose, CA, United States
12 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$120,000.0 - $150,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Backup Devices Bash Shell Cloud Computing Information Systems Data Transmissions Extract Transform Load (ETL) Desktop Computing Linux Desktop Virtualization
+31 more
File Systems Domain Name System (DNS) FlexNet Publisher Identity and Access Management Virtual Private Networks (VPN) Python (Programming Language) Key Management Routing Network Segmentation Citrix Systems Performance Tuning Red Hat Enterprise Linux Software Tools Ansible Runbook Software Deployment Data Logging Cloud Platform System Application Specific Integrated Circuits System Availability Software Troubleshooting Firewalls (Computer Science) Amazon Virtual Private Cloud (VPC) Gitlab-ci Information Technology Slurm Hardware Infrastructure Cloudwatch Physical Design Ansys Terraform

Job description

Ayar Labs is moving silicon engineering compute workloads into AWS. The Senior Cloud HPC Platform Engineer will design, build, and operate the cloud HPC platform that supports EDA, simulation, verification, physical design, AMS, and related engineering workflows.

You will lead the migration from the current RHEL-based compute environment to AWS while protecting engineering productivity, design data, license availability, and output correctness. You will work closely with IT, TFM, ASIC, AMS, verification, physical design, security, finance, and EDA vendors., * Lead the AWS migration: Inventory engineering workloads, dependencies, data, licenses, and performance requirements; define migration waves, cutover plans, rollback procedures, and acceptance criteria.

  • Build the cloud HPC platform: Design and operate scalable AWS compute using appropriate EC2 instance families, accelerated networking, autoscaling, placement strategies, and workload isolation.
  • Own scheduling and job execution: Make Slurm the primary scheduler and operational control plane for interactive, batch, regression, and multi-day simulation workloads. Deploy and operate Slurm directly and/or through AWS ParallelCluster where appropriate; configure partitions, QoS, priorities, fair-share, reservations, preemption, accounting, dependencies, job arrays, and policy-based autoscaling.
  • Plan engineering run capacity: Partner with design and verification teams before major regressions, simulations, and tapeout milestones to translate run manifests and workload forecasts into CPU/core, memory, GPU, wall-time, scratch and capacity I/O, network, license-token, Slurm partition/reservation, and budget requirements; publish capacity scenarios, reservations, and readiness risks.
  • Engineer storage and data movement: Design high-performance storage and tiering across services such as Amazon FSx, EFS, S3, and on-prem systems. Establish backup, lifecycle, replication, and recovery controls.
  • Enable EDA workloads: Build reproducible RHEL-compatible environments for Cadence, Synopsys, Ansys, and other engineering tools. Support PDKs, third-party IP, shared flows, and controlled releases.
  • Manage licenses: Design reliable FlexNet/FlexLM access across hybrid and cloud environments, monitor utilization, and prevent licensing from becoming a scaling bottleneck.
  • Automate the environment: Define infrastructure through Terraform or OpenTofu and automate images, configuration, patching, and application deployment with tools such as Packer and Ansible.
  • Prove performance and correctness: Benchmark representative workloads before and after migration. Validate runtime, queue time, storage performance, reliability, cost, and quality-of-results with engineering owners.
  • Deliver self-service access: Architect and operate secure virtual desktop infrastructure (VDI/DVI) for engineering workflows, including Citrix Virtual Apps and Desktops and/or NICE DCV/VNC. Own application publishing, golden images, patching, SSO/MFA, session brokering and policies, profile and storage integration, GPU/graphics support, clipboard and file-transfer controls, monitoring, capacity, high availability, and performance troubleshooting; provide documented self-service paths for launching jobs and remote sessions.
  • Own security and connectivity: Implement least-privilege IAM, network segmentation, encryption, secrets management, logging, vendor access controls, and secure connectivity through VPN and/or Direct Connect.
  • Operate for reliability: Establish observability, service objectives, incident response, runbooks, change controls, and disaster-recovery testing. Partner with engineering on run-demand forecasting and capacity planning for major regressions, simulations, and tapeout milestones.
  • Control cloud cost: Implement tagging, budgets, chargeback/showback, scheduling policies, idle-resource controls, and workload-specific cost/performance optimization. Use Slurm accounting and workload forecasts to provide engineering with resource and cost estimates before large runs.
  • Reduce operational fragility: Replace undocumented manual steps and one-off scripts with versioned, tested, supportable automation and clear documentation.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience.
  • 7+ years building and operating Linux infrastructure, including 3+ years in AWS or a comparable cloud environment.
  • Deep hands-on experience with Enterprise Linux in production as a System Administrator
  • Experience designing or operating HPC, batch compute, large-scale simulation, or similarly compute-intensive platforms.
  • Deep production experience administering Slurm as the primary HPC scheduler, including partitions, QoS, priorities, fair-share, reservations, preemption, accounting, job arrays and dependencies, failure recovery, upgrades, and integration with AWS ParallelCluster or equivalent cloud capacity.
  • Strong AWS experience across EC2, IAM, VPC, S3, CloudWatch, Systems Manager, KMS, and high-performance storage services.
  • Strong infrastructure-as-code skills using Terraform or OpenTofu, including reusable modules, state management, review, and testing.
  • Experience automating Linux images and configuration with Packer, Ansible, Python, and/or Bash.
  • Strong knowledge of high-performance and shared storage, Linux file systems, data transfer, backup, and recovery.
  • Production experience operating secure virtual desktop infrastructure (VDI/DVI) for engineering workloads, preferably Citrix Virtual Apps and Desktops, including application publishing, image and patch lifecycle, SSO/MFA, session brokering and policy, profile and storage integration, GPU/graphics, clipboard and file-transfer controls, monitoring, capacity and high availability, and performance troubleshooting.
  • Strong knowledge of cloud networking, DNS, routing, firewalls, VPN, and hybrid connectivity.
  • Experience supporting FlexNet/FlexLM or another network-license system.
  • Experience establishing monitoring, alerting, incident response, capacity management, and cost controls for production infrastructure.
  • Ability to partner directly with engineers, translate run manifests and workload forecasts into CPU/core, memory, GPU, wall-time, storage I/O and capacity, network, license-token, Slurm partition/reservation, and budget requirements, and communicate capacity and migration risks clearly.
  • Clear written documentation, design proposals, operating procedures, and post-incident reviews., * Experience supporting semiconductor EDA environments, including Cadence, Synopsys, Ansys, Siemens EDA, PDKs, and IP libraries.
  • Experience migrating EDA, HPC, simulation, or verification workloads from on-prem infrastructure to AWS.
  • Experience with Amazon FSx for Lustre, FSx for OpenZFS, EFA, AWS Batch, ParallelCluster, or equivalent HPC services.
  • Experience designing and operating Citrix Virtual Apps and Desktops or comparable virtual desktop infrastructure (VDI/DVI) for engineering workloads, including application publishing, image and patch lifecycle, SSO/MFA, session brokering and policies, GPU/graphics, profile and storage integration, high availability, monitoring, and performance troubleshooting.
  • Experience benchmarking workload runtime, queue time, storage I/O, scaling efficiency, quality-of-results, and cost.
  • Familiarity with GitLab CI, artifact repositories, observability platforms, and controlled release processes.
  • AWS Professional or Specialty certification.

About the company

Ayar Labs is shattering AI data bottlenecks by moving data at the speed of light. As pioneers of co-packaged optics (CPO), we are using light instead of electricity to move data faster, further, and with a fraction of the energy needed to fuel the explosive growth of AI models.

Backed by industry giants like NVIDIA, AMD and Intel and manufactured in partnership with the world’s leading semiconductor ecosystem, Ayar Labs’ co-packaged optics solution is key to unleashing next-generation AI scale-up architectures.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

55 sec

Generating ASCII art branding for AI agent interfaces

Chris Heilmann +2 · LIVE

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:24 min

Evaluating formal AWS certifications versus raw practical engineering experience

Jan Giacomelli · LIVE

2:31 min

Simplifying terminal output with new escape sequences

Ambesh Singh +1 · LIVE

Videos

See all

Related articles

See all