Network Simulation Engineer

Eridu Corporation
United States
7 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence C++ (Programming Language) Nvidia CUDA Computer Programming Distributed Computing Environment Ethernet High-Level Architecture Network Topologies InfiniBand Python (Programming Language) Tensorflow Systems Architecture
+8 more
Network Simulation Application Specific Integrated Circuits Pytorch Large Language Models Gpu Programming Data Center Networking Information Technology Modeling and Simulation

Job description

We are seeking a highly motivated Network Simulation Engineer to lead the simulation and analysis of AI communication workloads (e.g. collective communications) across various data center network topologies. In this role, you will apply network simulation tools to model real-world AI applications-including LLMs and DLRMs-to inform architectural decisions across Eridu’s product development lifecycle and to showcase our value to prospective customers and investors.

You will collaborate cross-functionally with customers, ASIC designers, and simulation tool providers to optimize performance, influence design, and deliver transformative AI networking solutions.

Responsibilities

  • Model AI Workloads: Simulate communication patterns of distributed AI workloads (e.g., LLMs, DLRMs) across diverse network topologies to analyze performance and scalability.
  • Drive Architecture Optimization: Work with customers to evaluate their AI workloads and provide recommendations for topology design, protocol tuning, and system architecture.
  • Influence ASIC Design: Collaborate with the internal ASIC and architecture teams by providing simulation-based insights that shape chip design for optimized AI traffic flows.
  • Tool Development & Partnership: Interface with simulation tool providers (who have optimized versions of NS-3, ASTRA-sim, etc.) to customize, tune, and enhance modeling frameworks for Eridu’s specific requirements and to operate these tools to run simulations.
  • Documentation & Communication: Create clear and compelling reports, documentation, and presentations to communicate insights to technical and non-technical stakeholders.

Requirements

  • MSc or PhD in Computer Science, Electrical Engineering, or a related field with some specialization in AI/ML communications or equivalent hands-on experience
  • Strong experience with network simulation tools such as NS3, OMNeT++, or custom-built simulators.
  • Familiarity with distributed training frameworks (e.g., PyTorch, TensorFlow), collective communication libraries (e.g., NCCL, RCCL), and GPU programming (CUDA or ROCm).
  • Deep understanding of frontier model architectures, parallelism approaches and operational functionality
  • Deep understanding of Ethernet, InfiniBand, and high-performance data center networking technologies.
  • Solid grasp of AI system architecture, including compute, memory, and interconnect bottlenecks in large-scale training/inference clusters.
  • Strong programming skills in C++ and Python.
  • Clear and confident communication skills, both written and verbal.
  • 2+ years of relevant experience preferred; exceptional early-career candidates will also be considered., Eridu does not accept unsolicited resumes or candidate profiles from staffing agencies or third-party recruiters. Any candidate submitted to Eridu without prior written authorization from our recruiting team will be considered unsolicited and will become the property of Eridu. Eridu reserves the right to pursue and hire such candidates without any obligation to pay fees. Recruiting agencies are expressly instructed not to contact hiring managers, employees, or executives regarding open positions.

Benefits & conditions

The starting base salary for the selected candidate will be established based on their relevant skills, experience, qualifications, work location, market trends, and the compensation of employees in comparable roles.

About the company

About Eridu

Eridu is a Silicon Valley-based hardware startup pioneering infrastructure solutions that accelerate AI data centers to deliver Faster AI. Today’s AI performance is frequently limited by communication bottlenecks. Eridu introduces multiple industry-first innovations across silicon, packaging, software, and systems to deliver an order of magnitude improvement in performance and unlock greater GPU utilization to speed training job completion times and tokens-per-second for more profitable inference. We do this while simultaneously reducing capital and power costs and improving reliability.

The company’s solutions and value proposition have been widely validated by leading hyperscalers.

Eridu has raised over $200M to date including its most recent, oversubscribed Series A round. The company is led by a veteran team of Silicon Valley executives who have delivered multiple billion dollar product lines and led multiple companies to billion dollar exits, including serial entrepreneur Drew Perkins, co-founder of Infinera (NASDAQ: INFN), Lightera (acq. by Ciena), Gainspeed (acq. by Nokia) and Mojo Vision (the world’s leading micro-LED company). The company is in execution mode and has a world-class engineering team with decades of experience in state-of-the-art silicon, packaging, optics, software, and systems. Eridu is working with best-in-class supply chain partners including silicon, packaging and systems., At Eridu, you’ll have the opportunity to shape the future of AI infrastructure, working with a world-class team on groundbreaking technology that pushes the boundaries of AI performance. Your contributions will directly impact the next generation of AI infrastructure solutions, transforming the performance of AI data centers.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:05 min

Acquiring Mellanox to build cohesive AI factories

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · World Congress 2025

4:52 min

Connecting namespaces with local virtual ethernet pairs

Oliver Seitz Oliver Seitz · World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

1:51 min

Bypassing the CPU stack with remote direct memory access

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

3:23 min

The AI workload technology stack and its components

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

Videos

See all

Related articles

See all