Principal AI and Machine Learning Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+11 more
Job description
This role has been designed as ‘Hybrid’ with a requirement that you will work on average 2 days per week from an HPE office., As a key contributor to AI/ML infrastructure initiatives, you will plan, execute, and analyze comprehensive benchmarks on switches, focusing on throughput, latency, congestion, incast, failover, path diversity, and workload performance to ensure optimal AI/ML network operations.
You will be guiding AI/ML workload deployments from initial scoping and test planning through execution and benchmark analysis, ensuring success criteria are met. Your role includes developing AI-driven automation workflows to streamline network development, operations, and implementations.
You will validate switch ASIC features including buffers, schedulers, QoS/queuing, ECMP behavior, telemetry, hashing, traffic distribution, and congestion visibility.
Owning switch OS configuration and automation, you will utilize SONiC, Junos, Ansible, Python, Bash, Git, and related tooling to implement and validate advanced features such as SRv6, segment routing, uSID, Adj-SID, and policy-based pathing as required. You will document PoC architecture, benchmark methodologies, topology diagrams, configurations, results, findings, and recommendations.
This role empowers you to shape the future of AI infrastructure networking by delivering scalable, high-performance, and resilient network fabrics that meet the stringent demands of AI/ML workloads, driving innovation and customer success at < >
Requirements
Bachelors + 7 years of related experience, or Masters + 4 years of related experience. Python for automation experience. Experience with L2/L3 network protocols such as BGP, OSPF, EVPN, VxLAN, IPv6 or similar. Experience with Traffic tools such as Spirent, IXIA or similar. Docker or Kubernetes experience. Experience with network testing and validation.
Preferred Qualifications
Clear written and verbal communication skills as well as documentation skills. SONiC, Junos, Linux or other open source network operating systems experience. Deep understanding of Leaf-spine fabric and troubleshooting them. Experience with Apstra and related automation tools for provisioning, managing and troubleshooting the fabric. Experience handling complex network segmentation, security policies, and multi-site fabric designs. Experience with RDMA, RoCEv2, PFC, ECN, congestion control, QoS, buffer behavior, and lossless Ethernet concepts.
Benefits & conditions
“The expected salary/wage range for this position is provided below. Actual offer may vary from this range based upon geographic location, work experience, education/training, and/or skill level.
- United States of America: Annual Salary USD 172,000 - 349,000 in California The listed salary range reflects base salary. Variable incentives may also be offered.”
About the company
Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this role…