Systems Design Eng - AI Cluster Software Engineer

Advanced Micro Devices, Inc.
Austin, TX, United States
about 1 month ago
Apply on diversityjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$133,200.0
Working hours
Regular working hours

Tech stack

JavaScript (Programming Language) Application Programming Interfaces (APIs) Agile Methodology Artificial Intelligence User Authentication Automation of Tests Bash Shell Code Review Computer Engineering Continuous Integration Software Debugging Linux
+27 more
File Systems Python (Programming Language) Performance Tuning Remote Direct Memory Access Ansible Web Application Security Software Engineering TypeScript Web Applications Scripting Graphics Processing Unit (GPU) High Performance Computing ReactJS Large Language Models Caching Backend Vue.js Perf (Linux) AngularJS Kubernetes Information Technology Slurm Front End Software Development Hardware Infrastructure Terraform Software Version Control Data Pipelines

Job description

This is a hands-on role for a full-stack developer to create and deliver a range of tools and applications focused on design and deployment of large-scale AI/ML clustered infrastructure. You will be working with the latest agentic tools and patterns to develop, deploy, and maintain these applications. You’ll join a growing team of multi-disciplined engineers that operates across industry verticals as subject matter experts in the AI stack and across the cluster., * Demonstrated use of AI coding assistants and LLM-powered developer tools: daily user of AI agents and tools

  • Professional software development experience, including substantial experience building and supporting web applications
  • Proficiency in modern frontend development using JavaScript or TypeScript and a framework such as React, Angular, or Vue
  • Experience developing backend services and APIs using a modern server-side language or framework
  • Experience deploying and operating applications using Linux, containers, CI/CD, and cloud or on-premises infrastructure
  • Experience with automated testing, source control, code review, debugging, and production support
  • Working knowledge of web application security, authentication, authorization, and secure secrets handling, * Partner with engineering peers, domain experts in adjacent teams, and business stakeholders to understand requirements and translate them into flexible, future-proof design solutions
  • Hands on development, iteration, and maintenance of tools and applications that codify various aspects of large-scale AI cluster design stages and cluster deployment activities
  • Design and development of cohesive interface code between disparate third party tools
  • Own features from requirements and design through deployment and ongoing maintenance
  • Work in an iterative software environment, including planning and delivering work in small increments, collaborating with stakeholders, often in different areas of domain expertise (Agile development practices)
  • Participate in code reviews and retros; adapt to changing requirements and priorities, AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

Requirements

  • Strong Linux fundamentals: Linux operating systems, networking, filesystems, containers, performance tooling (perf, flamegraphs, nvprof/rocprof, basic eBPF).
  • Clear communication: ability to turn complex systems into accessible, structured documentation with diagrams and reproducible steps
  • AMD ecosystem experience: ROCm, RCCL, Instinct GPUs, EPYC platforms, compiler/toolchain impacts, and performance tuning
  • Orchestration models: Slurm configuration patterns, Kubernetes for HPC/AI (GPU operators, device plugins), Apptainer/Singularity
  • Automation, IaC , and scripting tools/languages (Ansible, Terraform, Python, bash)
  • Storage/data: knowledge of or familiarity with parallel filesystems (Lustre, BeeGFS), object stores, RDMA, data pipeline throughput and caching strategies
  • Hands-on familiarity with on-premises infrastructure, particularly for AI/ML/HPC workloads would be beneficial

ACADEMIC CREDENTIALS:

Bachelors or Masters degree in computer science, or software/computer engineering

About the company

At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them.

AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger - technology that moves the world forward.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:10 min

Analyzing the out-of-the-box security posture of Vue.js

Philippe De Ryck · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:58 min

Verifying hardware access and exploring AI inference scaling

Piotr Zaniewski Piotr Zaniewski · World Congress 2026 Europe

1:48 min

Overview of the target real-time application

Abdelrahman Awad Abdelrahman Awad · World Congress 2025

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all