System Design Engineer - AI Cluster Software Engineer
Role details
Job location
Tech stack
Job description
This is a hands-on role for a full-stack developer to create and deliver a range of tools and applications focused on design and deployment of large-scale AI/ML clustered infrastructure. You will be working with the latest agentic tools and patterns to develop, deploy, and maintain these applications. You'll join a growing team of multi-disciplined engineers that operates across industry verticals as subject matter experts in the AI stack and across the cluster., * Demonstrated use of AI coding assistants and LLM-powered developer tools: daily user of AI agents and tools, * Partner with engineering peers, domain experts in adjacent teams, and business stakeholders to understand requirements and translate them into flexible, future-proof design solutions
- Hands on development, iteration, and maintenance of tools and applications that codify various aspects of large-scale AI cluster design stages and cluster deployment activities
- Design and development of cohesive interface code between disparate third party tools
- Own features from requirements and design through deployment and ongoing maintenance
- Work in an iterative software environment, including planning and delivering work in small increments, collaborating with stakeholders, often in different areas of domain expertise (Agile development practices)
- Participate in code reviews and retros; adapt to changing requirements and priorities, AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here.
Requirements
-
Five or more years of professional software development experience, including substantial experience building and supporting web applications
-
Proficiency in modern frontend development using JavaScript or TypeScript and a framework such as React, Angular, or Vue
-
Experience developing backend services and APIs using a modern server-side language or framework
-
Experience deploying and operating applications using Linux, containers, CI/CD, and cloud or on-premises infrastructure
-
Experience with automated testing, source control, code review, debugging, and production support
-
Working knowledge of web application security, authentication, authorization, and secure secrets handling, * Strong Linux fundamentals: Linux operating systems, networking, filesystems, containers, performance tooling (perf, flamegraphs, nvprof/rocprof, basic eBPF).
-
Clear communication: ability to turn complex systems into accessible, structured documentation with diagrams and reproducible steps
-
AMD ecosystem experience: ROCm, RCCL, Instinct GPUs, EPYC platforms, compiler/toolchain impacts, and performance tuning
-
Orchestration models: Slurm configuration patterns, Kubernetes for HPC/AI (GPU operators, device plugins), Apptainer/Singularity
-
Automation, IaC , and scripting tools/languages (Ansible, Terraform, Python, bash)
-
Storage/data: knowledge of or familiarity with parallel filesystems (Lustre, BeeGFS), object stores, RDMA, data pipeline throughput and caching strategies
-
Hands-on familiarity with on-premises infrastructure, particularly for AI/ML/HPC workloads would be beneficial
ACADEMIC CREDENTIALS:
- Bachelors or Masters degree in computer science or software/computer engineering