> Markdown version of [/jobs/ext/1317940-ai-systems-performance-engineer-new-graduate](https://www.wearedevelopers.com/jobs/ext/1317940-ai-systems-performance-engineer-new-graduate). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Systems Performance Engineer - New Graduate - **Company:** The Eeo - **Location:** San Jose, CA, United States - **Experience:** Starter - **Salary:** $135,000.0 - $165,000.0 - **Contract:** Internship / Graduate position - **Skills:** Artificial Intelligence, C++ (Programming Language), Compilers, Program Optimization, Nvidia CUDA, Computer Programming, Computer Engineering, Data Structures, Distributed Systems, Data Flow Control, Generalized Linear Model, Python (Programming Language), Machine Learning, OpenCL, Performance Tuning, Tensorflow, Software Systems, High Performance Computing, Pytorch, Large Language Models, Deep Learning, Caching, Parallel Computation, Information Technology, Low Latency, Free and Open-Source Software, Machine Learning Operations, TensorRT, GPT, Programming Languages - **Published:** July 17, 2026 - **Apply:** https://www.dice.com/job-detail/2efa62b1-997d-488b-a319-c3dbd16403c1 ## About the Role * Bachelor's or Master's degree in computer science, electrical engineering, computer engineering, or a related technical field (e.g., applied mathematics, physics, or statistics), completed or expected before the start date. * Strong programming skills in Python, C++, or a similar programming language. * Solid foundations in algorithms, data structures, computer architecture, operating systems, or parallel computing. * Familiarity with deep learning and at least one major ML framework, such as PyTorch, TensorFlow, or JAX. * Strong analytical and problem-solving skills, with an interest in understanding and optimizing system performance. * Ability and enthusiasm to learn across machine learning, software systems, and hardware., * Coursework, research, internship, or project experience in machine learning systems, computer architecture, compilers, distributed systems, or high-performance computing. * Hands-on experience with LLMs, multimodal models, or transformer architectures. * Familiarity with model inference, KV cache, batching, quantization, or distributed execution. * Experience with GPU or accelerator programming using CUDA, Triton, OpenCL, or similar technologies. * Familiarity with frameworks such as vLLM, DeepSpeed, Megatron, or TensorRT. * Understanding of memory hierarchy, caching, parallelism, or scheduling. * Experience profiling and optimizing the performance of software or ML workloads. * Research publications, open-source contributions, programming competitions, or technically challenging personal projects are a plus. We value strong technical fundamentals, curiosity, and the ability to learn quickly. Prior production experience with large-scale AI systems is not required. ## Description We are seeking a talented and highly motivated AI Systems Performance Engineer to bring up and optimize state-of-the-art foundation models on SambaNova's reconfigurable dataflow platform. You'll work hands-on with advanced AI models - such as DeepSeek, GLM, Kimi, GPT OSS, Llama, Qwen, and other frontier architectures - and learn how modern AI systems achieve high throughput, low latency, and efficient large-scale inference. In this role, you'll work at the intersection of machine learning and computer systems, collaborating with engineers across model, compiler, runtime, and hardware teams. This is an ideal opportunity for a new graduate who is passionate about understanding how AI models execute on real hardware and wants to help build the next generation of high-performance AI systems. Responsibilities * Bring up cutting-edge foundation models, including LLMs and multimodal models, on the SambaNova platform through the SambaNova software stack. * Analyze and profile model execution to identify performance bottlenecks across model, compiler, runtime, and hardware layers. * Optimize AI workloads for throughput, latency, memory efficiency, and scalability. * Collaborate with machine learning, compiler, runtime, and hardware engineers to develop high-performance AI applications. * Explore and integrate new techniques in model architecture, quantization, scheduling, caching, and memory optimization. * Develop tools, benchmarks, and performance analysis methodologies for large-scale AI inference. * Investigate new model architectures and translate research advances into efficient implementations on production AI systems. * Contribute ideas for dataflow, scheduling, and system optimizations for both single-node and distributed inference. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Event based cache invalidation in GraphQL](https://www.wearedevelopers.com/videos/433-event-based-cache-invalidation-in-graphql) - [Streaming AI Responses in Real-Time with SSE in Next.js & NestJS](https://www.wearedevelopers.com/videos/1630-streaming-ai-responses-in-real-time-with-sse-in-next-js-nestjs) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts)