GPU Performance Engineer
Genmo Inc.
San Francisco, CA, United States
4 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on startup.jobs
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Languages
English, Spanish
Job source
Tech stack
C++ (Programming Language)
Profiling
Nvidia CUDA
Software Debugging
InfiniBand
Python (Programming Language)
Linux Kernel
Remote Direct Memory Access
System Programming
Graphics Processing Unit (GPU)
Syntactically Awesome Style Sheets (SASS)
Information Technology
Job description
You’ll be our performance optimization expert, using advanced profiling tools to identify bottlenecks and implementing solutions that achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you’ll ensure our infrastructure delivers world-class performance. This role is perfect for someone who gets excited about microsecond optimizations and pushing hardware to its theoretical limits., * Profile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentation
- Write high-performance CUDA and Triton kernels for critical model operations
- Optimize cold start latency from seconds to milliseconds for our serving infrastructure
- Tune memory access patterns, kernel fusion, and GPU utilization
- Collaborate with ML engineers to optimize model implementations
- Debug performance issues across the full stack from application to hardware
- Implement custom memory pooling and allocation strategies
- Share optimization techniques and build performance culture across teams
Requirements
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or related field
- 5+ years systems programming experience with 3+ years focused on GPU optimization
- Expert proficiency with GPU profiling tools (Nsight Systems, nvprof)
- Strong CUDA programming skills with production kernel development
- Deep understanding of GPU architecture (memory hierarchy, SMs, warps)
- Track record of achieving significant performance improvements (5-10x)
- Experience with Python and C++ in production environments
We Value
- Experience with Triton kernel development
- Knowledge of CUTLASS or similar high-performance libraries
- Background in ML-specific optimizations (attention, transformers)
- RDMA/InfiniBand optimization experience
- Contributions to GPU libraries or frameworks
- Low-level debugging skills (PTX/SASS reading), Genmo is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law. Genmo, Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on startup.jobs
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
DA
Dr. Andy R. Terrel - NVIDIA
over 1 year ago
DC
Daniel Cranney
Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs
over 1 year ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
CH
Chris Heilmann
Dev Digest 102 - Race conditions
over 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
CH
Chris Heilmann
Dev Digest 121 - AI goes offline
over 2 years ago