Software Engineer, AI Inference Platform
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+7 more
Job description
- Break down LLM and transformer workloads into fine-grained primitives tailored to our proprietary compute hardware.
- Design and implement IR transformations, graph optimizations, kernel lowering, and code generation for novel hardware architectures.
- Collaborate with ML researchers to co-design algorithmic optimizations that yield real end-to-end performance gains.
- Work closely with hardware architects to refine microarchitectural features, instruction sets, memory hierarchies, and execution models.
- Build performance models, profiling tools, and benchmarking frameworks to identify bottlenecks and guide design decisions.
- Prototype and validate improvements across the entire stack - from PyTorch/XLA-level passes to custom kernel implementations.
- Contribute to shaping the overall system architecture of a first-of-its-kind inference engine.
Requirements
- BS/MS/PhD in Computer Science, Software Engineering, or a related field.
- Deep experience building compilers, optimizing kernels, or working with ML frameworks at a systems level.
- Strong proficiency in one or more programming languages such as Python and C++.
- Strong understanding of one or more of the following:
- LLM architectures and transformer internals
- MLIR, LLVM, XLA, TVM, Triton, or similar compiler infrastructures
- GPU/TPU/FPGA/ASIC compute models, memory hierarchies, and parallel execution
- Quantization, sparsity, or algorithmic optimization for deep learning
- Deep expertise on ML frameworks (e.g., PyTorch, TensorFlow, JAX) and understanding of ML model deployment challenges.
- Solid understanding of software engineering best practices, including data structures, algorithms, and testing.
- Thinking in terms of latency, cycles, memory bandwidth, and arithmetic intensity, not just algorithms.
- Excellent problem-solving abilities and a knack for tackling complex technical challenges.
- Excited to collaborate across ML, hardware, and software boundaries to invent something fundamentally new.
- Strong communication skills and a proven ability to collaborate effectively in a cross-functional team environment.
- Ability to thrive in a fast-paced, dynamic startup environment., * PhD in Computer Science, Software Engineering, or a related field.
- Experience with custom hardware accelerators for ML inference.
- Contributions to open-source compiler or ML systems projects.
- Prior startup experience or background building first-generation systems.
Benefits & conditions
- A chance to be a foundational engineer in an innovative AI startup
- A dynamic and collaborative work environment and the change to have a significant impact on new technology
- The opportunity to work on challenging problems at the intersection of ML, software, and systems.
- Competitive compensation and startup equity package
- Comprehensive medical, dental, and vision coverage (100% paid by employer)
- Life insurance and AD&D
- Flexible Time Off (FTO)
- 12-paid holidays
- Paid parental leave
- Gym or fitness benefit
- Commuter benefit
- Weekly catered lunches in the office
- Investment in employee learning & development
About the company
ElastixAI is an early-stage startup on a mission to reinvent AI inference infrastructure from the ground up. We’re building a next-generation inference platform that delivers unprecedented efficiency by tightly integrating machine learning, software stack, and custom hardware. Our philosophy is simple: the best performance comes from holistic co-design, where every layer, from model architecture to kernels to silicon, works in harmony. If you’re excited about pushing AI performance to physical limits, and about shaping the future of large-scale inference, we’d love to meet you.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Stephan Gillich - Bringing AI Everywhere
MLOps And AI Driven Development
What Are Large Language Models?
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production