Machine Learning Research Engineer (LLMs & AI Systems)

Tenstorrent Usa, Inc.
Boston, MA, United States
9 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$100,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Distributed Computing Environment Python (Programming Language) Machine Learning Pytorch Large Language Models Deep Learning Machine Learning Operations

Job description

  • Lead research and development efforts focused on LLM training and inference optimization.
  • Train, evaluate, and optimize state-of-the-art AI models on Tenstorrent hardware.
  • Improve performance through techniques such as speculative decoding, quantization, kernel fusion, flash attention, and distributed training.
  • Investigate system bottlenecks and collaborate cross-functionally to drive performance improvements.
  • Translate cutting-edge ML research into scalable, production-ready solutions.

What You Will Learn

  • How to optimize AI models on custom AI accelerators from application to silicon.
  • How large-scale ML systems are deployed, tuned, and scaled in production.
  • How hardware, compiler, kernel, and ML teams collaborate to maximize performance.
  • The challenges and tradeoffs of scaling modern AI workloads across custom hardware.

Requirements

  • Strong Python and PyTorch experience developing and training deep learning models.
  • Deep understanding of ML architectures, LLM training, and inference optimization.
  • Hands-on experience training large-scale machine learning models.
  • 4+ years of industry and/or academic experience in ML research and LLM development.
  • PhD, published research, or experience with speculative decoding is highly valued.

Benefits & conditions

Compensation for all engineers at Tenstorrent ranges from $100k - $500k including base and variable compensation targets. Experience, skills, education, background and location all impact the actual offer made.

Tenstorrent offers a highly competitive compensation package and benefits, and we are an equal opportunity employer.

About the company

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities.

Tenstorrent is building next-generation AI systems that push the boundaries of model training, inference, and large-scale distributed compute. The ML Models team sits at the intersection of cutting-edge AI research and high-performance hardware, bringing state-of-the-art machine learning models to life on Tenstorrent’s custom AI accelerators. From training large language models to optimizing inference performance at scale, this team works across the full stack to turn breakthrough research into production-ready AI systems. If you are passionate about advancing the frontier of AI research, inference and training optimizations, this is an opportunity to shape how future AI models are developed and deployed.

This role is hybrid, based out of Toronto, ON and Boston, MA.

We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Open-source community and machine learning frameworks

Gian Marco Iodice Gian Marco Iodice · World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

1:25 min

Distinguishing artificial intelligence from deep learning

Sam Witteveen · Coffee With Developers

2:36 min

Applying supervised machine learning for practical rule extraction

Katja Träumner

3:51 min

Overcoming hardware configuration barriers in machine learning

Jose Luis Latorre Millas · LIVE

2:08 min

Applying large language models to infrastructure tasks

Alfonso Sandoval Rosas Alfonso Sandoval Rosas · Europe 2026 Virtual

Videos

See all

Related articles

See all