Mid-Level Machine Learning Engineer

TETRAMEM INC
San Jose, CA, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
5 years minimum
Compensation
$110,000.0 - $300,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Audio Signal Processing C++ (Programming Language) Field-Programmable Gate Array (FPGA) Python (Programming Language) Machine Learning Tensorflow Systems Architecture Application Specific Integrated Circuits Pytorch Information Technology Low Latency
+3 more
ONNX (Open Neural Network Exchange) Format Hardware Acceleration TensorRT

Job description

  • Develop, optimize, and deploy lightweight machine learning models for edge AI applications, particularly for audio processing.
  • Implement and optimize ML models on embedded platforms, including FPGA and custom ASIC solutions.
  • Work closely with hardware and software teams to integrate ML models into production systems.
  • Research and implement state-of-the-art ML techniques to enhance model efficiency, latency, and power consumption for embedded AI applications.
  • Improve inference efficiency and model compression techniques, including quantization, pruning, and knowledge distillation.
  • Collaborate with cross-functional teams to drive innovation and contribute to the overall system architecture.
  • Provide technical leadership and mentorship to junior engineers.
  • Publish research findings, present at conferences, and contribute to open-source projects when applicable.

Requirements

  • 5+ years of experience or PhD in Computer Science, Electrical Engineering, or related fields.
  • Strong experience in machine learning, with a focus on edge AI and lightweight model deployment.
  • Expertise in ML frameworks such as PyTorch, TensorFlow, JAX.
  • Proficiency in programming languages such as C/C++, Python, and experience with ML model optimization.
  • Ability to work independently and collaboratively in a fast-paced startup environment.

Experience in one or more of the following areas considered a strong plus:

  • Understanding of ML compiler and runtime design.
  • Experience working with tools such as Optimum, ONNX, TensorRT, TFLite/LiteRT, ncnn, or CoreML.
  • Familiarity with hardware acceleration techniques.
  • Experience in embedded system development.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Open-source community and machine learning frameworks

Gian Marco Iodice Gian Marco Iodice · WWC 2025

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · WWC 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · WWC 2024

Videos

See all

Related articles

See all