Neural Network Kernel Software Development Engineer

Targeted Talent
Seattle, WA, United States
2 days ago
Apply on www.juju.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Artificial Neural Networks C++ (Programming Language) Code Generation Computer Programming Python (Programming Language) Performance Tuning Software Engineering Information Technology Hardware Acceleration C++14
+1 more
Recurrent Neural Networks

Job description

Our client is making substantial investments in software to enhance the seamless deployment of neural networks on their hardware, streamlining the experience for researchers and developers. The focus involves the optimization of various common neural networks for optimal performance on architectures, facilitated by the software optimization tool flow.

We are seeking software developers who are driven and naturally curious. The chosen candidate will contribute within agile teams, working closely with senior software engineers for mentorship. This role presents an opportunity to tackle novel challenges using cutting-edge technologies, as they build innovative systems from scratch.

As a key team member, you will specialize in constructing efficient implementations of practical neural net kernels tailored to their distinctive hardware architecture. Additionally, you will implement diverse computing algorithms, maximizing computation and communication throughput. This role involves developing a profound understanding of the architecture’s intricacies, working collaboratively with the architects and compiler engineers., * Experience writing kernels to accelerate Neural Network execution on custom hardware accelerators (not on CPU’s)

  • Design, prototype, and execute low-level, adaptable C++ programs (kernels) for various neural net operations.
  • Define, document, and communicate configuration APIs for these kernels to the compiler team.
  • Share performance optimization concepts with both compiler engineers and architects working on future product generations.
  • Develop comprehensive computation strategies spanning kernels for multichannel and multi-chip neural net implementations.

Requirements

  • Degree in Computer Science, Engineering, Math, Physics, or related field (preferably MS or PhD).
  • Profound knowledge of modern C++, with a focus on code generation and low-level compute optimizations.
  • Familiarity with fundamental Neural Network operator algorithms - Convolutions, Transformers, RNNs.
  • Demonstrated capability to independently navigate challenging, well-defined problems.
  • Aptitude and interest in both high-level conceptual understanding and intricate technical details.
  • Enthusiasm for problem-solving within highly structured and restricted environments.

Preferred Skills and Experience:

  • Proficiency in Python.
  • Experience with other AI accelerator programming.
  • Strong mathematical aptitude.
  • Enjoyment of solving complex problems.

Benefits & conditions

  • Competitive Salary
  • Unlimited sick leave.
  • Stock options.
  • Contribution to revolutionizing chip and software technologies with global impact.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

8:32 min

Benchmarking GitOps engine constraints for extensive multi-cluster environments

Artem Lajko · Europe 2026 Virtual

2:22 min

Reviewing and troubleshooting code generated by AI assistants

Cassidy Williams Cassidy Williams · Coffee With Developers

2:01 min

Refactoring software applications with domain specific SDKs

Ankit Patel Ankit Patel · World Congress 2024

4:10 min

Masking system latency through strategic character design

Ben Hopkins +1 · Coffee With Developers

3:47 min

Understanding the trade-offs of automated code generation

Marco Podien Marco Podien · World Congress 2025

4:20 min

Utilizing AI and hardware acceleration for application code optimization

Stephan Gillich Stephan Gillich · World Congress 2024

Videos

See all

Related articles

See all