> Markdown version of [/jobs/ext/1233731-software-engineer-runtime](https://www.wearedevelopers.com/jobs/ext/1233731-software-engineer-runtime). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Runtime - **Company:** Ollama USA Inc. - **Location:** Palo Alto, CA, United States - **Contract:** Permanent contract - **Skills:** Microsoft Windows, Apple Mac Systems, C++ (Programming Language), Nvidia CUDA, Databases, Linux, Memory Management, Game Engine, Open Source Technology, Graphics Processing Unit (GPU), Gpu Programming, Backend - **Published:** July 11, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=5c5b6010a6596a07 ## About the Role * You have strong systems fundamentals and are comfortable in Go, C, or C++ * You've worked close to the metal - GPU compute, inference, game engines, databases, OS, or networking. * You care about performance and have profiled and optimized real workloads. * You're comfortable shipping to millions of users and handling the long tail of hardware and OS combinations. * Bonus: experience with model quantization, GPU programming (CUDA/Metal/SYCL), Apple MLX ## Description You'll work on the heart of Ollama - the local runtime that runs open models on developers' own machines. It loads models, manages memory, drives GPU acceleration across NVIDIA, AMD, Intel, Qualcomm, and Apple Silicon (including our MLX integration), and makes all of it feel instant. You'll work in Go and C/C++ and touch the model formats and inference engines underneath, shipping to macOS, Linux, and Windows across an enormous range of hardware., * Make open models run fast and reliably on consumer and enterprise hardware - from a MacBook Pro to server-grade NVIDIA GPUs. * Own pieces of the runtime: model loading & scheduling memory management, quantization, GPU hardware backends. * Integrate new model architectures and quantization formats so the latest open models work on day one. * Improve cold-start, time-to-first-token, and throughput * Partner with model labs and hardware vendors on early access and deep integrations. * Ship in the open: Ollama is open source, and you'll work with the community Example projects * Add support for a new model family end-to-end - format parsing, weights loading, and the defaults that make it useful out of the box. * Cut cold-start for a popular model in half by streaming weights and lazy-loading layers. * Land a new quantization format so a 70B model runs on a single consumer GPU. * Wire up a new GPU backend and find a 2x throughput win with kernel selection and memory tuning. * Improve the "Auto" experience - picking the right model and settings for a machine's hardware without the user thinking about it. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 102 - Race conditions](https://www.wearedevelopers.com/magazine/386-dev-digest-102-race-conditions) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)