> Markdown version of [/jobs/ext/2114964-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/2114964-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - **Company:** UST Inc - **Location:** Portland, OR, United States - **Salary:** $82,000.0 - $123,000.0 - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Code Generation, Program Optimization, Computer Programming, Extract Transform Load (ETL), Software Debugging, Linux, Middleware, Github, Python (Programming Language), Unix Shell, Machine Learning, OpenCL, GitHub Copilot, Large Language Models, Multi-Agent Systems, Model Validation, Low Latency, ONNX (Open Neural Network Exchange) Format, Decoding, Docker - **Published:** August 19, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/88066959/1 ## About the Role * Hands-on development of multi-agent workloads * Experience with agent frameworks such as LangGraph/LangChain, AutoGen, CrewAI etc. * Models & Inference * Strong understanding of LLMs, SLMs and multimodal models, including Transformer architecture, attention, tokenization, context windows, and KV cache * Hands-on experience with model selection, evaluation, and deployment for different agentic workload requirements * Understanding of model formats and optimization, including ONNX, Open VINO IR, safe tensors, and quantization such as FP16/BF16/INT8/INT4 * Experience with inference engines/runtime frameworks such as Open VINO, ONNX Runtime, vLLM, llama.cpp or TGI * Understanding of prefill vs. decode, batching, continuous batching, speculative decoding, KV-cache management, and memory optimization * Understanding of CPU/GPU model execution, device placement, and heterogeneous inference * Ability to benchmark and compare different models, inference engines and runtime configurations and identify performance bottlenecks * Accelerator / Compute Stack * Understanding of GPU memory, kernel execution, synchronization, device selection and host/device data movement * Familiarity with middleware / accelerator compute runtimes such as OMIX/OneAPI/SYCL, Level Zero or OpenCL * Programming & Development * Strong Python * Working knowledge of C/C++ * Git/GitHub, Linux shell and debugging tools * Practical use of GitHub Copilot for development, debugging and code generation * Docker/container fundamentals * Performance Engineering * Ability to benchmark and profile AI workloads * Understanding of latency, throughput, tokens/sec, GPU utilization, memory bandwidth and CPU/GPU bottlenecks * Ability to identify whether a performance issue originates in the agent, model, inference engine, runtime or driver ## Description * Develop a multi-agent workload at the application level and understand, profile, and optimize its execution all the way down to the inference runtime, OMIX/middleware, compute runtime, Linux GPU driver, and accelerator hardware. * Practical use of GitHub Copilot for development, debugging and code generation * Optimize for latency, throughput, tokens/sec, memory footprint, and accelerator utilization * Select and tune models/inference engines based on agent workload characteristics, hardware capabilities, and performance requirements This position description identifies the responsibilities and tasks typically associated with the performance of the position. Other relevant essential functions may be required. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Building AI Applications with LangChain and Node.js](https://www.wearedevelopers.com/videos/1512-building-ai-applications-with-langchain-and-node-js) ## Related Articles - [A 5-Step Open-Source Setup for Agentic Engineering](https://www.wearedevelopers.com/magazine/738-a-5-step-open-source-setup-for-agentic-engineering) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)