> Markdown version of [/jobs/ext/1903630-staff-principal-machine-learning-scientist-ai-inference-optimization-organization](https://www.wearedevelopers.com/jobs/ext/1903630-staff-principal-machine-learning-scientist-ai-inference-optimization-organization). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff / Principal Machine Learning Scientist, AI Inference & Optimization Organization - **Company:** Netskope - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $182,500.0 - $260,500.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Memory Management, Python (Programming Language), Machine Learning, Performance Tuning, Large Language Models, Backend, Information Technology, ONNX (Open Neural Network Exchange) Format, Hardware Acceleration, TensorRT - **Published:** July 31, 2026 - **Apply:** https://www.careerbuilder.com/job-details/senior-staff-principal-machine-learning-scientist-ai-inference-optimization-santa-clara-ca--857254b8-a4e5-4cf3-88e7-2b2252e0ce75 ## About the Role * 10+ years of overall industry experience, with 4+ years hands-on in ML/AI (model development, fine-tuning, and inference optimization). * Hands-on with fine-tuning (e.g. LoRA/QLoRA), quantization (GGUF/AWQ/GPTQ), and inference runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime, llama.cpp, or MLX/CoreML). On-device or edge inference experience is a strong plus. * Strong Python; comfort reaching into C++ for low-level interop is a plus. * Solid grasp of transformer internals and the levers that move real inference performance and cost: KV cache, attention, batching, memory footprint. * Fluency with agentic coding systems and genuine curiosity about agent harnesses like Claude Code, Pi, and Codex, so you should already be building with them, or itching to. * Clear communication: able to distill a model or infra bottleneck into an actionable concept for cross-functional teammates. Education * MS in Computer Science, Machine Learning, Electrical Engineering, or equivalent technical degree required, with a focus in AI/ML research; PhD in a related field strongly preferred. ## Description As a Senior Staff Machine Learning Scientist, you own the inference and optimization layer that makes AI in agentic workflows fast, efficient, and production-grade. You fine-tune and evaluate models, push latency and throughput on real hardware, and build the runtime that executes bounded AI tasks, validated against usage from Netskope's large customer base so you optimize where the data points, not where you guess. What's in it for you * High-impact ownership. You own the model layer of a net-new product that changes the performance and economics of agentic AI. * Cutting-edge, unusual stack. The hard, interesting inference problems live here: quantization, KV-cache and memory management, sparsity, fine-tuning, and hardware acceleration under real-world resource constraints. * Real scale to build against. Netskope's customer footprint gives you production signals most teams never see, so you deploy, validate, and iterate fast. What you will be doing * Build and optimize the model inference path: quantization, KV-cache optimization, batching, and latency/memory/throughput tuning on constrained, commodity hardware. * Fine-tune and evaluate models for bounded tasks; build eval harnesses that gate a capability to release on real accuracy, latency, and security relevance. * Design and grow the task execution runtime (bounded sub-agents), pushing toward dynamic task generation and context compaction. * Drive hardware acceleration / sparsity and support for larger models as the platform matures. * Partner with the systems and backend engineers to ship capabilities end-to-end and iterate on real production signals. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Making neural networks portable with ONNX](https://www.wearedevelopers.com/videos/301-making-neural-networks-portable-with-onnx) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)