> Markdown version of [/jobs/ext/1312692-ai-technical-lead](https://www.wearedevelopers.com/jobs/ext/1312692-ai-technical-lead). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Technical Lead - **Company:** NIO - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $192,100.0 - $249,600.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Apache HTTP Server, Computer Vision, C++ (Programming Language), Cloud Computing, Computer Engineering, Distributed Systems, Memory Management, Python (Programming Language), Linux Kernel, Open Source Technology, Large Language Models, Information Technology, ONNX (Open Neural Network Exchange) Format, Machine Learning Operations, TensorRT - **Published:** July 17, 2026 - **Apply:** https://www.dice.com/job-detail/57feda4a-7d3d-44a5-aac5-7614715f031d ## About the Role * Education & Experience: Ph.D. in Computer Science, Computer Engineering, Artificial Intelligence, or a related field with 8+ years of relevant industry experience (or Master's degree with 12+ years), including proven experience leading technical teams or driving complex architectural roadmaps. * End-to-End Systems Leadership (T-Shaped Profile): Demonstrated capability to lead full-stack AI systems engineering. You possess deep, hands-on mastery in at least one or two of the following core domains, coupled with the comprehensive systemic breadth required to effectively lead engineers working across the others: + Distributed Systems & Hybrid Inference: Designing, scaling, and deploying production-grade distributed ML systems. Balancing cloud infrastructure with edge constraints using modern routing paradigms, such as cascading inference architectures and semantic routing. + Algorithmic & Inference Optimization: Proven experience optimizing state-of-the-art LLM/VLM inference pipelines. Deep understanding of model compression (e.g., PTQ, QAT, AWQ, FP8/INT4), hardware-aware compute optimizations (e.g., FlashAttention), and advanced memory management (e.g., PagedAttention, KV cache compression/eviction). + Advanced Systems & Compiler Engineering: C++ and production-grade Python proficiency. Deep understanding of edge/cloud model-serving frameworks (e.g., vLLM, TensorRT-LLM, ExecuTorch, MLC-LLM) and AI compilers (e.g., MLIR, Apache TVM, Triton) for compute graph optimization and custom kernel development. Preferred Qualifications * Privacy & Security: Deep understanding of privacy-preserving AI techniques (federated learning, differential privacy, secure enclaves) essential for processing sensitive data across edge and cloud environments. * Community Engagement & Open Source: Publications in relevant AI, ML, or systems conferences (e.g., NeurIPS, ICML, MLSys), or active contributions to open-source ML infrastructure projects (e.g., vLLM, ONNX Runtime, Apache TVM, llama.cpp). ## Description * Architect the Hybrid AI Vision: Lead the architectural design and strategic vision for hybrid inference systems, dynamically distributing Large Language Model (LLM) and Vision-Language Model (VLM) workloads across edge computing environments and cloud infrastructure. * Team Leadership & Innovation: Lead, mentor, and inspire a team of specialized engineers working across distributed systems orchestration, inference optimization, and AI compiler engineering. While you are not expected to be a hands-on master of every domain, you will drive the overarching technical roadmap, foster a culture of cutting-edge innovation, and guide domain experts in navigating complex system tradeoffs. * Design Dynamic Orchestration & Resilience: Oversee the architecture of high-availability orchestration engines that intelligently route inference tasks. Guide the team in developing cascading inference mechanisms, dynamic model fallback strategies, and robust telemetry to ensure continuous, steady-state inference under varying connectivity constraints. ## Related Videos - [Focoos AI: Building the Future of Computer Vision](https://www.wearedevelopers.com/videos/1659-focoos-ai-building-the-future-of-computer-vision) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)