> Markdown version of [/jobs/ext/733553-senior-software-engineer-inference](https://www.wearedevelopers.com/jobs/ext/733553-senior-software-engineer-inference). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, Inference - **Company:** PIKA, LLC - **Location:** Palo Alto, CA, United States - **Experience:** Expert - **Salary:** $185,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Code Review, Nvidia CUDA, Computer Programming, Distributed Computing Environment, Language Modeling, Rapid Prototyping Process, Software Deployment, Data Streaming, AI Infrastructure, Delivery Pipeline, Large Language Models, Deep Learning, Parallel Computation, Gpu Programming, AI Platforms, Free and Open-Source Software, Machine Learning Operations - **Published:** June 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=339ccaa6cc22446c ## About the Role Do you have experience in AI platforms (beyond public GPTs)?, * Experience: 5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale. * Inference Mastery: Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks. * GPU & Parallelism: Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference. * AI Domain Knowledge: Familiarity with video generation (videogen) models and large language models (LLMs). * Collaboration: Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions. * Ownership Mindset: Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment. * Bonus: Experience in enhancing training efficiency, stability, or resource optimization for large models., * Experience with high-throughput video or real-time streaming model deployment * Familiarity with distributed training and optimization toolkits * Contributions to open source projects in AI infrastructure or deep learning compilers * Startup or rapid prototyping experience ## Description We are seeking a Senior Inference Engineer to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale. You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what's possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika's video and language models., * Accelerate Inference: Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving. * Maximize GPU Parallelism: Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability. * Programming for Performance: Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL. * Advance AI Deployment: Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production. * Improve Training Efficiency: (Bonus) Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle. * Technical Excellence: Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming. ## Related Videos - [Developer Experience, Platform Engineering and AI powered Apps](https://www.wearedevelopers.com/videos/990-developer-experience-platform-engineering-and-ai-powered-apps) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [This App Reached 10,000 Users in One Week. Here's How.](https://www.wearedevelopers.com/videos/100329-this-app-reached-10-000-users-in-one-week-here-s-how) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [From AI Assistance to Agentic Systems: Scaling Sovereign AI in Banking](https://www.wearedevelopers.com/videos/100070-from-ai-assistance-to-agentic-systems-scaling-sovereign-ai-in-banking) - [Supercharge your cloud-native applications with Generative AI](https://www.wearedevelopers.com/videos/950-supercharge-your-cloud-native-applications-with-generative-ai) ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)