> Markdown version of [/jobs/ext/3456373-ai-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/3456373-ai-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Infrastructure Engineer - **Company:** Netpreme Corporation - **Location:** Boston, MA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Profiling, Nvidia CUDA, Computer Engineering, Python (Programming Language), Performance Tuning, AI Infrastructure, Pytorch, Large Language Models, Caching, Kubernetes, Information Technology, Machine Learning Operations, TensorRT - **Published:** September 3, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=6df871d6756fe5f9 ## About the Role * BS, MS, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience. * 2+ years of relevant experience in LLM inference, ML systems, GPU systems, or performance engineering. * Must have: hands-on experience deploying and performance-tuning vLLM and/or SGLang. * Strong understanding of LLM inference fundamentals, including prefill vs. decode, batching, KV cache, latency/throughput trade-offs, and distributed GPU execution. * Strong Python engineering skills. * Working knowledge of inference-serving concepts: continuous batching, KV cache handling, quantization, and serving SLAs. * Clear written and verbal communication skills to work effectively with a small, fully distributed team. [Preferred Qualifications (optional)] * Contributions to vLLM, SGLang, FlashInfer, TensorRT-LLM, etc. * Experience with MoE / long-context model deployment. * Experience with speculative decoding, prefix caching, P/D disaggregation, attention/KV optimization. * Experience with Nsight Systems / PyTorch Profiler. * Familiarity with Kubernetes / production GPU serving. * Previous startup experience. ## Description * Deploy and optimize large language and multimodal models using vLLM and SGLang or other inference engines. * Design and evaluate TP/EP/DP/PP and hybrid parallelism strategies across GPU systems. * Build reproducible benchmarks to evaluate TTFT, TPOT, throughput, concurrency scaling, GPU utilization, and memory utilization. * Analyze model architecture and its serving implications, including attention, KV cache, MoE, long context, and speculative decoding. * Tune vLLM and SGLang configurations such as continuous batching, max batched tokens, chunked prefill, prefix caching, KV-cache precision/capacity, speculative decoding, CUDA Graphs, and P/D disaggregation. * Profile and diagnose bottlenecks across GPU compute, memory, communication, scheduling, and serving runtime. * Compare deployment configurations and identify production operating points balancing latency, throughput, capacity, and stability. * Work with model/system engineers to bring newly released models into production efficiently. * Collaborate closely with our hardware/systems team (direct access to CTO-level technical leadership on a small team) to translate performance requirements into backend architecture decisions. * Contribute to defining next-generation benchmarks and service requirements as workloads evolve - multi-turn coding, agentic pipelines, RAG, and other long-context use cases.