> Markdown version of [/jobs/48420-llm-training-engineer](https://www.wearedevelopers.com/jobs/48420-llm-training-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # LLM Training Engineer - **Company:** Sciforium - **Location:** San Francisco, United States - **Experience:** Expert - **Salary:** €155,000.0 - €220,000.0 - **Skills:** Python - **Published:** August 21, 2026 - **Apply:** https://jobs.ashbyhq.com/Sciforium/2c72c7ad-9da9-4738-af58-2650d22ec6de ## About the Role Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications. ## About the Role As an **LLM Training Engineer** , you’ll work across the full foundation-model stack: **pretraining and scaling** , **post-training and Reinforcement Learning** , **sandbox environments for evaluation and agentic learning** , and **deployment + inference optimization**. You’ll build and iterate quickly on research ideas, contribute production-grade infrastructure, and help deliver models that can serve real-world use cases at scale. ## **What you’ll work on** This role spans multiple tracks - candidates may focus on one or contribute across several. Examples include: Pretraining & Scaling - Train large byte-native foundation models across massive, heterogeneous corpora - Design stable training recipes and scaling laws for novel architectures - Improve throughput, memory efficiency, and utilization on large GPU clusters - Build and maintain distributed training infrastructure and fault-tolerant pipelines Post-training & RL - Develop post-training pipelines (SFT, preference optimization, RLHF/RLAIF, RL) - Curate and generate targeted datasets to improve specific model capabilities - Build reward models and evaluation frameworks to drive iterative improvement - Explore inference-time learning and compute techniques to enhance performance Sandbox Environments & Evaluation - Build scalable sandbox environments for agent evaluation and learning - Create realistic, high-signal automated evals for reasoning, tool use, and safety - Design offline + online environments that support RL-style training at scale - Instrument environments for observability, reproducibility, and iteration speed Deployment & Inference Optimization - Optimize inference throughput/latency for byte-native architectures - Build high-performance serving pipelines (KV caching, batching, quantization, etc.) - Improve end-to-end model efficiency, cost, and reliability in production - Profile and optimize GPU kernels, runtime bottlenecks, and memory behavior ## Description ## **Ideal candidate credentials** Technical strength - Strong general software engineering skills (writing robust, performant systems) - Experience with training or serving large neural networks (LLMs or similar) - Solid grasp of deep learning fundamentals and modern literature - Comfort working in high-performance environments (GPU, distributed systems, etc.) Relevant experience (one or more) - Pretraining / large-scale distributed training (FSDP/ZeRO/Megatron-style systems) - Post-training pipelines (SFT, RLHF/RLAIF, preference optimization, eval loops) - Building RL environments, simulators, or agent frameworks - Inference optimization, model compression, quantization, kernel-level profiling - Building large ETL pipelines for internet-scale data ingestion and cleaning - Owning end-to-end production ML systems with monitoring and reliability Research orientation - Ability to propose and evaluate research ideas quickly - Strong experimental hygiene: ablations, metrics, reproducibility, analysis - Bias toward building — you can turn ideas into working code and results Education - MS or PhD in Computer Science, Machine Learning, AI, Mathematics, or related field ## About Sciforium [Company profile](https://www.wearedevelopers.com/companies/4292-sciforium) ### More Jobs at Sciforium - [LLM Dataset Engineer](https://www.wearedevelopers.com/jobs/48419-llm-dataset-engineer) - [Model Implementation Engineer](https://www.wearedevelopers.com/jobs/48421-model-implementation-engineer) - [Senior AI Serving Engineer, Backend](https://www.wearedevelopers.com/jobs/48414-senior-ai-serving-engineer-backend) - [GPU Kernel Engineer](https://www.wearedevelopers.com/jobs/48412-gpu-kernel-engineer) - [ML Engineer](https://www.wearedevelopers.com/jobs/48422-ml-engineer) ## Related Videos - [CUDA in Python](https://www.wearedevelopers.com/videos/1294-cuda-in-python) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Full Stack Web Apps With Nothing But Python](https://www.wearedevelopers.com/videos/417-full-stack-web-apps-with-nothing-but-python) - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/videos/1522-adding-knowledge-to-open-source-llms) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Automagic Configuration in Python](https://www.wearedevelopers.com/videos/363-automagic-configuration-in-python) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)