Unveiling the Magic: Scaling Large Language Models to Serve Millions
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Ditch the 20-minute cold starts. Treat self-hosted LLMs as compute-heavy REST APIs. Learn how NFS, Pingora, and intelligent rate-limiting seamlessly scale your infrastructure to serve millions.
Matching moments
More from World Congress 2025
Related videos
From learning to earning
Jobs that call for the skills explored in this talk.
about 1 month ago
•
Verified
LLM Training Engineer
Sciforium
San Francisco, United States
Expert
$155k–220k
Python
about 1 month ago
•
Verified
LLM Dataset Engineer
Sciforium
San Francisco, United States
Expert
$155k–210k
Python
about 1 month ago
•
Verified
Lead Software Engineer, Model Serving Platform
Sciforium
San Francisco, United States
Expert
$230k–300k
Python
about 1 month ago
•
Verified
Model Implementation Engineer
Sciforium
San Francisco, United States
Expert
$165k–220k
Python
about 1 month ago
•
Verified
Senior AI/ML Engineer
PagerDuty
Lisbon, Portugal
Expert
Remote
AI Frameworks
AI-assisted coding tools
about 1 month ago
•
Verified
Senior AI Serving Engineer, Backend
Sciforium
San Francisco, United States
Expert
$190k–250k
Rust
Python