> Markdown version of [/jobs/ext/2712467-software-engineer-infrastructure](https://www.wearedevelopers.com/jobs/ext/2712467-software-engineer-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Infrastructure - **Company:** Pal Inc. - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Nvidia CUDA, Uptime, Routing, Graphics Processing Unit (GPU), Large Language Models, Backend, Kubernetes, Video Streaming - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/software-engineer-infrastructure-tavus-io-8025970 ## About the Role * Hands-on GPU inference experience. You've deployed and optimized inference workloads on GPUs and know what it takes to build reliable systems on top of GPU cloud providers. * Kubernetes and EKS depth, including routing and scheduling. You're comfortable designing how work gets placed across a fleet, and writing the services that make it happen. * Deep AWS experience. You're at home spinning up new services and turning them into simple, repeatable processes others can build on. * A senior track record of ownership. You've set technical direction, made decisions others built on, and carried ambiguous work over the finish line. You explain complex ideas clearly, to engineers and non-engineers alike. Nice to have * Experience with GCP * Experience with video streaming infrastructure * Experience with training infrastructure or LLM serving * Experience with SOC2 or security compliance ## Description We're hiring a Senior Software Engineer (Infrastructure) to own the systems behind CVI, our real-time conversational product. Every live conversation between a person and a PAL runs on infrastructure your team owns. You'll take goals like uptime, latency, and cost and chase them wherever they lead, including into backend services and product code., * CVI's inference deployments. The GPU infrastructure serving live conversations across multiple providers and regions. You'll join as an early senior member of a growing infra team, working on projects like tuning the newest GPU generations and cutting cold-start and model load times so users wait less. * Expanding our GPU footprint. You'll bring on new providers and regions, stand up clusters on EKS, and build the routing, scheduling, and throughput needed for fast weight loading. * Uptime. You'll be one of the people pushing our uptime bar higher, along with the security and SOC2 work that keeps our infrastructure trustworthy. * Fix what you find. When you see a problem, you have the trust and the mandate to fix it or flag it. Reworking our deploy pipeline so shipping is fast and boring is exactly the kind of thing you'd take on., * Multi-provider, multi-region inference infrastructure: routes live conversations across GPU providers and regions, so one provider's outage never becomes a user's problem * CUDA optimizations for Phoenix, our video rendering model: doubled the frame rate by tracing and optimizing hot paths with our researchers * Parallel conversations on a single GPU: several live conversations sharing one card, multiplying what the fleet can serve, * You're energized by unfamiliar problems. If the next thing that matters is standing up a training deployment you've never touched, you jump in and learn on the fly. * You adapt as priorities evolve. In a space moving this fast, the most important thing to build can change as we learn. When it does, you adjust course without losing momentum. * You care about this problem. Keeping large-scale, real-time systems fast and reliable is something you think about unprompted. ## Related Videos - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [WWC24 - Ankit Patel - Unlocking the Future Breakthrough Application Performance and Capabilities with NVIDIA](https://www.wearedevelopers.com/videos/920-wwc24-ankit-patel-unlocking-the-future-breakthrough-application-performance-and-capabilities-with-nvidia) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data)