> Markdown version of [/jobs/ext/3536219-software-engineer](https://www.wearedevelopers.com/jobs/ext/3536219-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - **Company:** Cisco Systems, Inc. - **Location:** Richardson, TX, United States - **Salary:** $128,600.0 - $184,900.0 - **Contract:** Permanent contract - **Skills:** Nvidia CUDA, Information Systems, Software Debugging, Software Design Documents, Linux, Release Management, Secure Coding, Software Engineering, AI Infrastructure, Software Organization, Retrieval-Augmented Generation, Large Language Models, Distributed Inference, Agentic-AI, Kubernetes, Information Technology, ROCm, Nim (Programming Language), VLLM, Cisco - **Published:** October 1, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3413479421&tx=KP7575FFQ&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * Bachelor's degree in Computer Science, Information Systems, or a related field with 3 years of related experience; master's degree with 1 year of related experience; PhD; or equivalent practical work experience. * Experience developing and operating production services on Kubernetes and Linux, including exposure to GPU based AI infrastructure. * Hands on experience serving models with vLLM, NVIDIA NIM, Triton, or a comparable inference runtime, including building supporting services. * Experience leading technical work, mentoring engineers, documenting systems, and supporting critical services through an on call rotation. Preferred Qualifications: * Experience evaluating and benchmarking models to support production release decisions. * Experience operating GPU workloads with NVIDIA CUDA or AMD ROCm. * Familiarity with distributed inference, disaggregated serving, KV cache aware routing, capacity planning, or cost optimization. * Experience fine tuning transformer models or evaluating RAG and agent systems using relevant frameworks. ## Description We run the platform that serves foundational models to Cisco IT. Our Foundational Model Service, gives engineering teams across the company access to small, large language, and embedding models. youWe serve those models on Kubernetes clusters, with Nim, Vllm and other runtimes. Beyond serving, We benchmark, evaluate, monitor, and release new models as improvements and demand warrant. Our customers depend on the platform under a 99.9% uptime SLA, and we build and operate accordingly. As a Software Engineer on the FMS team you'll keep that platform running and make it easier to run. You'll support model onboarding and releases, take your turn on call, help automate the validation that gates every deployment, and contribute to the monitoring that shows us and our customers how the service is behaving. You'll also apply Agentic solutions to our own operations to catch problems earlier and remediate them automatically. Your Impact You'll develop software consistent with Cisco Design Thinking Principles, with simplification and user experience at its core, using secure coding practices, protecting user privacy, and following software development best practices. You'll partner with design, product management, and other engineering teams to build the right solution for our customers. You'll create technical design documentation for the team, contribute to the documentation end users rely on, and debug platform issues both in development and in production. This is a strong role for a systems or infrastructure engineer who has worked with LLMs, whether through academic coursework or by building and running them on the job, and who understands how inference serving works. We're looking for a self starter who takes on unfamiliar work and grows into it. You'll work with engineers who have built these systems from the ground up. New runtimes and new models arrive frequently, so the platform is always evolving, and you'll develop a deep understanding of developing AI infrastructure, model runtimes, and delivering products at scale. On call is shared across the team. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Nemotron: NVIDIA's open model strategy for developers](https://www.wearedevelopers.com/videos/100064-nemotron-nvidia-s-open-model-strategy-for-developers) - [Computer Vision from the Edge to the Cloud done easy](https://www.wearedevelopers.com/videos/263-computer-vision-from-the-edge-to-the-cloud-done-easy) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [ I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose.](https://www.wearedevelopers.com/magazine/773-i-gave-a-video-editor-more-autonomy-than-a-trading-bot-on-purpose) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)