> Markdown version of [/jobs/ext/3543027-software-engineer](https://www.wearedevelopers.com/jobs/ext/3543027-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - **Company:** Cisco Systems, Inc. - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $149,100.0 - $282,000.0 - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Application Programming Interfaces (APIs), Artificial Intelligence, Application Lifecycle Management, Unit Testing, Software Documentation, Code Review, Encodings, Nvidia CUDA, Information Systems, Disaster Recovery, Distributed Computing Environment, Github, Monitoring of Systems, Python (Programming Language), Linux System Administration, Load Testing, Nginx, Open Source Technology, Ansible, Prometheus, Runbook, Software Engineering, AI Infrastructure, QLoRA, Load Balancing, Pytorch, DeepSpeed, Retrieval-Augmented Generation, Transfer Learning, Large Language Models, Grafana, Distributed Inference, Kubernetes, Information Technology, Low Latency, HuggingFace, ROCm, Machine Learning Operations, Mixed Precision, DeepEval, Nim (Programming Language), Api Gateway, Terraform, Splunk, Cisco - **Published:** October 1, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=6fb706820b4a71cf ## About the Role * 7 or more years of related systems, platform, or software engineering experience, or equivalent practical experience, with solid knowledge across related technologies * A production background developing and operating AI infrastructure * A solid understanding of LLM, SLM, embedding, and reranker model internals, including context length, batching, token throughput, and memory use * Hands on work serving models on inference runtimes including vLLM, NVIDIA NIM, or Triton * Proven results building services around models, including the APIs, gateways, and operational tooling that make them consumable * Depth in evaluating and benchmarking models, and using the results to make release decisions * Familiarity moving model artifacts through evaluation, optimization, packaging, and production serving * Deep understanding of supervised fine tuning, parameter efficient fine tuning, distillation, and quantization * Command of distributed GPU training concepts, mixed precision, and parallelism strategies * Production Kubernetes work running GPU workloads at scale * Linux administration and troubleshooting * Programming in Python or Go * Fluency with CI/CD pipelines and Infrastructure as Code, for example Terraform or Ansible * mastery of monitoring and observability tooling, including Prometheus, Grafana, or Splunk * A track record of mentoring engineers and setting technical standards * A demonstrated pattern of learning new systems and the initiative to lead unfamiliar work * Ability to take part in an on call rotation for a service the company depends on * Clear written communication and the habit of documenting what you develop Preferred Knowledge and Experience * Work with AMD accelerators and ROCm alongside NVIDIA Cuda * Exposure to distributed inference, disaggregated serving, or KV cache aware routing * Familiarity with evaluation frameworks including lm-eval or DeepEval, and with safety and capability suites * A background evaluating RAG or agent systems * Time spent with API gateways, ingress, or load balancing, for example Envoy, APISIX, or NGINX * Fluency with GitOps and Helm, particularly ArgoCD * Capacity planning, traffic pattern understanding, and cost optimization for GPU fleets * Disaster recovery for stateful platform services * Contributions to open source inference or evaluation projects * Certified Kubernetes Administrator (CKA) or an equivalent cloud certification * Training or fine tuning transformer models with PyTorch and Hugging Face * Practical use of LoRA, QLoRA, and PEFT * Distributed training with PyTorch FSDP, DeepSpeed, or comparable frameworks * Model registries and experiment tracking systems, for example MLflow or Weights and Biases * Multi node GPU training and collective communication libraries Education * Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent practical experience ## Description As a Senior Software Engineer, you'll guide model serving, runtime tuning, accelerator optimization, and release evaluation. You'll lead incident response and improve reliability, performance, and operability. You'll also use agentic workflows to identify problems earlier and automate remediation. You'll follow Cisco Design Thinking Principles, simplify user experience, apply secure coding practices, and protect privacy. You'll work across design, product, and engineering to improve customer solutions, documentation, development practices, and production reliability. This role calls for production experience in AI infrastructure, model serving, evaluation, and related services. We're looking for a self starter who works independently, turns ambiguous problems into plans, mentors engineers, and raises technical standards. The platform evolves with new runtimes, accelerators, and models. On call is shared across the team. Responsibilities * Contribute to model serving direction and roadmap, including runtime selection and tuning across vLLM, NVIDIA NIM, and other runtimes. * Guide quantization and accelerator optimization across GPU vendors, validating performance, quality, capacity, and cost with data * Develop and enhance platform services, APIs, gateways, and operational tooling around model inference * Evolve evaluation, benchmarking, and load test frameworks that gate model releases against service level objectives * Define model promotion criteria across quality, safety, latency, throughput, and resource use * Evaluate how fine tuning, distillation, and quantization affect production behavior * Shape routing and capacity behavior, including prefix caching, KV aware routing, and prefill and decode separation * Improve model registry, packaging, evaluation, release, and development workflows using Infrastructure as Code, GitHub Actions, and agentic workflows * Refine model artifacts when runtime compatibility or performance requires changes * Develop observability that shows platform health, model performance, capacity, customer adoption, and usage * Monitor production, serve as an escalation point for on call issues, lead postmortems and root cause analyses, and drive durable improvements * Coordinate across customer, product, design, and engineering teams to gather input, forecast capacity, track milestones, and guide platform direction * Apply AI to platform operations through anomaly detection, automated remediation, and predictive operations * Lead features and projects from technical design through completion, working with minimal guidance and driving results through delegation and review * Write clean code and unit tests independently, and review code for quality, threat models, scale, reliability, and release velocity * Act as a technical resource, mentor engineers, run design reviews, and share knowledge across teams * Create technical designs, runbooks, user documentation, project updates, and remediation plans ## Related Videos - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [Computer Vision from the Edge to the Cloud done easy](https://www.wearedevelopers.com/videos/263-computer-vision-from-the-edge-to-the-cloud-done-easy) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Your Infrastructure Is Not a Playground: AI Agents for Infra Done Right](https://www.wearedevelopers.com/videos/2084-your-infrastructure-is-not-a-playground-ai-agents-for-infra-done-right) ## Related Articles - [ I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose.](https://www.wearedevelopers.com/magazine/773-i-gave-a-video-editor-more-autonomy-than-a-trading-bot-on-purpose) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)