> Markdown version of [/jobs/ext/3540467-principal-software-engineering](https://www.wearedevelopers.com/jobs/ext/3540467-principal-software-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Software Engineering - **Company:** Cisco Systems, Inc. - **Location:** Milpitas, CA, United States - **Experience:** Expert - **Salary:** $250,600.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, BIOS, Computer Clusters, Computer Engineering, Continuous Integration, Dynamic Host Configuration Protocol, Ethernet, Firmware, InfiniBand, Python (Programming Language), Load Testing, Operational Data Store, Queueing Systems, Software Deployment, Software Engineering, Web Services, Workflow Management Systems, Model-Driven Development, High Performance Computing, Performance Testing, System Availability, Backend, Event Driven Architecture, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Slurm, Cisco - **Published:** October 2, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/18504504?backUrl=%2Fcareer%2F18504504%2FPrincipal-Software-Engineering-California-Milpitas ## About the Role * Bachelor's degree in computer science, computer engineering, electrical engineering, or a related field plus 15 years of related experience; a master's degree plus 12 years; a PhD plus 8 years; or equivalent related work experience. * Experience serving as an architect for at least one platform, from technical design and implementation through production deployment and adoption by [one or more teams. * Experience automating labs, networks, devices, or infrastructure at scale, and running that automation on Kubernetes or an equivalent control plane. * Experience identifying and resolving system-performance constraints using profiling, benchmarking, monitoring, or load testing, with documented improvements in areas such as latency, throughput, capacity, or resource utilization. * 10 + years of software-design and development experience using Python and at least one additional production language, such as Go, Java, or Rust. * Experience building production APIs, workflow or orchestration services, and the operational UI or dashboards engineers use to run them. * Experience delivering documented improvements in at least one area-such as system capacity, delivery cycle time, reliability, or user adoption-and coordinating implementation across engineering, operations, and customer-facing teams., * Workflow engines, versioned automation and event-driven architecture with message queues and operational data stores. * Device or service orchestration, model-driven automation, or scale testing of large managed environments. * Performance baselines, capacity testing, resilience testing, fault injection, or failure-mode analysis. * AI-assisted triage, diagnosis, or remediation validated in real environments. * Large GPU clusters, high-performance computing, or customer cluster bring-up. * Slurm or other HPC and AI workload schedulers. * PXE, DHCP, Redfish, firmware and BIOS automation, or infrastructure as code. * High-performance Ethernet, InfiniBand, or RoCE automation, and high-performance storage. * Cisco UCS, Cisco switching, NVIDIA platforms, Supermicro, VAST Data, NVIDIA Base Command Manager, or comparable platforms. * CI/CD and automated quality gates for infrastructure or orchestration software. ## Description Cisco Validated Infrastructure Services (CVIS) establishes and proves the performance of large AI clusters before they are handed over to customers. We deploy each cluster, validate it against performance baselines, and deliver the evidence that it is ready for service. The platform behind that work is full-stack. It combines workflow and API services, network and infrastructure automation, and the UI used to run validation across labs and customer environments. Your Impact As a Principal Software Engineer, you will be the hands-on architect and lead developer of this platform. You will take it from design into daily use: services, workflow automation, environment orchestration, diagnostics, and the operational experience. You will turn expert deployment practice into software that many teams can run, and you will stay accountable for the behavior of the whole system. * Define the architecture and build the core platform: APIs, workflow orchestration, backend services, and the engineer-facing experience. * Orchestrate lab and customer environments at scale, including concurrent runs, scheduling, throttling, cleanup, high availability, and failure handling. * Automate discovery, provisioning, configuration, and validation across compute, network, and storage, including firmware, operating systems, and cluster software when the deployment requires them. * Provide versioned workflows and multi-tenant access control so deployment and product teams can share one platform. * Automate performance validation during bring-up: run functional, resilience, and performance tests across compute, network, and storage, compare results with baselines, and produce the evidence used for handover. * Build the surfaces engineers use daily: dashboards, health and performance views, diagnostics, reporting, and triage. * Add AI-assisted diagnostics and triage, with results validated in the real environment. * Detect configuration drift and support policy-controlled remediation. * Lead design reviews, mentor senior engineers, and turn customer and field experience into reusable platform capabilities. ## Related Videos - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [10M Data Records Lost, Underwater Computing, and Psychedelic Fish - Matthias Geniar](https://www.wearedevelopers.com/videos/1908-10m-data-records-lost-underwater-computing-and-psychedelic-fish-matthias-geniar) - [Reducing Cognitive Overload Through Platform Engineering](https://www.wearedevelopers.com/videos/679-reducing-cognitive-overload-through-platform-engineering) - [Computer Vision from the Edge to the Cloud done easy](https://www.wearedevelopers.com/videos/263-computer-vision-from-the-edge-to-the-cloud-done-easy) - [How to Submit CFPs and Get into Public Speaking - Moran Weber](https://www.wearedevelopers.com/videos/2112-how-to-submit-cfps-and-get-into-public-speaking-moran-weber) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering) - [Is Software Engineering Hard?](https://www.wearedevelopers.com/magazine/448-is-software-engineering-hard) - [Top Characteristics of a Software Engineer](https://www.wearedevelopers.com/magazine/166-top-characteristics-of-a-software-engineer) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)