> Markdown version of [/jobs/ext/3397468-software-engineer-imc-platform-and-cluster-infrastructure](https://www.wearedevelopers.com/jobs/ext/3397468-software-engineer-imc-platform-and-cluster-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - IMC Platform and Cluster Infrastructure - **Company:** KLA-Tencor - **Location:** Milpitas, CA, United States - **Experience:** Experienced - **Salary:** $136,300.0 - $231,700.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, Intelligent Platform Management Interface, Bash Shell, BIOS, Configuration Management, Nvidia CUDA, Computer Networks, Computer Engineering, System Configuration, Linux, File Systems, Distributed Systems, Domain Name System (DNS), Ethernet, Firmware, InfiniBand, Python (Programming Language), Linux System Administration, Linux-Powered Devices, Networking Basics, Network Diagnostics, Routing, Package Management Systems, Performance Tuning, Remote Direct Memory Access, Release Management, Ansible, Prometheus, Software Engineering, Software Systems, Systems Integration, TCP/IP, Scripting, Graphics Processing Unit (GPU), Computer Networking Systems, Computer Network Technologies, Grafana, Git, Containerization, Kubernetes, Infrastructure Automation Frameworks, Storage Technologies, Information Technology, Machine Learning Operations, Hardware Infrastructure, Server Operating Systems & Platforms - **Published:** September 7, 2026 - **Apply:** https://kla.wd1.myworkdayjobs.com/Search/job/Milpitas-CA/Software-Engineer---IMC-Platform-and-Cluster-Infrastructure_2639272 ## About the Role + Experience with Linux system administration, platform integration, and infrastructure support in enterprise or production environments, including Linux networking, file systems, storage technologies, and performance tuning. + Hands-on experience with Ansible, cluster provisioning solutions, infrastructure automation, configuration management, and deployment orchestration. + Strong scripting and software development skills using Python, Bash, or similar languages to automate installation, validation, monitoring, and operational workflows. + Experience deploying, administering, and troubleshooting Kubernetes, containerized applications, and distributed computing environments. + Understanding of Linux device drivers, kernel modules, operating system internals, hardware-software dependencies, and platform integration. + Experience evaluating and qualifying server platforms, including BIOS, BMC/IPMI firmware, hardware health, and system configurations across enterprise server technologies. + Knowledge of GPU-accelerated computing environments, including NVIDIA GPU drivers, CUDA, AI/ML infrastructure, and performance optimization for compute-intensive workloads. + Experience with high-performance networking technologies, including Mellanox/NVIDIA networking solutions, RDMA, RoCE, InfiniBand, and high-speed Ethernet infrastructures. + Experience with observability, monitoring, and cluster health management frameworks such as Prometheus, Grafana, ELK, OpenTelemetry, or similar platforms. + Familiarity with modern software delivery practices including Git, CI/CD pipelines, Helm, Infrastructure-as-Code, and automated release management methodologies. + Experience supporting secure, mission-critical, or air-gapped production environments, including customer deployments, platform upgrades, and critical issue resolution. + Strong troubleshooting and root cause analysis skills across software, infrastructure, networking, operating systems, GPUs, and hardware platforms, with the ability to collaborate effectively with vendors, cross-functional engineering teams, and customers. + Willingness and ability to travel domestically and internationally to support platform qualification, deployment, installation, and customer-facing activities., + Bachelor's degree in Computer Science, Computer Engineering, Information Technology, Electrical Engineering, or a related technical field. + 2+ years of experience in Linux system administration, infrastructure engineering, platform engineering, or deployment engineering. + Strong understanding of Linux operating systems, system services, package management, software installation, and fault isolation /resolution. + Experience with scripting and automation using Python, Bash, or similar scripting languages. + Experience with Ansible, infrastructure automation, cluster provisioning, or configuration management tools. + Knowledge of Linux networking fundamentals, including TCP/IP, DNS, routing, switching, and network diagnostics. + Understanding of Linux file systems, storage technologies, and system-level analysis. + Familiarity with server hardware, operating system deployment, hardware-software dependencies, and platform integration concepts. + Experience installing, configuring, validating, and supporting software solutions in Linux-based environments. + Strong analytical, problem solving, and root-cause analysis skills. + Ability to work effectively in a multi-functional engineering environment and collaborate with software, systems, hardware, and field teams. + Willingness to travel domestically and internationally to support platform deployments, customer installations, upgrades, and critical field customer concerns.