Senior Cloud & Linux Infrastructure Engineer
Role details
Job location
Tech stack
Job description
We're looking for a Senior Cloud & Linux Infrastructure Engineer to design, build, and operate scalable, secure, high-performance infrastructure for compute-intensive semiconductor design workloads (EDA, simulation, verification, and engineering services). You'll own key parts of our AWS cloud platform and Linux estate, with a strong focus on Linux-based user environments, High Performance Computing (HPC) engineering, shared services, and automation-first operations.
You'll work closely with CAD, design, and verification teams to improve throughput, reliability, and cost efficiency while maintaining strong security and operational standards. This role requires deep hands-on Linux knowledge across servers, tooling, identity, storage, scheduling, and troubleshooting in a mixed cloud and on-prem environment., Cloud Architecture & Linux Platform Engineering
- Design and operate AWS-based environments optimized for semiconductor and EDA workloads, covering both interactive user access and batch compute execution.
- Build, maintain, and improve Linux-based engineering environments for end users, including performance, scalability, lifecycle management, and operational support.
- Engineer and support Linux platform services across compute, storage, identity, access, and shared tooling layers.
- Architect HPC solutions across compute, storage, and data movement, with a focus on high-throughput and low-latency workload patterns used in simulation and verification environments.
- Deploy, tune, and support compute clusters using SLURM, including capacity planning, queue policies, scheduling behavior, node lifecycle management, and Linux image consistency.
Linux Systems Administration & Shared Services
- Administer and troubleshoot Linux systems across RHEL, Amazon Linux, Ubuntu, or similar enterprise distributions, including patching, hardening, performance analysis, and service reliability.
Manage core Linux services such as centralized authentication, access control, DNS, NFS, package repositories, shared software locations, and configuration standards.
- Support shared engineering tool deployments on Linux, including version control of installations, permissions, module-based access patterns, and multi-user usability.
- Create and maintain reusable Linux base images, templates, and standards for engineering workstations, utility hosts, and cluster nodes.
Infrastructure as Code & Automation
- Implement and version infrastructure using Terragrunt and related automation tooling, with Ansible or equivalent configuration management where appropriate.
- Build and maintain automated CI/CD workflows for infrastructure provisioning, platform changes, and Linux configuration consistency.
- Develop reusable modules, templates, and standards to improve consistency, repeatability, and speed of delivery across cloud and Linux environments.
- Automate operational tasks using scripting languages such as Bash and Python, especially for Linux platform administration and support workflows.
Operations & Maintenance
- Monitor, troubleshoot, and continuously improve availability and performance across cloud infrastructure, Linux systems, and cluster services.
- Establish operational best practices including alerting, logging, runbooks, incident response, patching strategy, change control, and service recovery procedures.
- Act as an escalation point for complex Linux, cloud, and HPC incidents, driving root cause analysis and permanent corrective actions.
- Partner with CAD and engineering teams to optimize workflows, usage patterns, software delivery, and job throughput across engineering environments.
Security & Compliance
- Apply secure-by-design principles across Linux and AWS environments, including hardening, least privilege, segmentation, secure administration, and service isolation.
- Implement and maintain IAM strategy, Linux access controls, secrets handling, encryption, and secure network design.
- Participate in security reviews, vulnerability remediation, audit preparation, and operational compliance activities.
Requirements
- 10+ years in Cloud Engineering, Linux Infrastructure, DevOps, SRE, or similar roles.
- Strong hands-on AWS experience, including designing for performance, resilience, and operational supportability.
- Deep Linux administration experience in production environments, including troubleshooting, system internals, networking, storage, identity integration, and shell-based operations.
- Strong infrastructure-as-code skills with Terraform; Ansible or equivalent configuration automation is a plus.
- Experience with HPC schedulers and batch systems, with SLURM strongly preferred.
- Solid understanding of networking, distributed systems, compute provisioning, Linux performance tuning, and storage architectures.
- Strong knowledge of security across cloud and Linux environments, including IAM, SSH access models, encryption, least-privilege design, and secure networking.
- Strong scripting skills in Bash, Python, or similar languages for automation and systems operations.
Desirable Experience
- Exposure to semiconductor design workflows, EDA toolchains, or other HPC-heavy engineering industries.
- Experience supporting Linux-based engineering desktops, virtual workstations, or remote design environments.
- Familiarity with shared Linux identity and authentication platforms such as FreeIPA, LDAP, Active Directory integration, or SSSD.
- Experience with centralized software deployment patterns, shared filesystems, and environment module frameworks for Linux.
- Observability tooling experience across infrastructure and Linux services, such as CloudWatch, Prometheus, or Grafana.
- Proficiency with Git-based workflows and structured infrastructure documentation practices.