Site Reliability, Principal

Synopsys
Sunnyvale, CA, United States
2 months ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$185,000.0 - $278,000.0
Working hours
Regular working hours
Job source

Tech stack

Ubuntu (Operating System) CentOS Computer Clusters Software Debugging Linux Linux Distribution Red Hat Enterprise Linux Reliability Engineering Site Reliability Engineering Practices Network Storage SUSE Linux Computer Network Technologies
+2 more
Containerization Slurm

Job description

  • Applying SRE practices to identify, monitor, communicate, and resolve issues in the environment, while also collaborating with internal teams and customers on post-mortem analysis to deliver root cause insights.
  • Following up on issues reported and looking for procedures to prevent similar occurrences.
  • Reviewing current processes and transforming them into scalable solutions.
  • Debugging OS and engineering issues within our provided Linux environment.
  • Collaborating on internal projects across different time zones and teams.
  • Following up with customers and handing over tasks/issues with team members to utilize time zones efficiently.

The Impact You Will Have:

  • Enhancing the reliability and performance of our engineering environment.
  • Streamlining processes to ensure scalability and efficiency.
  • Resolving complex OS and engineering issues, contributing to smoother operations.
  • Driving successful project outcomes through effective collaboration across time zones.
  • Improving customer satisfaction by addressing and resolving issues promptly.
  • Foster SRE practices within multifunctional teams and identify gaps for resolution.

Requirements

You are a person looking to work in an intercultural and global team. You thrive on solving challenges in a large-scale HPC environment, right at the heart of technology. You are passionate about creating scalable processes and enjoy working collaboratively across time zones. Your excellent problem-solving skills and ability to work through issues and challenges make you a valuable team member. You are excited about joining an innovative team that values continuous learning, great leadership, and being part of a growing organization., * 10+ years of SRE processes and related skills required

  • Capability to understand complex engineering implementations and their inter dependencies for troubleshooting.
  • Deep Knowledge with Linux distributions (CentOS, RedHat, Ubuntu, SuSE).
  • Deep Knowledge of virtualization and containerization technologies.
  • Extensive knowledge of storage solutions, including network storage and associated protocols.
  • Good Experience in network technologies.
  • Good Experience in load sharing facilities such as LSF, Slurm and various workload scheduling technologies.
  • Good interpersonal, communication and leadership skills, * Part of a global Team supporting one of the biggest scaled environments that includes multiple HPC clusters, High performance Storage, Large scale private cloud implementation as well as GPU clusters for HPC/GenAI workloads.
  • Challenging yourself to work with the latest state of Art technologies
  • Being part of one of the biggest private clouds in the world.
  • Embrace and implement SRE best practices.
  • An individual who monitors and comprehends complex environments.
  • Able to break down complex issues into relevant areas and independently coordinate follow-ups with internal teams. A good communicator with interpersonal skills.
  • A proactive problem solver with a keen eye for detail.
  • A collaborative team player who thrives in a global, intercultural environment.
  • Adept at multitasking and managing multiple priorities effectively.
  • Self-motivated and capable of working independently.
  • Passionate about continuous learning and professional development.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

4:41 min

Scale and diversity of software development teams

Bastian Heilemann Bastian Heilemann +1 · World Congress 2025

Videos

See all

Related articles

See all