IT Manager
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+26 more
Job description
Among the key duties of this position are the following:
- Leads the team responsible for Rutgers High-Performance Computing (HPC) platforms and closely related infrastructure.
- Oversees day-to-day reliability and long-term evolution of clustered compute, GPU resources, high-performance storage, and the scheduler/tooling that support research at scale.
- Drives a culture of operational excellence and collaboration with researchers and campus IT partners.
- Ensures secure, performant, and cost-effective HPC services aligned to university priorities and compliance standards, while planning capacity and guiding technology roadmaps.
- Coordinates vendor engagements to build a future-ready research computing ecosystem., Full Time Working Hours Standard Hours 37.50 Daily Work Shift Work Arrangement Consistent with the current application of Rutgers Policy 60.3.22 or the applicable provisions of relevant collective negotiations agreements, this position may be eligible for a hybrid work arrangement. Flexible work arrangements are not permanent, subject to change or discontinuation, and contingent on the employee receiving approval in the FlexWork@RU Application System.
Requirements
Do you have experience in Vulnerability management?, Do you have a Bachelor’s degree?, * Bachelor’s degree in Computer Science, Information Technology, or a related field.
- Minimum of 7 years of experience in IT systems administration/engineering, including direct experience with hardware, virtualization technologies, various operating systems, and production support of clustered or large-scale environments. This experience must include at least 3 years of team lead experience., * Demonstrated expertise with HPC technologies: Linux at scale, job schedulers, workload accounting, and performance tuning.
- Experience with high-performance research storage for HPC, including parallel filesystems, with knowledge of quotas, snapshots, and DR concepts.
- Familiarity with high-speed interconnects, GPUs/accelerators, and node provisioning in clustered environments.
- Proficiency with automation and configuration management, and with common scripting languages for operations.
- Hands-on experience with IBM Spectrum Scale (GPFS), Slurm workload manager, InfiniBand high-speed interconnects (including switches, fabric configuration, and performance optimization), HPC software environment management using module systems (e.g., Lmod), and container technologies (e.g., Apptainer).
- Strong understanding of security best practices for shared compute/storage services, including identity/integrations (LDAP/AD), access controls, and patch/vulnerability management.
- Excellent communication and collaboration skills with the ability to work effectively across central IT, researchers, and vendors.
- Proven project management skills in planning, execution, risk management, and delivery., * Advanced degree in Computer Science, Information Systems, or a related field.
- Prior experience managing a multi-tenant HPC environment in higher education or a similar complex research setting.
- Experience designing/operating GPU-capable and heterogeneous clusters; container workflows and module systems.
- Experience with cost modeling and utilization reporting.
- Familiarity with compliance sensitive research workflows and data governance.
- Budget planning, hardware lifecycle, and vendor/SOW management.
Equipment Utilized
- HPC Cluster Hardware: Experience with HPC server platforms including dense compute, GPU/accelerator, and large-memory nodes from vendors such as Lenovo, Dell, and HPE, along with rack-scale power, cooling, and cabling considerations common in research data centers.
- High-Speed Interconnects and GPUs: Hands-on experience with InfiniBand fabrics (subnet management, fabric configuration, and performance tuning) and NVIDIA GPUs/accelerators with associated drivers and libraries (CUDA, NCCL) in clustered environments.
- Linux at Scale and Cluster Provisioning: Proficiency managing enterprise Linux (Red Hat Enterprise Linux, etc.) and cluster provisioning and imaging tools such as Warewulf, combined with Ansible-based configuration management and automation for consistent node deployment and lifecycle management.
- Parallel and Research Storage: Experience with parallel filesystems such as IBM Spectrum Scale (GPFS) and Lustre, along with NFS for home and project directories, tiered storage, snapshots, quotas, and research data backup and disaster recovery tools.
- Schedulers and Workload Management: Expertise with Slurm (including accounting, QoS, fair-share, partitions, and reservations) and familiarity with alternatives such as PBS Pro/OpenPBS and LSF used in research computing environments.
- Research Software Environments and Containers: Proficiency with scientific software environment tools such as Lmod, EasyBuild, and Spack, and container runtimes for HPC such as Apptainer/Singularity to support reproducible research workflows.
- Monitoring, Automation, and Identity Services: Proficiency with monitoring and utilization tools (e.g., Prometheus/Grafana, XDMoD, Ganglia), configuration management and automation (e.g., Ansible, Git), and shared authentication services (LDAP/Active Directory, Kerberos) used to operate multi-tenant research computing environments.
- Collaboration and Ticketing Tools: Proficiency with office productivity, ticketing, and collaboration platforms such as Microsoft Office Suite, Microsoft Teams, Slack, Jira or ServiceNow, and documentation and knowledge-base systems used to coordinate operations and engage with the research community.
Physical Demands and Work Environment
- Ability to work with personal computer for 8 hours a day.
- Participate in online and in-person meetings.
- Light local travel to visit locations housing Enterprise Infrastructure equipment., * * Do you have a bachelor’s degree in computer science, information technology, or a related field?
- Yes
-
No
-
- Do you have a minimum of 7 years of experience in IT systems administration/engineering, including direct experience with hardware, virtualization technologies, various operating systems, and production support of clustered or large-scale environments; including at least 3 years of team lead experience?
Benefits & conditions
Pulled from the full job description
- Employee discount
- Dental insurance
- Life insurance
- Paid holidays, Rutgers provides a comprehensive benefits package to eligible employees. The specific benefits vary based on the position and may include:
- Medical, prescription drug, and dental coverage
- Paid vacation, holidays, and various leave programs
- Competitive retirement benefits, including defined contribution plans and voluntary tax-deferred savings options
- Employee and dependent educational benefits (when applicable)
- Life insurance coverage
- Employee discount programs
About the company
Rutgers, The State University of New Jersey, is a leading national research university and the State of New Jersey’s preeminent, comprehensive public institution of higher education. As one of the largest employers in the State of New Jersey, Rutgers University is committed not only to the students and the State that we serve, but also to the faculty and staff who work on our campuses. For two consecutive years, Rutgers is ranked on Forbes’ list of America’s Best Large Employers. Rutgers holds #64 of 500 employers and is the #1 New Jersey employer on the publication’s 2023 list. Rutgers’ commitment to its employees includes maintaining and fostering a safe, diverse, and respectful workplace environment, creating employment opportunities for our nation’s military veterans, and ensuring accessibility and accommodation for individuals with disabilities. Posting Summary Rutgers, The State University of New Jersey is seeking an IT Manager for the OIT - Enterprise Infrastructure. This position reports to the Director, Enterprise Infrastructure (OIT).
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on indeed.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Résumé-Driven Development: How IT trends affect the job market for software developers
A Guide to Green Tech and Green IT Careers
Dev Digest 121 - AI goes offline
7 Cloud Computing Trends Coming in 2025 for Developers