Manager, Linux Platform
Role details
Job location
Tech stack
Job description
Our client is seeking an experienced Manager, Linux Platform to lead a team of Linux System Engineers and Administrators responsible for supporting and modernizing enterprise infrastructure across corporate and distribution center environments.
This is a hands-on leadership role for a seasoned technology manager who can balance people leadership, operational excellence, and technical strategy. The ideal candidate will bring deep Linux expertise, strong Azure cloud experience, and a proven track record of leading teams through operational and technical transformation.
You will be responsible for managing a team of 5-9 engineers and administrators, driving automation and modernization initiatives, ensuring platform reliability and security, and serving as the senior escalation point for critical production issues.
Key Responsibilities
Leadership & Team Development
- Lead, mentor, and develop a team of 5-9 System Engineers and Administrators, including both full-time employees and consultants.
- Establish clear goals, performance expectations, and accountability measures.
- Foster a culture of collaboration, technical excellence, innovation, and continuous improvement.
- Drive organizational and operational change initiatives to improve team effectiveness and platform maturity.
- Manage workload prioritization, resource planning, succession planning, and on-call rotations.
- Partner with vendors and service providers to ensure service quality and SLA compliance.
- Collaborate with Finance and Procurement to manage budgets, licensing, renewals, vendor contracts, and cloud spend.
Linux Platform Engineering & Operations
- Oversee the administration, configuration, patching, maintenance, and optimization of enterprise Linux environments, including RHEL and SUSE.
- Lead infrastructure modernization efforts across on-premises and cloud platforms.
- Drive automation using Ansible, Terraform, Bash scripting, and CI/CD tools.
- Support cloud-native technologies, containers, platform services, and infrastructure-as-code initiatives.
- Serve as the primary escalation point for complex production incidents and critical outages.
- Lead root cause analysis (RCA) efforts and implement permanent corrective actions.
- Manage platform lifecycle planning, capacity management, and technology upgrades.
- Oversee enterprise backup, recovery, storage, NAS, and SAN environments.
Security, Compliance & Governance
- Ensure compliance with SOX, PCI-DSS, and internal security standards.
- Manage access controls, auditing, encryption, vulnerability remediation, and security best practices.
- Maintain accurate technical documentation, architecture diagrams, operational procedures, and support runbooks.
Operational Excellence
- Ensure platform availability, resilience, and performance across mission-critical environments.
- Lead disaster recovery planning, testing, and business continuity initiatives.
- Drive observability, monitoring, and proactive issue prevention.
- Support 24x7 operational environments with defined service-level objectives and response processes.
Requirements
- Bachelor's degree in Computer Science, Information Technology, or a related discipline, or equivalent experience.
- 7+ years of hands-on experience supporting enterprise Linux environments, including RHEL and/or SUSE.
- 3+ years of people leadership experience managing System Engineers, Platform Engineers, Infrastructure Engineers, or DevOps teams.
- Demonstrated success leading teams through operational or technology transformation initiatives.
- Strong experience with Microsoft Azure cloud services and hybrid infrastructure environments.
- Deep expertise in Linux administration, performance tuning, troubleshooting, and system optimization.
- Hands-on experience with infrastructure automation and configuration management tools such as Ansible, Terraform, Puppet, or similar technologies.
- Experience supporting virtualization platforms such as VMware vSphere.
- Strong understanding of networking protocols and services including TCP/IP, DNS, DHCP, NTP, VPN, and load balancing technologies.
- Experience with monitoring and observability platforms such as Datadog, Dynatrace, Nagios, SolarWinds, or equivalent.
- Knowledge of enterprise backup and recovery solutions including Veeam, Rubrik, or NetBackup.
- Experience supporting highly available, 24x7 production environments.
- Excellent communication and stakeholder management skills, with the ability to engage technical and executive audiences., * Experience within retail, distribution, logistics, or other high-volume transactional environments.
- Experience with containers, Kubernetes, serverless technologies, and cloud modernization programs.
- Familiarity with databases and middleware platforms.
- ITIL certification or practical knowledge of Incident, Problem, Change, and Service Management processes.
- Experience managing compliance-driven environments subject to SOX and PCI requirements.