Linux / SRE Infrastructure Engineer
Drunix Solution Inc
Chandler, AZ, United States
2 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source
Tech stack
Bash Shell
DevOps
Distributed File Systems
Domain Name System (DNS)
Python (Programming Language)
Kerberos (Protocol)
Lightweight Directory Access Protocols (LDAP)
Linux System Administration
Linux Servers
Logical Volume Manager
Simple Mail Transfer Protocols
Network Service
+17 more
Network Time Protocols
Network Protocols
Performance Tuning
Site Reliability Engineering Practices
Ansible
Sendmail
TCP/IP
Oracle Linux
Software Vulnerability Management
Scripting
Data Storage Management
System Availability
Delivery Pipeline
Reliability of Systems
Postfix
OpenSSH
Infrastructure Automation Frameworks
Job description
- Lead the administration, monitoring, and performance tuning of Oracle Enterprise Linux (OEL) environments in a large-scale enterprise ecosystem.
- Oversee the design, build, and lifecycle management of Linux servers, including storage, virtualization, and associated infrastructure.
- Manage high availability (HA) configurations, clustering, and load-balanced environments to ensure minimal downtime.
- Drive capacity planning, performance optimization, and system scalability initiatives.
Reliability & Automation (SRE Practices)
- Define and implement SRE principles, including SLIs, SLOs, and error budgets.
- Lead initiatives for infrastructure automation (provisioning, configuration, patching) using tools such as Ansible.
- Build and maintain self-healing systems, reducing manual intervention and improving system resilience.
- Develop automation for system installation, configuration, and deployment pipelines.
System Administration & Infrastructure Management
- Install, configure, and maintain Oracle Enterprise Linux (OEL) operating systems and related software stacks.
- Manage Logical Volume Manager (LVM) configurations, including volume groups and filesystem expansion.
- Administer distributed file systems, NFS servers/clients, and automount configurations.
- Maintain network services such as DNS, NTP, LDAP/Kerberos, SMTP (sendmail/postfix), and OpenSSH.
- Troubleshoot and support network protocols (TCP/IP, HTTP, HTTPS, RPC).
Monitoring, Incident Management & Support
- Implement and enhance monitoring, alerting, and observability frameworks for proactive issue detection.
- Lead incident response, root cause analysis (RCA), and postmortem reviews.
- Drive continuous improvement by identifying systemic issues and implementing preventive solutions.
- Oversee break/fix operations, ensuring timely resolution and minimal business impact.
Security & Compliance
- Ensure systems are secure, hardened, and compliant with enterprise security standards.
- Manage patching, vulnerability remediation, and OS upgrades.
- Partner with security teams to implement best practices for access control, auditing, and encryption.
- Leadership & Collaboration
- Provide technical leadership and mentorship to SRE and infrastructure teams.
- Collaborate with application, DevOps, and platform teams to improve system reliability and deployment processes.
- Define and enforce operational standards, runbooks, and best practices.
- Drive cross-functional initiatives to enhance platform stability and efficiency.
Documentation & Governance
- Maintain comprehensive documentation for architecture, processes, and operational procedures.
- Ensure adherence to change management and incident governance frameworks.
- Standardize operational workflows across environments.
Requirements
- 5+ years of experience in Linux system administration in enterprise environments.
- Strong expertise in Oracle Enterprise Linux (OEL) systems and FPP.
- Proven experience in high availability systems, virtualization, and storage management.
- Hands-on experience with automation and configuration management tools (Ansible preferred).
- Proficiency in at least one scripting/programming language ( Bash, Python preferred).
- Strong experience with infrastructure troubleshooting, performance tuning, and incident management.
- Solid understanding of enterprise infrastructure (compute, storage, network).
- Excellent analytical, problem-solving, and organizational skills.
- Strong communication and collaboration skills in a global team environment.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
over 2 years ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
about 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
LM
Luis Minvielle
Why Upskilling And Reskilling is Important For Developers
over 2 years ago
EM
Eli McGarvie
DevOps Engineer Salary [2023]
over 3 years ago
JF
Jonas Fritzsch
Résumé-Driven Development: How IT trends affect the job market for software developers
almost 5 years ago