Linux / SRE Infrastructure Engineer
SHREE NARAYANI INC
United States
2 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source
Tech stack
Bash Shell
Computer Programming
DevOps
Distributed File Systems
Domain Name System (DNS)
Monitoring of Systems
Python (Programming Language)
Kerberos (Protocol)
Lightweight Directory Access Protocols (LDAP)
Linux System Administration
Linux Servers
Logical Volume Manager
+18 more
Simple Mail Transfer Protocols
Network Service
Network Time Protocols
Performance Tuning
Reliability Engineering
Site Reliability Engineering Practices
Ansible
TCP/IP
Oracle Linux
Software Vulnerability Management
Curam Configuration Tools
Data Storage Management
Load Balancing
System Availability
Delivery Pipeline
Software Troubleshooting
Reliability of Systems
OpenSSH
Job description
Administer, monitor, and tune Oracle Enterprise Linux (OEL) environments in large-scale enterprises.
- Oversee design, build, and lifecycle management of Linux servers, storage, and virtualization infrastructure.
- Manage high availability (HA), clustering, and load-balancing for minimal downtime.
- Lead capacity planning and performance optimization initiatives.
Reliability & Automation (SRE Practices)
- Define and implement Site Reliability Engineering (SRE) principles (SLIs, SLOs, error budgets).
- Lead infrastructure automation using tools like Ansible for provisioning, configuration, and patching.
- Build self-healing systems to reduce manual intervention and improve resilience.
- Automate system installation, configuration, and deployment pipelines.
System Administration & Infrastructure Management
- Install, configure, and maintain OEL operating systems and related software.
- Manage Logical Volume Manager (LVM) configurations and distributed file systems.
- Administer network services (DNS, NTP, LDAP/Kerberos, SMTP, OpenSSH) and troubleshoot protocols (TCP/IP, HTTP/S, RPC).
Monitoring, Incident Management & Support
- Implement and enhance system monitoring, alerting, and observability.
- Lead incident response, root cause analysis, and postmortem reviews.
- Drive continuous improvement and oversee break/fix operations.
Security & Compliance
- Ensure systems are secure, hardened, and compliant with security standards.
- Manage patching, vulnerability remediation, and OS upgrades.
- Collaborate with security teams on access control, auditing, and encryption best practices.
Leadership & Collaboration
- Provide technical leadership and mentorship to SRE and infrastructure teams.
- Collaborate with application, DevOps, and platform teams for system reliability.
- Define and enforce operational standards, runbooks, and best practices.
Documentation & Governance
- Maintain comprehensive documentation for architecture and operational procedures.
- Ensure compliance with change management and incident governance frameworks.
- Standardize operational workflows across environments.
Requirements
5+ years’ Linux system administration in enterprise settings.
- Strong expertise in Oracle Enterprise Linux (OEL) and FPP.
- Experience in high availability systems, virtualization, and storage management.
- Hands-on with automation/configuration tools (Ansible preferred).
- Proficient in scripting/programming (Bash, Python preferred).
- Strong troubleshooting, performance tuning, and incident management skills.
- Solid understanding of enterprise compute, storage, and networking.
- Excellent analytical, problem-solving, communication, and collaboration skills.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
EM
Eli McGarvie
over 3 years ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
about 2 years ago
Learning Kubernetes made easy with KubeCampus
over 2 years ago
EM
Eli McGarvie
DevOps Engineer Salary [2023]
over 3 years ago
LM
Luis Minvielle
Why Upskilling And Reskilling is Important For Developers
over 2 years ago