> Markdown version of [/jobs/ext/2029152-linux-sre-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2029152-linux-sre-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Linux / SRE Infrastructure Engineer - **Company:** SHREE NARAYANI INC - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Bash Shell, Computer Programming, DevOps, Distributed File Systems, Domain Name System (DNS), Monitoring of Systems, Python (Programming Language), Kerberos (Protocol), Lightweight Directory Access Protocols (LDAP), Linux System Administration, Linux Servers, Logical Volume Manager, Simple Mail Transfer Protocols, Network Service, Network Time Protocols, Performance Tuning, Reliability Engineering, Site Reliability Engineering Practices, Ansible, TCP/IP, Oracle Linux, Software Vulnerability Management, Curam Configuration Tools, Data Storage Management, Load Balancing, System Availability, Delivery Pipeline, Software Troubleshooting, Reliability of Systems, OpenSSH - **Published:** August 11, 2026 - **Apply:** https://www.dice.com/job-detail/c488a8f7-2f1a-4bc5-ae5a-7502866001cc ## About the Role 5+ years' Linux system administration in enterprise settings. - Strong expertise in Oracle Enterprise Linux (OEL) and FPP. - Experience in high availability systems, virtualization, and storage management. - Hands-on with automation/configuration tools (Ansible preferred). - Proficient in scripting/programming (Bash, Python preferred). - Strong troubleshooting, performance tuning, and incident management skills. - Solid understanding of enterprise compute, storage, and networking. - Excellent analytical, problem-solving, communication, and collaboration skills. ## Description Administer, monitor, and tune Oracle Enterprise Linux (OEL) environments in large-scale enterprises. - Oversee design, build, and lifecycle management of Linux servers, storage, and virtualization infrastructure. - Manage high availability (HA), clustering, and load-balancing for minimal downtime. - Lead capacity planning and performance optimization initiatives. Reliability & Automation (SRE Practices) - Define and implement Site Reliability Engineering (SRE) principles (SLIs, SLOs, error budgets). - Lead infrastructure automation using tools like Ansible for provisioning, configuration, and patching. - Build self-healing systems to reduce manual intervention and improve resilience. - Automate system installation, configuration, and deployment pipelines. System Administration & Infrastructure Management - Install, configure, and maintain OEL operating systems and related software. - Manage Logical Volume Manager (LVM) configurations and distributed file systems. - Administer network services (DNS, NTP, LDAP/Kerberos, SMTP, OpenSSH) and troubleshoot protocols (TCP/IP, HTTP/S, RPC). Monitoring, Incident Management & Support - Implement and enhance system monitoring, alerting, and observability. - Lead incident response, root cause analysis, and postmortem reviews. - Drive continuous improvement and oversee break/fix operations. Security & Compliance - Ensure systems are secure, hardened, and compliant with security standards. - Manage patching, vulnerability remediation, and OS upgrades. - Collaborate with security teams on access control, auditing, and encryption best practices. Leadership & Collaboration - Provide technical leadership and mentorship to SRE and infrastructure teams. - Collaborate with application, DevOps, and platform teams for system reliability. - Define and enforce operational standards, runbooks, and best practices. Documentation & Governance - Maintain comprehensive documentation for architecture and operational procedures. - Ensure compliance with change management and incident governance frameworks. - Standardize operational workflows across environments. ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)