Linux / SRE Infrastructure Engineer

SHREE NARAYANI INC
United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Bash Shell Computer Programming DevOps Distributed File Systems Domain Name System (DNS) Monitoring of Systems Python (Programming Language) Kerberos (Protocol) Lightweight Directory Access Protocols (LDAP) Linux System Administration Linux Servers Logical Volume Manager
+18 more
Simple Mail Transfer Protocols Network Service Network Time Protocols Performance Tuning Reliability Engineering Site Reliability Engineering Practices Ansible TCP/IP Oracle Linux Software Vulnerability Management Curam Configuration Tools Data Storage Management Load Balancing System Availability Delivery Pipeline Software Troubleshooting Reliability of Systems OpenSSH

Job description

Administer, monitor, and tune Oracle Enterprise Linux (OEL) environments in large-scale enterprises.

  • Oversee design, build, and lifecycle management of Linux servers, storage, and virtualization infrastructure.
  • Manage high availability (HA), clustering, and load-balancing for minimal downtime.
  • Lead capacity planning and performance optimization initiatives.

Reliability & Automation (SRE Practices)

  • Define and implement Site Reliability Engineering (SRE) principles (SLIs, SLOs, error budgets).
  • Lead infrastructure automation using tools like Ansible for provisioning, configuration, and patching.
  • Build self-healing systems to reduce manual intervention and improve resilience.
  • Automate system installation, configuration, and deployment pipelines.

System Administration & Infrastructure Management

  • Install, configure, and maintain OEL operating systems and related software.
  • Manage Logical Volume Manager (LVM) configurations and distributed file systems.
  • Administer network services (DNS, NTP, LDAP/Kerberos, SMTP, OpenSSH) and troubleshoot protocols (TCP/IP, HTTP/S, RPC).

Monitoring, Incident Management & Support

  • Implement and enhance system monitoring, alerting, and observability.
  • Lead incident response, root cause analysis, and postmortem reviews.
  • Drive continuous improvement and oversee break/fix operations.

Security & Compliance

  • Ensure systems are secure, hardened, and compliant with security standards.
  • Manage patching, vulnerability remediation, and OS upgrades.
  • Collaborate with security teams on access control, auditing, and encryption best practices.

Leadership & Collaboration

  • Provide technical leadership and mentorship to SRE and infrastructure teams.
  • Collaborate with application, DevOps, and platform teams for system reliability.
  • Define and enforce operational standards, runbooks, and best practices.

Documentation & Governance

  • Maintain comprehensive documentation for architecture and operational procedures.
  • Ensure compliance with change management and incident governance frameworks.
  • Standardize operational workflows across environments.

Requirements

5+ years’ Linux system administration in enterprise settings.

  • Strong expertise in Oracle Enterprise Linux (OEL) and FPP.
  • Experience in high availability systems, virtualization, and storage management.
  • Hands-on with automation/configuration tools (Ansible preferred).
  • Proficient in scripting/programming (Bash, Python preferred).
  • Strong troubleshooting, performance tuning, and incident management skills.
  • Solid understanding of enterprise compute, storage, and networking.
  • Excellent analytical, problem-solving, communication, and collaboration skills.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · World Congress 2022

7:31 min

Essential foundational skills and concepts for infrastructure roles

Megha Kadur · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all