Site Reliability Engineer - Container Platform

Stellent IT LLC
Jersey City, NJ, United States
10 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$96,800.0 - $145,200.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Application Lifecycle Management Cloud Computing Linux System Administration OpenShift Reliability Engineering Cloud Services Software Deployment Google Cloud Containerization Kubernetes Information Technology
+3 more
Rancher Drilldown Azure AKS

Job description

We are looking for an experienced Site Reliability Engineer (SRE) to support and maintain enterprise container platforms across on-premises and public cloud environments, including OpenShift, Rancher/RKE, Azure AKS, AWS, and Google Cloud. The ideal candidate will have strong hands-on experience with Kubernetes/container platforms, Linux administration, cloud infrastructure, automation, monitoring, security, and incident management. This role will focus on improving platform reliability, troubleshooting complex infrastructure issues, reducing operational toil, and driving automation and operational excellence., * Monitor, troubleshoot, and maintain container platforms including OpenShift, Rancher/RKE, and Azure AKS.

  • Troubleshoot platform performance, connectivity, availability, and security issues.
  • Perform deep-dive analysis of systemic and latent reliability issues.
  • Participate in incident and problem management activities.
  • Identify, analyze, and remediate infrastructure vulnerabilities and application deployment issues.
  • Conduct blameless Root Cause Analysis (RCA) and work with engineering and operations teams to implement permanent fixes.
  • Support application onboarding and provide troubleshooting throughout the application lifecycle.
  • Identify opportunities to automate repetitive operational tasks and reduce TOIL.
  • Partner with Risk and Compliance teams to implement controls and remediate vulnerabilities.
  • Ensure platform resiliency during implementations and proactively identify and resolve resiliency gaps.
  • Work with Architecture, Engineering, and Product teams as a key stakeholder in cloud service design.
  • Support highly available, multi-datacenter environments.
  • Participate in 24x7 on-call support using a follow-the-sun model.

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related technical field.

About the company

JPMorgan Chase

  • Plano, TX You belong to the top echelon of talent in your field. At one of the world’s most iconic financial institutions, where infrastructure is of paramount importance, you can play a piv…

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:47 min

Comparing declarative GitOps tooling alternatives and interface priorities

Davide Imola Davide Imola · LIVE

4:47 min

Exploring OpenShift and Red Hat Developer Sandbox resources

Markus Eisele Markus Eisele · World Congress 2024

3:49 min

Container hosting options available on Amazon Web Services

Federico Fregosi · World Congress 2022

1:35 min

Accessing software containers and developer training platforms

Paul Graham Paul Graham · World Congress 2024

3:17 min

Deploying applications directly to OpenShift using the shift tool

Don Schenck Don Schenck · World Congress 2024

2:37 min

Deploying the augmented chatbot on OpenShift AI platforms

Cedric Clyburn Cedric Clyburn · World Congress 2024

Videos

See all

Related articles

See all