Site Reliability Engineer (SRE)

I8IS INC.
Jersey City, NJ, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$114,400.0 - $124,800.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Batch Processing Database Queries Identity and Access Management Python (Programming Language) Reliability Engineering System Availability Amazon Virtual Private Cloud (VPC) Functional Programming Cloudwatch

Job description

Job description: The production support consultant role delivers BAU and production support for AWS infrastructure and batch processing workloads, ensuring high availability, operational stability, and on-time completion of scheduled runs.

Support Requirements

  • Monitoring & Incident Management Continuous monitoring of AWS environments and batch schedules. Rapid triage and resolution of incidents and job failures, including execution of reruns/recoveries within approved procedures.
  • Coordination & Service Restoration Engage and coordinate with engineering and dependent teams to restore service during complex incidents. Maintain clear, structured communication throughout.
  • Change & Release Support Support deployment activities including pre/post-implementation checks and rollback execution as needed.
  • Documentation Create and maintain runbooks, SOPs, troubleshooting guides, and escalation paths to ensure operational continuity across shifts.
  • SLA Accountability: Ensure critical batch jobs complete within agreed cutoffs and incident response/restoration targets are met.
  • Actively track and improve through rapid diagnosis.
  • Reduce repeat incidents through rigorous RCA and implementation of preventive actions.

Requirements

  • AWS Platform Support Hands-on experience supporting production workloads on AWS. This includes working knowledge of core services such as EC2, S3, Lambda, CloudWatch, IAM, VPC networking, and the ability to interpret AWS console/CLI outputs for troubleshooting infrastructure issues in real time.
  • Batch Scheduling & Job Management Solid understanding of enterprise batch scheduling concepts including job dependencies, sequencing, retry logic, processing windows, cutoff management, and failure handling. Ability to interpret job logs, identify upstream/downstream impacts, and execute controlled reruns.
  • SQL Proficiency Ability to write and execute basic SQL queries for operational validation and incident troubleshooting.
  • Python & automation experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

1:32 min

Comparing stream processing architecture to traditional batch processing

Bobur Umurzokov · LIVE

3:13 min

Navigating the GenAI observability dashboard in Amazon CloudWatch

Yasemin Aktürk Yasemin Aktürk · Europe 2026 Virtual

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

1:26 min

Meeting service level agreements for batch processing

Tomasz Woznica Tomasz Woznica · WWC Europe 2026

55 sec

Resolving DynamoDB hot keys using in-memory caching

Irina Branovic Irina Branovic · WWC Europe 2026

Videos

See all

Related articles

See all