Site Reliability Engineer (SRE)
I8IS INC.
Jersey City, NJ, United States
about 1 month ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$114,400.0 - $124,800.0
Working hours
Regular working hours
Job source
Tech stack
Amazon Web Services
Amazon Elastic Compute Cloud
Amazon S3
Batch Processing
Database Queries
Identity and Access Management
Python (Programming Language)
Reliability Engineering
System Availability
Amazon Virtual Private Cloud (VPC)
Functional Programming
Cloudwatch
Job description
Job description: The production support consultant role delivers BAU and production support for AWS infrastructure and batch processing workloads, ensuring high availability, operational stability, and on-time completion of scheduled runs.
Support Requirements
- Monitoring & Incident Management Continuous monitoring of AWS environments and batch schedules. Rapid triage and resolution of incidents and job failures, including execution of reruns/recoveries within approved procedures.
- Coordination & Service Restoration Engage and coordinate with engineering and dependent teams to restore service during complex incidents. Maintain clear, structured communication throughout.
- Change & Release Support Support deployment activities including pre/post-implementation checks and rollback execution as needed.
- Documentation Create and maintain runbooks, SOPs, troubleshooting guides, and escalation paths to ensure operational continuity across shifts.
- SLA Accountability: Ensure critical batch jobs complete within agreed cutoffs and incident response/restoration targets are met.
- Actively track and improve through rapid diagnosis.
- Reduce repeat incidents through rigorous RCA and implementation of preventive actions.
Requirements
- AWS Platform Support Hands-on experience supporting production workloads on AWS. This includes working knowledge of core services such as EC2, S3, Lambda, CloudWatch, IAM, VPC networking, and the ability to interpret AWS console/CLI outputs for troubleshooting infrastructure issues in real time.
- Batch Scheduling & Job Management Solid understanding of enterprise batch scheduling concepts including job dependencies, sequencing, retry logic, processing windows, cutoff management, and failure handling. Ability to interpret job logs, identify upstream/downstream impacts, and execute controlled reruns.
- SQL Proficiency Ability to write and execute basic SQL queries for operational validation and incident troubleshooting.
- Python & automation experience.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
over 2 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
about 2 years ago
LM
Luis Minvielle
Why Upskilling And Reskilling is Important For Developers
over 2 years ago
KM
Kaleb McKelvey
The Best Software Developer Blogs to Read
over 3 years ago