Sr. SRE Engineer (W2) || 6+ months || Irving, TX (5-day onsite)
Amazon.com, Inc.
Irving, TX, United States
about 1 month ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.careerjet.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$90,000.0 - $110,000.0
Working hours
Regular working hours
Job source
Tech stack
Query Performance
Amazon Web Services
Amazon Elastic Compute Cloud
Amazon S3
Application Services
Backup Devices
Bash Shell
Cloud Engineering
Continuous Integration
Data Integrity
Linux
DevOps
+27 more
Distributed Systems
Failover
Github
Identity and Access Management
Python (Programming Language)
MongoDB
Performance Tuning
Queueing Systems
RabbitMQ
Reliability Engineering
Scripting
Data Storage Management
Postman
Firebase
Amazon Virtual Private Cloud (VPC)
Microsoft InTune
Amazon Relational Database Service
Gitlab-ci
Kubernetes
Apache Kafka
Functional Programming
Cloudwatch
Amazon Simple Queue Service (SQS)
New Relic (SaaS)
Jenkins
Servicenow
Microservices
Job description
Client is seeking a highly skilled Site Reliability Engineer (SRE) to join our Production Support team. This role is responsible for ensuring the reliability, performance, and stability of our production systems across AWS, MongoDB, and related application services. The ideal candidate has strong operational instincts, deep troubleshooting skills, and a passion for building resilient systems., Production Support & Incident Management
- Serve as a primary responder for production incidents, ensuring rapid triage, mitigation, and resolution.
- Lead root cause analysis (RCA) and drive long term corrective actions.
- Maintain and improve incident response processes, runbooks, and escalation paths.
- Collaborate with engineering, QA, and product teams to prevent recurrence of issues.
AWS Infrastructure Operations
- Support and optimize AWS services such as EC2, ECS/EKS, Lambda, S3, CloudWatch, IAM, RDS, and VPC networking.
- Monitor system health, performance, and capacity across cloud environments.
- Implement infrastructure best practices around reliability, scalability, and cost efficiency.
- Assist with deployments, environment configuration, and CI/CD pipelines.
Database & Storage Support
- Manage and troubleshoot MongoDB clusters, including performance tuning, replication, backups, and failover.
- Diagnose query performance issues and collaborate with developers on schema optimization.
- Ensure data integrity, availability, and recovery readiness.
Monitoring, Observability & Alerting
- Use New Relic, CloudWatch, and other observability tools to monitor application and infrastructure performance.
- Build dashboards, alerts, and telemetry that provide actionable insights.
- Continuously refine monitoring thresholds to reduce noise and improve signal quality
Requirements
- Experience with on-call rotations and 24/7 production environments.
- Work cross-functionally with the various teams in the organization and help establish SLOs and achieve those SLOs.
- 5+ years of experience in SRE, DevOps, Production Support, or similar operational roles.
- Strong hands-on experience with AWS services and cloud-native architectures.
- Proficiency with MongoDB administration and troubleshooting.
- Experience with New Relic or similar APM/observability platforms.
- Experience using additional tools like Postman, Intune, and Firebase, Service Now, Cloudwatch.
- Strong understanding of Linux systems, networking, and distributed systems.
- Solid scripting skills (Python, Bash, or similar).
- 5+ years Monitoring and Alarming in all environments and familiar with tools like Mongo Charts, New Relic, Cloudwatch, Service Now.
- Proven experience managing high-severity incidents and driving RCA processes.
- Familiarity with CI/CD tools (Jenkins, GitHub Actions, GitLab CI, etc.).
Additional:
- Experience with container orchestration (ECS, EKS, Kubernetes).
- Knowledge of message queues (Kafka, SQS, RabbitMQ).
- Exposure to microservices architectures.
- Certifications such as AWS Solutions Architect, AWS SysOps, or MongoDB DBA.
- Working experience with IoT devices, and Microsoft Intune., FIRE PROTECTION ENGINEER JOB DESCRIPTION Position Summary: Allied Fire Protection is seeking a Fire Protection Engineer with a minimum of 5 years of experience in fire protec…
- 7 days ago
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.careerjet.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
over 2 years ago
EM
Eli McGarvie
Find a Developer Job: 12 Best Job Sites For Developers
over 3 years ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
EM
Eli McGarvie
Software Engineer Salary London
about 3 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
24 days ago