SRE- DevOps Engineer

Altitude Technology Solutions Inc
Englewood Cliffs, NJ, United States
3 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Bash Shell Continuous Integration DevOps Identity and Access Management Python (Programming Language) PostgreSQL Linux System Administration MySQL Performance Tuning
+14 more
Reliability Engineering Cloud Services Ansible Prometheus Datadog Grafana Reliability of Systems Amazon Virtual Private Cloud (VPC) Git Amazon Relational Database Service Cloudwatch Terraform Splunk Jenkins

Job description

Automate operational tasks and infrastructure management using Shell, Python, Ansible, or Terraform.

Manage and support AWS services including EC2, RDS, S3, IAM, VPC, CloudWatch, and related cloud services.

Perform Linux server administration, troubleshooting, patching, and performance tuning.

Monitor application and infrastructure health using tools such as Grafana, Prometheus, CloudWatch, Datadog, Splunk.

Participate in incident management, root cause analysis (RCA), and problem management activities.

Define and maintain SLIs, SLOs, and SLAs to ensure service reliability.

Support PostgreSQL and MySQL databases for operational and basic administration tasks.

Collaborate with development, QA, cloud, and support teams to improve system reliability and deployment processes.

Drive automation, observability, capacity planning, security, and operational best practices.

Requirements

6-7 years of experience in Site Reliability Engineering, Production Support, DevOps, or Infrastructure Operations.

Strong understanding of Linux administration and troubleshooting.

Hands-on experience with AWS cloud services (EC2, RDS, IAM, VPC, CloudWatch, S3).

Experience with monitoring, alerting, and observability tools.

Knowledge of incident management, problem management, and RCA processes.

Experience with automation and scripting using Shell and/or Python.

Working knowledge of PostgreSQL and MySQL databases.

Experience with Git version control.

Understanding of CI/CD concepts and tools such as Jenkins.

Roles & Responsibilities

Design, implement, and maintain highly available and reliable production systems.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

1:48 min

Analyzing network packets with database protocol tools

Daniël van Eeden Daniël van Eeden · World Congress 2026 Europe

Videos

See all

Related articles

See all