AWS L3 SME / Cloud Specialist (AWS Primary, Google Cloud Platform Secondary

RIVAGO INFOTECH INC.
Warren, NJ, United States
4 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Kubernetes Security Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Cloud Computing Cloud Computing Security Cloud Engineering Cloud Storage Configuration Management Software Code Optimization DevOps Domain Name System (DNS)
+36 more
Identity and Access Management IP Routing Subnetting Virtual Private Networks (VPN) Python (Programming Language) Knowledge Management Network Security Routing Ansible Prometheus Shell Script Amazon Simple Notification Service (SNS) Software Vulnerability Management Google Cloud Load Balancing Cloud Platform System Cloud Monitoring Autoscaling System Availability Grafana Multi-Cloud Amazon Virtual Private Cloud (VPC) Cloudformation Containerization Kubernetes Infrastructure Automation Frameworks Information Technology Low Latency Deployment Automation Performance Monitor Route53 Cloud Migration Cloudwatch Amazon Simple Queue Service (SQS) Terraform Docker

Job description

We are seeking an experienced AWS L3 SME / Cloud Specialist to provide advanced operational support, engineering expertise, and platform optimization for enterprise cloud environments. The role will be primarily responsible for managing, troubleshooting, and enhancing AWS infrastructure while supporting workloads deployed on Google Cloud Platform.

The ideal candidate should possess strong hands-on expertise in AWS services, cloud operations, incident management, automation, security, networking, and cloud governance. The individual will act as the highest level technical escalation point for complex cloud incidents and contribute to continuous improvement initiatives., This role is critical to ensuring:

High availability and reliability of cloud platforms

Fast resolution of business-critical incidents

· Cloud security and compliance adherence

· Infrastructure automation and operational efficiency

· Optimized cloud performance and cost management

Key Responsibilities

  1. AWS Platform Operations & Engineering

· Serve as the L3 technical SME for AWS cloud services.

· Manage and support AWS environments including:

o EC2, EBS, S3, RDS, EFS

o VPC, Load Balancers, Route53

o Lambda, CloudWatch, SNS, SQS

· Troubleshoot critical production incidents and perform root cause analysis.

· Drive platform stability, reliability, and performance improvements.

· Support cloud migrations and modernization activities.

  1. Google Cloud Platform Knowledge & Support (Basic knowledge is enough)

Provide operational support for Google Cloud Platform services.

· Assist in administration of:

o Compute Engine

o Cloud Storage

o Cloud SQL

o VPC Networking

o IAM

· Collaborate with Google Cloud Platform SMEs on issue resolution and optimization initiatives.

· Support multi-cloud connectivity between AWS and Google Cloud Platform.

  1. Cloud Networking

· Manage and troubleshoot:

o VPCs

o Subnets

o Route Tables

o Security Groups

o Network ACLs

o VPN Connectivity

· Resolve routing, DNS, latency, and connectivity issues.

· Support hybrid connectivity with on-premises environments.

· Implement secure network segmentation and access controls.

  1. Security & Compliance

· Manage IAM users, roles, policies, and federation.

· Implement and maintain security best practices.

· Support vulnerability remediation and compliance initiatives.

· Manage encryption services including KMS and Secrets Manager.

· Assist with compliance requirements such as SOC2, ISO 27001, HIPAA, and NIST.

  1. Automation & Infrastructure as Code

· Develop and maintain infrastructure automation solutions.

· Work with:

o Terraform

o CloudFormation

o Ansible

o Python

o Shell Scripting

· Automate provisioning, configuration management, patching, and operational tasks.

· Contribute to self-healing and operational efficiency initiatives.

  1. Cloud Monitoring

· Implement and support monitoring platforms.

· Configure and maintain:

o CloudWatch

o Grafana

o Prometheus

o Logging Solutions

· Create operational dashboards and alerts.

· Drive proactive monitoring and incident prevention.

  1. Kubernetes & Container Platforms

· Support containerized workloads on:

o EKS

o GKE (basic operational support)

· Troubleshoot cluster, networking, ingress, and workload issues.

· Support container security and platform upgrades.

  1. Incident, Problem & Change Management

· Act as escalation point for Priority 1 and Priority 2 incidents.

· Perform detailed root cause analysis and corrective action planning.

· Participate in change reviews and implementation planning.

· Drive service improvement initiatives.

· Prepare post-incident and problem management reports.

  1. Cost Optimization & Governance

· Identify opportunities for cloud cost reduction.

· Support:

o Rightsizing

o Reserved Instances

o Savings Plans

o Storage Optimization

· Monitor cloud consumption and usage trends.

· Ensure adherence to tagging and governance standards.

  1. Collaboration & Knowledge Management

· Work closely with application, infrastructure, security, and DevOps teams.

· Create and maintain operational runbooks and technical documentation.

· Mentor L1 and L2 support engineers.

· Participate in on-call and major incident support rotations

Requirements

Experience

6-10 years of overall IT infrastructure and cloud experience.

4+ years hands-on AWS experience.

Experience supporting production cloud environments.

Strong background in incident management and operational support.

Experience in enterprise or regulated environments preferred.

Technical Skills:

· AWS (Mandatory)

· EC2

· VPC

· S3

· IAM

· RDS

· Route53

· CloudWatch

· Lambda

· Auto Scaling

· Load Balancers

· Systems Manager (SSM)

Google Cloud Platform (Good to Have)

· Compute Engine

· Cloud Storage

· Cloud SQL

· IAM

· VPC

Automation

Terraform

CloudFormation

Ansible

Python

Shell Scripting

Containers:

· Kubernetes

· Docker

· EKS

· GKE

Monitoring

Grafana, Leadership & Soft Skills

· Strong troubleshooting and analytical capabilities

· Excellent communication and stakeholder management skills

· Ability to lead technical incident bridges

· Documentation and knowledge-sharing mindset

· Mentoring and coaching capability __________________

Preferred Certifications

· AWS Certified Solutions Architect Associate / Professional

· AWS Certified SysOps Administrator

· AWS Advanced Networking Specialty (preferred)

· AWS Security Specialty (preferred)

· Google Associate Cloud Engineer (good to have)

· Terraform Associate Certification

· Kubernetes (CKA/CKAD) certification

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:13 min

Defining cloud proficiency by technical role

Piet Van Dongen · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:22 min

Analyzing differences between mobile and traditional backend DevOps

Mete Baydar Mete Baydar · World Congress 2025

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

1:24 min

Evaluating formal AWS certifications versus raw practical engineering experience

Jan Giacomelli · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all