Platform Engineer / SRE (Site Reliability Engineer) Cloud & AI Automation

StarTechs Inc.
United States
11 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Amazon Web Services Amazon Cloudfront Amazon Elastic Compute Cloud Amazon S3 Microsoft Azure BigQuery Continuous Integration Extract Transform Load (ETL) DevOps Federated Identity Management
+41 more
Github Identity and Access Management Python (Programming Language) PCI Data Security Standards Performance Tuning Role-Based Access Control Reliability Engineering Prometheus Azure Machine Learning Systems Integration Trusted Systems Datadog Data Logging Scripting Google Cloud Cloud Platform System Okta Autoscaling Snowflake Grafana Multi-Cloud Reliability of Systems AWS Lambda Infrastructure as Code (IaC) Amazon Virtual Private Cloud (VPC) Cloudformation Amazon Relational Database Service Gitlab-ci Kubernetes Infrastructure Automation Frameworks Information Technology Data Management Virtual Agents Functional Programming Kibana Terraform Serverless Computing Docker Jenkins Amazon Redshift Microservices

Job description

Design, implement, and manage cloud-native platforms on AWS using Terraform, CloudFormation, and IaC best practices.

  • Lead infrastructure automation initiatives to enable zero-touch provisioning, self-service environments, and CI/CD pipeline integration.

  • Architect and operate highly available, scalable, and secure systems using Kubernetes (EKS/ECS), Docker, and microservices.

  • Implement and maintain robust observability stacks using Grafana, Datadog, Prometheus, ELK, and OpenTelemetry for real-time monitoring, alerting, and performance tuning.

  • Own incident management lifecycle: lead on-call rotations, conduct RCA (Root Cause Analysis), and implement preventive measures to improve system reliability.

  • Drive cloud cost optimization strategies across multi-cloud environments (AWS, Azure, Google Cloud Platform) through right-sizing, auto-scaling, tagging policies, and usage analytics.

  • Enforce security compliance standards including PCI, PII, HIPAA, GDPR, ISO 27001/27701, and SOC 2, ensuring IAM policies, encryption, and audit readiness.

  • Integrate Okta, AWS IAM, and identity federation for secure access control and role-based access management.

  • Spearhead AI automation and agentic AI use cases for infrastructure operations automating deployments, incident response, policy enforcement, and resource provisioning.

  • Collaborate with DevOps, SRE, Security, and Product teams to deliver platform capabilities that accelerate time-to-market and improve developer experience.

  • Mentor junior engineers and contribute to engineering best practices, documentation, and knowledge sharing.

Requirements

We are seeking a highly skilled and proactive Senior Platform Engineer / SRE to lead the design, automation, and operational excellence of cloud platforms and AI-driven systems.

The ideal candidate will be a seasoned engineer with hands-on experience in AWS, Terraform, Kubernetes, CI/CD, Infrastructure as Code (IaC), and Observability, who thrives in fast-paced, high-impact environments.

You will play a pivotal role in building and maintaining scalable, secure, and self-healing platforms that support enterprise-grade applications and AI-powered workflows., Bachelor s or Master s degree in Computer Science, Engineering, or related field.

6+ years of hands-on experience in DevOps, SRE, or Platform Engineering roles.

Expertise in AWS cloud services (EC2, S3, Lambda, RDS, VPC, IAM, CloudFront, etc.) and Terraform/CloudFormation for infrastructure provisioning.

  • Proven experience with Kubernetes (EKS, AKS, GKE), Docker, and container orchestration.

  • Strong proficiency in CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or similar tools.

  • Deep understanding of observability tools: Grafana, Datadog, Prometheus, Kibana, and logging frameworks.

  • Experience with incident management, RCA, and post-mortem processes in production environments.

  • Solid knowledge of security compliance frameworks: PCI-DSS, HIPAA, GDPR, CCPA, ISO 27001/27701.

  • Experience integrating identity providers (Okta, AWS IAM) and managing RBAC across cloud environments.

  • Hands-on experience with Python for automation, scripting, and tooling.

  • Familiarity with Agentic AI, AI automation, and intelligent operations (AIOps) for infrastructure and platform management.

  • Strong communication, collaboration, and leadership skills.

Preferred Qualifications:

  • Experience with multi-cloud environments (AWS, Azure, Google Cloud Platform).

  • Knowledge of serverless architectures (AWS Lambda, Azure Functions).

  • Experience with data platforms (Snowflake, Redshift, BigQuery) and ETL pipelines (Airflow, Glue).

  • Exposure to AI/ML platforms and GenAI integration in DevOps workflows.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Introduction to security advocacy and automation testing

Chris Heilmann +2 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:21 min

Deploying a primary Elasticsearch and Kibana cluster configuration

Philipp Krenn · World Congress 2022

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:37 min

Architecting single sign-on flows across multiple application domains

Gift Egwuenu · World Congress 2023

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all