AI Platform Engineer (Hybrid in NYC or CT)

Insight Global
Stamford, CT, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Bash Shell Cloud Computing Computer Programming Continuous Integration DevOps Distributed Systems Identity and Access Management Python (Programming Language)
+18 more
Reliability Engineering Prometheus Software Engineering Datadog Data Logging Scripting Grafana AWS Lambda Amazon Virtual Private Cloud (VPC) Cloudformation Amazon Relational Database Service AI Platforms Kubernetes Infrastructure Automation Frameworks Deployment Automation Cloudwatch Terraform Docker

Job description

We are seeking a Platform Engineer to help design, build, and operate the foundational cloud and application platforms that power our AI digital products and services. In this role, you will focus on creating reliable, secure, scalable and quality assured platforms that enable application teams to deliver software quickly and safely.

You will work closely with infrastructure, security, and application teams to provide self-service capabilities, standardized tooling, ensure quality and strong operational practices across environments.

Responsibilities

Platform & Cloud Infrastructure

  • Build and operate cloud-based platform services that support application development and runtime workloads.

  • Design and maintain infrastructure using AWS services such as EC2, EKS, ECS, S3, RDS, IAM, VPC, Lambdas, Bedrock AI services and CloudWatch.

  • Implement and manage Infrastructure as Code (IaC) using Terraform, CDK, CloudFormation, or similar tools.

  • Support containerized and non-containerized workloads across development, staging, and production environments.

Reliability, Operations & Observability

  • Ensure platform reliability, availability, and performance using DevOps and SRE best practices.

  • Implement and maintain monitoring, logging, and alerting for platform services.

  • Participate in on-call rotations and incident response, contributing to root cause analysis and continuous improvement.

  • Develop operational runbooks and automation to reduce manual workload.

Quality Assurance, Security & Governance

  • Build platforms that are secure by default, following least-privilege access and defense-in-depth principles.

  • Partner with security and compliance teams to implement required controls, policies, and auditability.

Requirements

  • 3-6+ years of experience in platform engineering, DevOps, SRE, or infrastructure engineering roles.

  • Hands-on experience with AWS in production environments.

  • Experience with Infrastructure as Code tools (Terraform preferred).

  • Familiarity with containers and orchestration (Docker, Kubernetes, or ECS).

  • Understanding of monitoring, logging, and alerting concepts.

  • Experience with scripting or programming (Python, Bash, or similar).

  • Solid understanding of networking, security, and distributed systems fundamentals. Preferred Qualifications

  • Experience operating Kubernetes platforms (EKS).

  • Familiarity with CI/CD systems and deployment automation.

  • Exposure to observability tools such as CloudWatch, Prometheus, Grafana, Datadog, or similar.

  • Experience working in large-scale or enterprise environments.

  • Interest in improving developer productivity and platform usability.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all