Platform Engineer IV

Insight Global
Greenwood Village, CO, United States
7 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$124,800.0 - $145,600.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Data Analysis Continuous Integration Information Engineering Data Infrastructure Data Security Linux Network Address Translation
+36 more
DevOps Distributed Computing Environment Domain Name System (DNS) Graph Database Network Topologies Identity and Access Management Key Management Nagios Neo4j Networking Basics Network Monitoring Peering Performance Tuning Prometheus Shell Script Data Streaming Management of Software Versions Computer Network Operations Grafana Apache Spark Firewalls (Computer Science) Amazon Virtual Private Cloud (VPC) Cloudformation Data Lakes Gitlab-ci Kubernetes Deployment Automation Apache Kafka Machine Learning Operations Virtual Agents Cloudwatch Terraform Splunk Data Pipelines Docker Artifactory

Job description

An Insight Global client’s Infrastructure Intelligence and Analytics (IIA) team builds and operates the data platform and AI agent infrastructure that powers proactive network monitoring and autonomous investigation for network operations. As part of this group, the Platform Engineer IV is responsible for designing, building, and maintaining the AWS infrastructure that underpins the IIA Data Lake, agent runtime environments, CI/CD pipelines, and graph database systems. This role ensures production environments are stable, scalable, and secure while enabling data science and agentic AI workloads to operate reliably at scale.

Infrastructure and Data Lake

Design and manage AWS infrastructure for the IIA Data Lake including S3 storage, Glue data catalog, Athena query engine, and EMR compute clusters.

Manage cross-account connectivity, VPC networking, security groups, and IAM roles/policies to enable secure data flow between IIA, upstream data providers, and downstream consumers.

Build and maintain infrastructure for AI agent runtime environments, including compute resources for LangGraph agents deployed via LangSmith Deployments.

Support deployment and operation of AWS Neptune for the network topology graph (digital twin), including capacity planning, schema design support, and performance tuning.

Implement and manage infrastructure-as-code (Terraform, CloudFormation) for repeatable, auditable environment provisioning.

Manage IAM access key rotations, secrets management (AWS Secrets Manager, Delinea), and security compliance for on-premises and cloud integrations (e.g., Splunk Edge Processor).

CI/CD and Agent Deployments

Build and maintain CI/CD pipelines for AI agent deployments using GitLab CI/CD, Docker, and Artifactory.

Manage container lifecycle for agents deployed via LangSmith Deployments, including image builds, versioning, and rollback procedures.

Automate deployment workflows to enable rapid, reliable promotion of agents from development through production.

Coordinate with SpecGPT platform team on AI Gateway integration, cross-account deployment, and connectivity requirements.

Production Operations

Ensure production environment stability through monitoring, alerting, and incident response. Maintain SLAs for data pipeline availability and agent uptime.

Implement production monitoring and alerting for deployed agents (health checks, error rates, latency, resource utilization).

Coordinate with upstream data teams and platform teams (SpecGPT, Splunk, Public Cloud) on connectivity, firewall requests, and integration requirements.

Support data engineering team with infrastructure needs for new data source onboarding (storage provisioning, access controls, pipeline compute).

Perform other duties as required.

This role will pay between $60- $70 per hour.

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global’s Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.

Requirements

Strong communication skills with ability to explain infrastructure decisions to non-infrastructure stakeholders

Expert-level experience with AWS services: EC2, S3, IAM, VPC, Glue, Athena, EMR, Secrets Manager, CloudWatch

Strong experience with infrastructure-as-code (Terraform preferred, CloudFormation acceptable)

Experience managing cross-account AWS architectures, VPC peering, PrivateLink, and transit gateway configurations

Experience with IAM policy design, least-privilege access patterns, and service account management

Experience with containerization (Docker) and container orchestration

Experience with CI/CD pipelines (GitLab CI preferred)

Proficiency with Linux-based operating systems and shell scripting

Experience with monitoring and alerting tools (CloudWatch, Prometheus, Grafana, or similar)

Understanding of networking fundamentals: DNS, CIDR, NAT, firewalls, security groups

Demonstrated ability to work across teams and coordinate with external platform owners on connectivity and access requirements Experience with graph databases (AWS Neptune, Neo4j) including deployment, scaling, and operational management

Experience with Apache Kafka or similar streaming platforms

Experience with Apache Spark (Scala preferred) for distributed data processing

Experience with Airflow or similar workflow orchestration platforms

Experience in the telecommunications industry or other large-scale network operations environments

Familiarity with AI/ML infrastructure requirements (model serving, GPU/CPU compute, artifact management via MLflow or similar)

Experience with Splunk integration, particularly Edge Processor and MCP connectivity

AWS certifications (Solutions Architect, DevOps Engineer, or similar)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:24 min

Comparing Neo4j and GraphQL conceptual models

William Lyon · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:30 min

Introduction to Neo4j and remote developer relations work

1:24 min

Evaluating formal AWS certifications versus raw practical engineering experience

Jan Giacomelli · LIVE

Videos

See all

Related articles

See all