Infrastructure Operations Engineer

Insight Global
Vienna, VA, United States
2 months ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon Cloudfront Amazon Elastic Compute Cloud Amazon S3 Command-Line Interface Cloud Computing Security Cloud Engineering Computer Networks Continuous Integration Domain Name System (DNS) Drupal Monitoring of Systems
+23 more
Identity and Access Management Subnetting Virtual Private Networks (VPN) Routing Ansible Datadog Data Logging Delivery Pipeline Amazon Virtual Private Cloud (VPC) Cloudformation Amazon Relational Database Service Containerization Gitlab-ci Kubernetes Route53 Cloudwatch Terraform Splunk New Relic (SaaS) Devsecops Docker Elk Stack Jenkins

Job description

  • Provide Tier 3 support for complex infrastructure and application-related incidents

  • Monitor system health, performance metrics, application logs, and infrastructure telemetry

  • Troubleshoot and resolve production issues across AWS infrastructure and Drupal-based platforms

  • Support AWS cloud services including compute, storage, networking, and security components

  • Investigate and diagnose performance bottlenecks, resource constraints, and configuration issues

  • Support CI/CD pipeline operations and troubleshoot deployment or release failures

  • Perform root cause analysis for recurring incidents and implement preventive measures

  • Coordinate incident response and resolution with development, DevSecOps, security, and infrastructure teams

  • Execute routine maintenance tasks including patching, scaling, backups, and system updates

  • Support deployment activities and release verification in production environments

  • Manage user support tickets and ensure timely resolution within SLA requirements

  • Maintain and update technical documentation for operational procedures and known issues

  • Implement and maintain monitoring alerts, logging, and automated health checks

  • Support disaster recovery testing and business continuity planning

  • Ensure compliance with federal security requirements and audit controls

  • Interface with federal stakeholders on operational status, issue escalation, and resolution

  • Collaborate with AWS support and third-party vendors for escalated technical issues

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global’s Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.

Requirements

  • 8+ years of total professional experience

  • 5+ years in infrastructure operations, cloud engineering, or production support roles

  • Prior or current experience supporting government programs

  • Strong technical knowledge and expertise in:

  • AWS core services (EC2, S3, RDS, VPC, ELB/ALB, CloudFront, Route53)

  • Cloud security services (IAM, Security Groups, KMS, CloudTrail, GuardDuty)

  • Infrastructure monitoring and observability (CloudWatch, Datadog, New Relic, or similar)

  • Infrastructure as Code (Terraform, CloudFormation, Ansible)

  • CI/CD pipeline operations (Jenkins, GitLab CI, AWS CodePipeline)

  • Linux/Unix system administration and command-line tools

  • Networking concepts (VPCs, subnets, routing, VPNs, DNS)

  • Log aggregation and analysis (CloudWatch Logs, ELK stack, Splunk)

  • Container technologies (Docker, ECS, EKS, Kubernetes)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:29 min

Transitioning from content management to cloud infrastructure

Matt Butcher · World Congress 2023

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:02 min

Asking about infrastructure and engineering culture operations

Karol Rogowski · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all