Apache Airflow Support/Operational Engineer- Remote

Palni Incorporated
United States
about 2 months ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
1 year minimum
Working hours
Regular working hours
Job source

Tech stack

Airflow Amazon Web Services Amazon Elastic Compute Cloud Continuous Integration Directed Acyclic Graph (Directed Graphs) Software Debugging Python (Programming Language) Role-Based Access Control Prometheus Grafana Software Troubleshooting Cloudformation
+4 more
Kubernetes Cloudwatch Terraform Docker

Requirements

  • 4+ years hands-on with Apache Airflow in production (not just “used it in a pipeline”) - scheduler, executors, DAG lifecycle, backfills

  • Strong troubleshooting of Airflow failures: restart loops, scheduler lag, task queue issues, executor problems, dependency conflicts

  • AWS production experience, specifically ECS on EC2 (containerized Airflow) - sizing, deployment, networking

  • Security hardening: RBAC, secrets backends, TLS, and reducing Airflow’s provider surface / CVE footprint - this is FedPoint’s core pain, so weight it heavily

  • Comfortable with CVE triage and vulnerability justification - assessing exploitability, documenting findings for compliance review

  • Python proficiency (Airflow is Python-native; needed to read/debug DAGs and providers)

  • Docker/containerization fundamentals

  • Strong written communication - client-facing support with SLA-bound response/resolution

Strongly preferred

  • Experience supporting federal / government contractor environments; familiarity with FISMA/NIST 800-53/800-171 concepts (support context, not as a certifier)

  • Kubernetes (some clients run Airflow there even if FedPoint is ECS)

  • CI/CD, infrastructure-as-code (Terraform/CloudFormation)

  • Prior SLA-based support or managed-services experience (vs. pure engineering) - knows what “2-hour response” discipline actually means

  • Monitoring: CloudWatch, Prometheus, Grafana

Soft requirements that matter for support (not engineering)

  • Calm under production-incident pressure; good bedside manner with a stressed client

  • Documentation discipline - runbooks, justification write-ups, ticket hygiene

  • Able to work a defined coverage window reliably (federal clients notice SLA misses)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all