Senior ML Observability Engineer

Everforth Ecs
Fairfax, VA, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Comptia Cloud+ Application Programming Interfaces (APIs) Artificial Intelligence Systems Engineering Cloud Computing Data Infrastructure Machine Learning NIPRNet Prometheus Zero Trust Network Access AI Infrastructure Istio
+5 more
Delivery Pipeline Grafana Cloudwatch Splunk Data Pipelines

Job description

The Senior ML Observability Engineer architects and governs the instrumentation and telemetry infrastructure needed to ensure production AI and machine learning models deployed across WDP’s multi-enclave environment perform reliably and securely at mission scale. This role is essential to maintaining real-time visibility into model behavior, pipeline execution, and cross-domain access interactions in direct support of Combatant Command and Joint Staff decision-making needs.

  • Designs, implements, and governs observability and instrumentation architectures supporting AI and machine learning model-serving operations across Unclassified, Secret, and Top Secret enclaves within the War Data Platform (WDP) Core Integration enterprise.
  • Develops semantic conventions, runtime instrumentation patterns, and telemetry pipelines that generate latency metrics, error signatures, throughput indicators, model-specific performance signals, and operational readiness measurements for deployed models and serving surfaces.
  • Integrates observability capabilities into existing data pipelines, model-deployment workflows, API access patterns, and serving runtime frameworks to provide mission-relevant monitoring aligned with Combatant Command and Joint Staff decision-support needs.
  • Configures and validates instrumentation using platforms such as OpenTelemetry, Prometheus, Grafana, Elastic, Splunk, Amazon CloudWatch, and service mesh telemetry components to deliver real-time visibility into model behavior, cross-domain access interactions, and pipeline execution characteristics.
  • Conducts observability readiness reviews, supports test and evaluation gates, and collaborates with cybersecurity personnel to embed anomaly-detection signals aligned with Zero Trust and DoW cyber standards.
  • Works with serving engineers, pipeline engineers, platform teams, and external provider integration engineers to maintain observability consistency across enclaves and resolve domain-specific telemetry constraints.
  • Produces observability standards, instrumentation specifications, dashboards, alerting configurations, and performance analysis reports that strengthen reliability, accelerate incident response, and reinforce mission assurance for production model access across all security networks.
  • Performs other duties as assigned.

Requirements

Do you have experience in Tooling?, * Current Secret security clearance with the ability to obtain and maintain a Top Secret (TS) security clearance with Sensitive Compartmented Information (SCI).

  • 10 or more years of progressive experience in systems engineering, platform operations, or ML/AI infrastructure roles, with a demonstrated focus on observability, telemetry, and monitoring in classified or federal government cloud environments.
  • Hands-on experience designing and implementing observability pipelines using industry-standard tooling such as OpenTelemetry, Prometheus, Grafana, Elastic, Splunk, or Amazon CloudWatch, including instrumentation of AI/ML model-serving runtimes and data pipelines.
  • Experience operating across multi-enclave environments, including NIPRNet, SIPRNet, and JWICS, with demonstrated ability to adapt telemetry and observability architectures to cross-domain constraints and multi-level security requirements.
  • CompTIA Cloud+ certification or equivalent, demonstrating foundational knowledge of cloud infrastructure, security, and operational monitoring standards.
  • Strong problem-solving and decision-making capabilities, with a proven ability to weigh the relative costs and benefits of potential actions and identify the most appropriate solution.
  • Highly developed interpersonal and oral/written communication skills, with the ability to effectively and professionally interact with a diverse set of stakeholders (from peers to end-users to executive management).

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · WWC Europe 2026

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

Videos

See all

Related articles

See all