Senior Observability Engineer | McLean, VA 3 to 4-day onsite role

Momento USA LLC
McLean, VA, United States
about 1 month ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
9 years minimum
Working hours
Regular working hours

Tech stack

JavaScript (Programming Language) Artificial Intelligence Amazon Elastic Compute Cloud Application Performance Management Cloud Computing Cloud Engineering Elasticsearch HP Systems Insight Manager Python (Programming Language) Log Analysis Prometheus Enterprise Software Applications
+9 more
Data Ingestion Spring Cloud Reliability of Systems Performance Monitor BIG-IP Access Policy Manager (APM) Splunk Data Pipelines Dynatrace Jenkins

Job description

  • We are seeking a highly skilled Senior Observability Engineer with 9+ years of experience in designing, implementing, and optimizing enterprise observability solutions. The ideal candidate will possess deep expertise in modern observability platforms, application performance monitoring, Open Telemetry implementation and cloud technologies, with a strong focus on improving system reliability, operational efficiency, and user experience., * Analyze the existing observability solution deployed in Elastic cloud and understand the gap.
  • Document ideal scenario versus existing deployment and recommend the changes required to bring the Observability solution to improve overall application monitoring
  • implement end-to-end observability solutions for distributed and cloud-native applications; Work with development, Infrastructure and application support team to streamline the application monitoring using Elastic Cloud
  • Develop comprehensive monitoring strategies covering infrastructure, applications, logs, traces, metrics, and user experience.
  • Migrate application monitoring from legacy monitoring platforms (Dynatrace, Splunk, Prometheus ) to modern observability platforms such as Dynatrace and Elastic.
  • Design and implement Elastic-based monitoring architectures, including data pipelines, storage, APM, dashboards, and advanced analytics.
  • Build custom extensions, automated workflows, and synthetic monitoring solutions using Python and JavaScript.
  • Integrate observability platforms with CI/CD pipelines (Jenkins and related tools) to automate monitoring, alerting, and incident management.
  • Configure OpenPipeline, Business Events, anomaly detection, and AI-driven analytics to improve operational visibility.
  • Optimize observability platform licensing, data ingestion, and storage costs while maintaining monitoring effectiveness.
  • Collaborate with Development, Infrastructure, and Operations teams to improve application reliability, performance, and operational excellence.
  • Conduct dashboard reviews, monitoring assessments, and observability maturity improvements across enterprise applications.
  • Support proactive monitoring, root cause analysis, incident response, and continuous service improvement initiatives.

Requirements

  • 9+ years of experience in Application Performance Monitoring (APM), Observability, and Monitoring Engineering.
  • Excellent knowledge about Deploying observably using Open Telemetry framework
  • Strong expertise in onboarding application in Elastic Search (

) * Strong knowledge of:

  • Distributed tracing
  • Log analytics
  • Infrastructure and application monitoring
  • Synthetic monitoring
  • Real User Monitoring (RUM)
  • Experience developing automation using Python and JavaScript.
  • Experience integrating monitoring platforms with Jenkins and CI/CD pipelines.
  • Hands on experience in implementing Observability for Container based application
  • Strong understanding of observability architecture, SRE principles, and cloud-native monitoring practices.
  • Hands-on experience with AWS cloud platforms.
  • Knowledge of anomaly detection, Open Pipeline configuration, Business Events, and AI-assisted observability.
  • Strong analytical, troubleshooting, and root cause analysis skills.
  • Experience designing scalable enterprise observability architectures.
  • Knowledge of license optimization and observability cost management.
  • Experience implementing AI-driven observability and automated incident management.
  • Excellent communication, stakeholder management, and cross-functional collaboration skills.
  • Passion for driving operational excellence through automation, proactive monitoring, and observability best practices.

About the company

Momento USA is a global technology consulting, talent acquisition and creative development firm that addresses clients most pressing needs and challenges., National Minority Certified by NMSDC One of the fastest growing company in NJ Awarded fastest growing Asian American business by Diversitybusiness.com E-verified Company Information transmitted by this e-mail is proprietary to Momento USA and/ or its Customers and is intended for use only by the individual or entity to which it is addressed, and may contain information that is privileged, confidential or exempt from disclosure under applicable law. If you are not the intended recipient or it appears that this mail has been forwarded to you without proper authority, you are notified Note: Momento USA is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, pregnancy, sexual orientation, gender identity, national origin, age, protected veteran status, or disability status.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · World Congress 2023

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · World Congress 2025

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

45 sec

Introduction to easy mode observability and ShiftMon

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

Videos

See all

Related articles

See all