Senior Monitoring/SRE Engineer (REMOTE)

Koniag Services, Inc.
Chantilly, VA, United States
4 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Bash Shell Cloud Computing Cloud Engineering DevOps Monitoring of Systems Python (Programming Language) Windows PowerShell Reliability Engineering Cloud Services Datadog Scripting
+7 more
Grafana Mttr Information Technology SolarWinds (Software) Cloudwatch Splunk Pagerduty

Job description

The Senior Monitoring / SRE Engineer will be responsible for designing, implementing, and owning the enterprise monitoring and observability architecture, leading incident response and root cause analysis for high-severity outages, and supporting continuous monitoring (ConMon) reporting requirements under FISMA/NIST SP 800-53., * Design, implement, and own the enterprise monitoring and observability architecture spanning infrastructure, applications, and cloud services (e.g., Splunk, SolarWinds, Grafana, Datadog, or CloudWatch).

  • Define and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets for mission-critical financial and case-management systems.
  • Lead incident response and root cause analysis for high-severity outages, driving cross-team remediation and after-action reviews.
  • Build automated alerting, dashboards, and runbooks to reduce mean time to detect (MTTD) and mean time to resolve (MTTR).
  • Support monitoring and validation activities for the customer’s business disaster continuity and recovery (BDCR) program, including synthetic transaction monitoring and failover verification.
  • Partner with server, cloud, storage, and database engineering teams to instrument systems and integrate telemetry into a unified observability platform.
  • Mentor junior SRE/monitoring engineers and establish best practices for capacity planning and performance baselining.
  • Support continuous monitoring (ConMon) reporting requirements under FISMA/NIST SP 800-53 in coordination with the security team.
  • Present operational health, reliability metrics, and improvement roadmaps to program and government leadership.

Requirements

Koniag IT Systems (KITS) is seeking a Senior Monitoring / SRE Engineer with a minimum of 8 years of experience to lead enterprise observability, monitoring, and reliability engineering practices for a federal civilian customer’s hybrid on-premises and AWS cloud environment. The ideal candidate has designed and operated enterprise monitoring platforms at scale, has strong incident management experience, and can drive site reliability practices across a large, multi-team infrastructure program., * Bachelor’s degree in Computer Science, Information Technology, or related field, or equivalent professional experience

  • Minimum of 8 years of experience in systems monitoring, site reliability engineering, or a related operations discipline
  • Hands-on experience designing and administering enterprise monitoring/observability platforms (Splunk, SolarWinds, Grafana, Datadog, or similar)
  • Strong experience with cloud-native monitoring in AWS (CloudWatch, CloudTrail, or equivalent)
  • Demonstrated experience leading incident response and root cause analysis for enterprise production environments
  • Experience with AIOps, auto-remediation/self-healing workflows or OpenTelemetry, * Working knowledge of scripting/automation (Python, PowerShell, or Bash) for monitoring and alerting integration
  • Strong understanding of ITIL-aligned incident, problem, and availability management practices
  • Excellent written and verbal communication skills, including experience briefing technical and program leadership
  • Ability to work collaboratively in a fast-paced environment
  • Excellent communication skills and the ability to convey complex technical concepts to non-technical stakeholders
  • Ability to obtain public trust clearance

Desired Skills and Competencies:

  • Experience working in a federal government IT environment
  • Splunk Certified Architect/Admin, AWS Certified DevOps Engineer, or equivalent monitoring/SRE certification
  • Experience supporting BDCR/COOP monitoring and DR failover validation for financial systems
  • Experience with PagerDuty, Opsgenie, or similar alert-management/on-call platforms
  • Familiarity with FedRAMP continuous monitoring (ConMon) reporting requirements

Benefits & conditions

We offer competitive compensation and an extraordinary benefits package including health, dental and vision insurance, 401K with company matching, flexible spending accounts, paid holidays, three weeks paid time off, and more.

About the company

Koniag IT Systems, LLC, a Koniag Government Services company , is seeking a Senior Monitoring/SRE Engineer to support KITS and our government customer in Washington, DC. The position is remote. This position requires the candidate to be able to obtain a Public Trust., Koniag Government Services (KGS) is an Alaska Native Owned corporation supporting the values and traditions of our native communities through an agile employee and corporate culture that delivers Enterprise Solutions, Professional Services and Operational Management to Federal Government Agencies. As a wholly owned subsidiary of Koniag, we apply our proven commercial solutions to a deep knowledge of Defense and Civilian missions to provide forward leaning technical, professional, and operational solutions. KGS enables successful mission outcomes for our customers through solution-oriented business partnerships and a commitment to exceptional service delivery. We ensure long-term success with a continuous improvement approach while balancing the collective interests of our customers, employees, and native communities. For more information, please visit www.koniag-gs.com.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:15 min

Overcoming integration hurdles in consolidated monitoring platforms

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

Videos

See all

Related articles

See all