> Markdown version of [/jobs/ext/2188628-senior-monitoring-sre-engineer-remote](https://www.wearedevelopers.com/jobs/ext/2188628-senior-monitoring-sre-engineer-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Monitoring/SRE Engineer (REMOTE) - **Company:** Koniag Services, Inc. - **Location:** Washington, DC, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Bash Shell, Cloud Computing, Cloud Engineering, DevOps, Monitoring of Systems, Python (Programming Language), Windows PowerShell, Reliability Engineering, Cloud Services, Datadog, Scripting, Grafana, Mttr, Information Technology, SolarWinds (Software), Cloudwatch, Splunk, Pagerduty - **Published:** August 22, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9109767/senior-monitoringsre-engineer-remote ## About the Role Koniag IT Systems (KITS) is seeking a Senior Monitoring / SRE Engineer with a minimum of 8 years of experience to lead enterprise observability, monitoring, and reliability engineering practices for a federal civilian customer's hybrid on-premises and AWS cloud environment. The ideal candidate has designed and operated enterprise monitoring platforms at scale, has strong incident management experience, and can drive site reliability practices across a large, multi-team infrastructure program., * Bachelor's degree in Computer Science, Information Technology, or related field, or equivalent professional experience * Minimum of 8 years of experience in systems monitoring, site reliability engineering, or a related operations discipline * Hands-on experience designing and administering enterprise monitoring/observability platforms (Splunk, SolarWinds, Grafana, Datadog, or similar) * Strong experience with cloud-native monitoring in AWS (CloudWatch, CloudTrail, or equivalent) * Demonstrated experience leading incident response and root cause analysis for enterprise production environments * Experience with AIOps, auto-remediation/self-healing workflows or OpenTelemetry, * Working knowledge of scripting/automation (Python, PowerShell, or Bash) for monitoring and alerting integration * Strong understanding of ITIL-aligned incident, problem, and availability management practices * Excellent written and verbal communication skills, including experience briefing technical and program leadership * Ability to work collaboratively in a fast-paced environment * Excellent communication skills and the ability to convey complex technical concepts to non-technical stakeholders * Ability to obtain public trust clearance Desired Skills and Competencies: * Experience working in a federal government IT environment * Splunk Certified Architect/Admin, AWS Certified DevOps Engineer, or equivalent monitoring/SRE certification * Experience supporting BDCR/COOP monitoring and DR failover validation for financial systems * Experience with PagerDuty, Opsgenie, or similar alert-management/on-call platforms * Familiarity with FedRAMP continuous monitoring (ConMon) reporting requirements ## Description The Senior Monitoring / SRE Engineer will be responsible for designing, implementing, and owning the enterprise monitoring and observability architecture, leading incident response and root cause analysis for high-severity outages, and supporting continuous monitoring (ConMon) reporting requirements under FISMA/NIST SP 800-53., * Design, implement, and own the enterprise monitoring and observability architecture spanning infrastructure, applications, and cloud services (e.g., Splunk, SolarWinds, Grafana, Datadog, or CloudWatch). * Define and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets for mission-critical financial and case-management systems. * Lead incident response and root cause analysis for high-severity outages, driving cross-team remediation and after-action reviews. * Build automated alerting, dashboards, and runbooks to reduce mean time to detect (MTTD) and mean time to resolve (MTTR). * Support monitoring and validation activities for the customer's business disaster continuity and recovery (BDCR) program, including synthetic transaction monitoring and failover verification. * Partner with server, cloud, storage, and database engineering teams to instrument systems and integrate telemetry into a unified observability platform. * Mentor junior SRE/monitoring engineers and establish best practices for capacity planning and performance baselining. * Support continuous monitoring (ConMon) reporting requirements under FISMA/NIST SP 800-53 in coordination with the security team. * Present operational health, reliability metrics, and improvement roadmaps to program and government leadership. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Monoskope: Developer Self-Service Across Clusters](https://www.wearedevelopers.com/videos/329-monoskope-developer-self-service-across-clusters) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Best Job Boards for Remote Work for Developers](https://www.wearedevelopers.com/magazine/290-best-job-boards-for-remote-work-for-developers)