> Markdown version of [/jobs/ext/2142054-monitoring-sre-engineer-remote](https://www.wearedevelopers.com/jobs/ext/2142054-monitoring-sre-engineer-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Monitoring/SRE Engineer (REMOTE) - **Company:** Koniag Services, Inc. - **Location:** Washington, DC, United States (Remote available) - **Experience:** Expert - **Salary:** $120,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Bash Shell, Cloud Computing, CompTIA Network+, Databases, Monitoring of Systems, Information Technology Operations, Python (Programming Language), Windows PowerShell, Reliability Engineering, Datadog, Scripting, Grafana, Information Technology, SolarWinds (Software), Cloudwatch, Splunk, Servicenow - **Published:** August 20, 2026 - **Apply:** https://dejobs.org/x/x/90DEDE7AD4714FBA9E4B4F47C5679EA0/job/ ## About the Role Koniag IT Systems (KITS) is seeking a Monitoring / SRE Engineer with a minimum of 5 years of experience to support enterprise monitoring, alerting, and incident response for a federal civilian customer's hybrid infrastructure environment., * Associate's or Bachelor's degree in Information Technology or related field, or equivalent professional experience * Minimum of 5 years of experience in IT operations, monitoring, or a related support role * Familiarity with enterprise monitoring/observability tools (Splunk, SolarWinds, Grafana, Datadog, or similar) * Basic understanding of cloud infrastructure concepts (AWS preferred) Required Skills and Competencies: * Strong attention to detail and ability to follow incident response procedures * Understanding of ITSM ticketing and escalation processes * Basic scripting ability (Python, PowerShell, or Bash) is a plus * Ability to obtain Public Trust clearance * Ability to work collaboratively in a fast-paced environment * Excellent communication skills and the ability to convey complex technical concepts to non-technical stakeholders Security Requirement: * Ability to obtain public trust clearance Desired Skills and Competencies: * Experience working in a federal government IT environment * CompTIA Network+, Splunk Core Certified User, or AWS Certified Cloud Practitioner * Exposure to ServiceNow or similar ITSM/on-call platforms * Prior experience on a federal government IT support contract * Interest in growing toward a Site Reliability Engineering (SRE) career path ## Description This is a strong opportunity for an operations-minded engineer to build site reliability engineering skills while supporting a large-scale infrastructure modernization program., The Monitoring / SRE Engineer will be responsible for monitoring enterprise infrastructure, application, and cloud dashboards, responding to system alerts, and supporting the build and maintenance of dashboards and alerting rules across the customer's hybrid on-premises and AWS cloud environment., * Monitor enterprise infrastructure, application, and cloud dashboards (e.g., Splunk, SolarWinds, Grafana, or CloudWatch) to identify and triage issues. * Respond to system alerts, perform initial troubleshooting, and escalate incidents per established runbooks and ServiceNow processes. * Assist in building and maintaining dashboards, alert rules, and automated notifications for infrastructure and application health. * Support root cause analysis and after-action documentation for production incidents. * Assist with synthetic monitoring and validation checks supporting the customer's business disaster continuity and recovery (BDCR) testing. * Maintain accurate incident tickets, monitoring runbooks, and knowledge base articles. * Collaborate with server, cloud, storage, and database engineers to ensure systems are properly instrumented and monitored. * Participate in an on-call rotation to support after-hours incident response. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Keycloak case study: Making users happy with service level indicators and observability](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Best Job Boards for Remote Work for Developers](https://www.wearedevelopers.com/magazine/290-best-job-boards-for-remote-work-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)