> Markdown version of [/jobs/ext/2697256-l2-noc-engineer](https://www.wearedevelopers.com/jobs/ext/2697256-l2-noc-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # L2 NOC Engineer - **Company:** American CyberSystems - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $85,000.0 - $99,500.0 - **Contract:** Permanent contract - **Skills:** Cloud Computing, Databases, System Configuration, Linux, Error Codes, Monitoring of Systems, Intrusion Detection and Prevention, Log Analysis, Network Architecture, Network Diagnostics, Backup and Restore, Data Logging, Network Routers, Sysadmin, Grafana, Firewalls (Computer Science), Gauge (Software), Wikis, Network Server, Dynatrace, Elk Stack - **Published:** September 3, 2026 - **Apply:** https://public-rest34.bullhornstaffing.com/rest-services/2WXP1S/query/JobBoardPost?where=id=1002631&fields=id,title,publishedCategory(id,name),address(city,state),employmentType,dateLastPublished,publicDescription,isOpen,isPublic,isDeleted ## About the Role * Experience monitoring infrastructure using various monitoring tools. Proven experience facilitating outage bridges or worked in Major Incident Management. * A minimum of 5 to 7 years of experience as an L2 Monitoring/Major Incident Management or similar role. * Good network diagnostic skills. * Basic Linux CLI and Basic sysadmin skills. * Preferred working knowledge on tools like Dynatrace, Max Gauge, Grafana, ELK Stack, and Log Management Systems. * Experienced in running outage bridges for closure or worked in Major Incident Management. · Willing to work rotational shifts including night shifts. * Ability to assess and prioritise faults and respond or escalate accordingly. * Experienced implementing service improvement techniques and procedures. * Good communicator with a natural aptitude for dealing with issues to resolution. ## Description * Manage and maintain the Client's Monitoring Systems for on-premises/Cloud entities like Network infrastructure (routers, switches, firewalls). Servers (physical and virtual), Storage systems, Applications and databases, Cloud resources, Security systems and Backup systems * Monitor Dynatrace, Max Gauge, Grafana, ELK Stack, and Log Management System * Incident detection, logging, classification, and prioritization * Incident response and resolution according to defined SLA * Proven experience facilitating outage bridges or worked in Major Incident Management. * Regular reporting on monitoring activities and incident metrics * Escalate incidents as needed to client POC. * Conduct root cause analysis for major incidents. * Recommend preventative measures * Configure and maintain log collection agents * Develop and refine log parsing rules and alert thresholds * Create and maintain error code detection rules * Maintain on-call manager rotation · Manage incident bridge infrastructure. * Document all incidents according to procedures * Correlation of logs across multiple systems and applications. * Maintenance of WIKI and technical documentation (for NOC) of processes and procedures used throughout normal operations. * Development of knowledge and skills in network and system administration, particularly about Client's architecture and platforms. * Participate in a 24x7 call-out rotation including Weekend support. Continuous Service Improvement: * Regular review of incident patterns and trends * Identification of recurring issues and root causes * Recommendations for preventative measures * Quarterly service improvement meetings * Ongoing optimization of Dynatrace, Max Gauge, Grafana, ELK Stack, and Log Management System configurations * Refinement of threshold values and error code detection rules * Log analysis pattern improvements * Incident bridge process refinement * Escalation procedure effectiveness review ## Related Videos - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Living Documentation That Can't Die](https://www.wearedevelopers.com/videos/2025-living-documentation-that-can-t-die) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Keycloak case study: Making users happy with service level indicators and observability](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)