NOC Engineer / SRE

NICE Ltd.
Palakkad, United States
5 days ago
Apply on www.thejobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Shift work

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Computer Programming Domain Name System (DNS) Python (Programming Language) Linux System Administration Networking Basics Reliability Engineering Ansible Prometheus
+15 more
TCP/IP Datadog Scripting Google Cloud Load Balancing Computer Network Operations Grafana Mttr Containerization Kubernetes Cloudwatch Terraform Splunk Docker Golang

Job description

At NiCE, we don't limit our challenges. We challenge our limits. Always. We're ambitious. We're game changers. And we play to win. We set the highest standards and execute beyond them. And if you're like us, we can offer you the ultimate career opportunity that will light a fire within you.

So, what's the role all about?

The SRE - NOC role sits at the intersection of traditional Network Operations Center (NOC) responsibilities and engineering-driven reliability practices. This role focuses on 24/7 service reliability, incident response, operational automation, and observability, while actively reducing operational toil through software and automation.

Unlike a traditional NOC analyst, an SRE-NOC is expected to engineer problems away, not just respond to alerts.

How will you make an impact?

Incident Response & Operations

  • Act as a primary or escalation responder in a 24x7 on-call rotation
  • Lead or support Major Incident (MI) response, including triage, mitigation, and resolution
  • Coordinate across Engineering, Infrastructure, Security, and Product teams
  • Execute and improve runbooks, playbooks, and escalation paths
  • Drive blameless post-incident reviews (PIRs) and track corrective actions

Monitoring, Alerting & Observability

  • Own service health monitoring across infrastructure, applications, and dependencies
  • Design and maintain alerting strategies that align with SLIs/SLOs
  • Reduce alert fatigue through signal-to-noise improvements
  • Build dashboards using tools such as:
    • Grafana
    • Prometheus
    • Datadog / Splunk / CloudWatch

Reliability Engineering & Automation

  • Automate repetitive operational tasks to reduce manual toil
  • Improve mean time to detect (MTTD) and mean time to resolve (MTTR)
  • Develop scripts and tools (Python, Bash, Go, etc.) to support NOC/SRE workflows
  • Implement self-healing and auto-remediation where possible
  • Partner with engineering teams to improve system design for reliability

Platform & Infrastructure Support

  • Support and troubleshoot:
    • Linux-based systems
    • Cloud platforms (AWS, Azure, GCP)
    • Kubernetes / containerized environments
  • Assist with capacity planning and availability reviews
  • Ensure operational readiness for production releases

Have you got what it takes?

Technical

  • Strong Linux systems administration
  • Experience with incident management and production support
  • Familiarity with:
    • Cloud infrastructure (AWS preferred)
    • Containers & orchestration (Docker, Kubernetes)
    • Monitoring/alerting platforms
  • Scripting or programming experience in Python, Bash, Go, or similar
  • Understanding of networking fundamentals (DNS, TCP/IP, load balancing)

Operational

Requirements

  • Experience working in 24x7 NOC or production operations environments
  • Ability to handle high-pressure incidents calmly and effectively
  • Strong written and verbal communication for incident coordination
  • Comfort working from runbooks-but improving them when they fall short
  • </ul>

    Preferred / Differentiators

    • Experience defining or operating to SLOs / SLIs
    • Prior migration from traditional NOC * SRE model
    • Infrastructure as Code experience (Terraform, Ansible, etc.)
    • Exposure to security, compliance, or regulated environments

    About the company

    NICE Ltd. (NASDAQ: NICE) software products are used by 25,000+ global businesses, including 85 of the Fortune 100 corporations, to deliver extraordinary customer experiences, fight financial crime and ensure public safety. Every day, NiCE software manages more than 120 million customer interactions and monitors 3+ billion financial transactions.

    Known as an innovation powerhouse that excels in AI, cloud and digital, NiCE is consistently recognized as the market leader in its domains, with over 8,500 employees across 30+ countries.

    Apply for this position

    This job is hosted externally. Click below to view the full posting and apply.

    Apply on www.thejobnetwork.com
    Prepare application

    Good distractions

    Talks and stories from around this role — technically off-topic, practically not.

    3:50 min

    Navigating specialized roles and toolsets across engineering teams

    Nele Uhlemann · World Congress 2023

    2:07 min

    Inspecting default bridge architectures and custom Docker networks

    Oliver Seitz Oliver Seitz · World Congress 2025

    1:08 min

    Building solutions with open source GoLang infrastructure tools

    Jad Wahab · LIVE

    3:08 min

    Aligning engineering processes with core business impact metrics

    Chris Riley · World Congress 2021

    2:34 min

    Docker sandbox architecture and microVM environment integration

    Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

    6:16 min

    Event-driven Golang backend architecture and cloud deployment

    Irina Branovic Irina Branovic · World Congress 2026 Europe

    Videos

    See all

    Related articles

    See all