Director - Splunk Platform Engineering & SRE

Ampcus Inc
New York, NY, United States
10 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Data Analysis Cyber Security Computer Programming Linux Disaster Recovery Distributed Systems Identity and Access Management Python (Programming Language) Machine Learning Packet Analyzer
+21 more
Performance Tuning Role-Based Access Control Reliability Engineering Ansible Security Information and Event Management Syslog TCP/IP Scripting Enterprise Software Applications Data Ingestion Git Kubernetes Infrastructure Automation Frameworks Information Technology Low Latency Deployment Automation Data Management Splunk Data Pipelines Golang Programming Languages

Job description

You will own a large-scale, mission-critical Splunk platform supporting enterprise observability and cybersecurity. This role requires deep expertise across Linux, networking, distributed systems, data ingestion, automation, and Site Reliability Engineering. You will serve as the highest technical escalation point for complex platform issues and drive architectural improvements that enhance scalability, resilience, and operational efficiency., * Own the engineering, architecture, operations, and lifecycle management of the enterprise Splunk SIEM platform.

  • Design and scale high-throughput log and event ingestion pipelines.
  • Serve as the highest technical escalation point for critical production incidents.
  • Troubleshoot complex issues involving Linux/Unix systems, networking, distributed systems, and data ingestion pipelines.
  • Lead platform reliability, observability, capacity planning, and performance engineering initiatives.
  • Architect integrations with Kubernetes, cloud platforms, enterprise systems, and Syslog/event collection frameworks.
  • Design and maintain authentication, RBAC, and enterprise access control models.
  • Develop automation solutions using Git, Ansible, and scripting languages to reduce operational overhead.
  • Lead incident response, root cause analysis, and long-term remediation using SRE principles including SLAs, SLOs, and error budgets.
  • Plan and execute platform upgrades, resilience improvements, and disaster recovery strategies.
  • Evaluate emerging technologies including AI/ML-driven analytics and contextual data platforms.
  • Develop custom engineering solutions using Python, Go, Java, or similar programming languages.
  • Influence enterprise engineering strategy through technical leadership and hands-on expertise.
  • Mentor engineers by providing technical guidance and best practices., * Be the primary technical authority for a critical enterprise Splunk platform.
  • Resolve complex production incidents across infrastructure, operating systems, networking, and applications.
  • Improve platform scalability, reliability, and operational efficiency.
  • Reduce manual operational effort through automation.
  • Strengthen enterprise security posture and governance.
  • Elevate engineering standards through technical leadership and innovation.

Requirements

  • Bachelor’s degree in Computer Science or related discipline (or equivalent experience); advanced degree preferred.
  • 12+ years of experience in cybersecurity, information security, or related technology disciplines.
  • Financial services or securities industry experience is a plus.
  • Deep hands-on experience administering and engineering enterprise Splunk platforms and SIEM environments.
  • Strong knowledge of Site Reliability Engineering (SRE) practices and distributed systems.
  • Expert-level Linux/Unix administration including performance tuning and troubleshooting.
  • Strong networking expertise including TCP/IP, packet analysis, latency troubleshooting, and Syslog.
  • Experience designing and operating high-volume data ingestion platforms.
  • Advanced proficiency with Splunk SPL and data analysis.
  • Strong programming skills in Python, Go, Java, or similar languages.
  • Experience with Git, Ansible, and DevOps automation practices.
  • Hands-on Kubernetes and container platform experience.
  • Strong understanding of enterprise security, RBAC, identity management, and governance.
  • Proven ability to independently resolve highly complex production issues.
  • Passion for automation, scalability, and operational excellence.
  • Technical understanding of AI technologies and Model Context Protocol (MCP) concepts.

Preferred Skills

  • Enterprise observability platforms
  • Performance engineering
  • Capacity planning
  • Infrastructure automation
  • Disaster recovery
  • Platform resilience engineering
  • AI/ML technologies for security analytics
  • Distributed system architecture

About the company

Ampcus Inc. is a certified global provider of a broad range of Technology and Business consulting services. We are in search of a highly motivated candidate to join our talented Team.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · WWC 2022

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all