> Markdown version of [/jobs/ext/1245252-sr-engineer-ii-epics-ng-siem-hybrid](https://www.wearedevelopers.com/jobs/ext/1245252-sr-engineer-ii-epics-ng-siem-hybrid). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Engineer II - EPICS, NG-SIEM (Hybrid) - **Company:** CrowdStrike, Inc. - **Location:** UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon Web Services, Microsoft Azure, Bash Shell, C++ (Programming Language), Software as a Service, Cloud Engineering, Cyber Security, Data as a Services, Disaster Recovery, Distributed Systems, Global Positioning Systems (GPS), Python (Programming Language), Reliability Engineering, Service-Oriented Architecture, Security Information and Event Management, Software Engineering, System Programming, Scripting, Google Cloud, Deployment Automation, Apache Kafka, Serverless Computing - **Published:** July 12, 2026 - **Apply:** https://www.totaljobs.com/job/senior-engineer/crowdstrike-job107675082 ## About the Role * A passion for reliability engineering and curiosity about how large-scale running systems behave under pressure; * 10+ years of experience in software engineering, site reliability engineering, or platform engineering, with significant time spent on large-scale distributed systems, and the ability to make pragmatic tradeoffs between short-term delivery needs and long-term platform goals; * Strong proficiency in at least one systems programming language (Go, Java, Rust, or C++) and one scripting language (Python, Bash); * Deep experience with end-to-end observability - building monitoring pipelines, defining SLIs/SLOs, and creating dashboards that drive actionable insights across multi-service architectures; * Demonstrated ability to diagnose and resolve complex incidents spanning multiple distributed components operating 24/7; * Experience with coordinated capacity planning and scaling for systems with significant infrastructure footprints; * Hands-on experience with streaming platforms (Kafka or similar) and understanding of backpressure, partition management, and consumer group dynamics at scale; * Familiarity with infrastructure-as-code, CI/CD pipelines, and automated deployment practices; * A can-do attitude - you thrive collaborating in a team and are not afraid of taking on responsibilities; * Strong written and verbal communication skills - you will lead incident communications and produce post-incident analyses that drive lasting improvements; * Comfort working across time zones with globally distributed teams., * Experience in a similar reliability or platform engineering role at a hyperscaler (AWS, Azure, GCP) or large-scale SaaS provider; * Track record of building automated remediation and self-healing infrastructure; * Experience with cost modeling and unit economics for large compute and storage footprints; * Familiarity with cloud-native architectures and serverless computing paradigms; * Hands-on experience operating platforms processing over 1 trillion events per day or more than 10 PB of data per day; * Exposure to or experience with Log Management, cybersecurity products, or security operations workflows; * Experience with disaster recovery planning and execution for multi-region systems. ## Description As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn't changed - we're here to stop breaches, and we've redefined modern security with the world's most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're also a mission-driven company. We cultivate a culture that gives every CrowdStriker both the flexibility and autonomy to own their careers. We're always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you., * End-to-end observability: Design, build, and maintain monitoring and synthetic test suites that provide deep visibility into the health of the entire NG-SIEM pipeline - from ingest through search and workflow execution - enabling rapid root cause analysis across component boundaries. * Coordinated scaling: Engineer orchestrated scaling solutions that treat the NG-SIEM pipeline as a unified system, proportionally increasing resources across all dependent components (Kafka, ingest pipelines, downstream services) to eliminate cascading bottleneck patterns. * Incident response engineering: Serve as a subject matter expert during platform-wide incidents (P2 and above), applying cross-service knowledge to diagnose and resolve multi-component failures. Partake in follow-the-sun on-call rotations, providing incident commander coordination for critical platform-wide events. * Capacity planning and cost management: Build and refine models for end-to-end capacity forecasting that account for all pipeline dimensions, including partner team dependencies (data services, GPS). Develop tooling to continuously track and surface cost drivers across the platform. * Automation and runbooks: Transform manual standard operating procedures into automated remediation workflows - including pipeline-wide scaling responses, CID rebalancing, and infrastructure healing - with the goal of resolving issues before customers are impacted. * Cross-team collaboration: Partner with cell-level teams, product engineering, GDI/3PI, and external stakeholders (e.g., CSM) to triage SLO breaches, drive problem management for large reliability efforts, and ensure consistent communication during incidents. * Platform improvements: Use your broad NG-SIEM knowledge to identify and drive systemic improvements across teams, contributing to the platform's long-term resilience and efficiency. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [The Cloud is Calling: Answer with In-Demand Skills](https://www.wearedevelopers.com/videos/945-the-cloud-is-calling-answer-with-in-demand-skills) - [Old tools, new tricks](https://www.wearedevelopers.com/videos/1916-old-tools-new-tricks) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Cloud Run- the rise of serverless and containerization](https://www.wearedevelopers.com/videos/106-cloud-run-the-rise-of-serverless-and-containerization) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)