Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)

CrowdStrike
Redmond, WA, United States
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Acceptance Test-Driven Development Software as a Service Continuous Delivery Data Structures Shard (Database Architecture) Distributed Systems Python (Programming Language) Node.Js Open Source Technology Performance Tuning
+12 more
Reliability Engineering Scala (Programming Language) Multithreading Cloud Platform System Concurrency Parallel Computation Backend Kotlin Falcon Platform Information Technology Programming Languages Microservices

Job description

Experteer Overview As a Principal SRE, you drive reliability and scalability across CrowdStrike Falcon’s core platform, partnering with product groups to define multi-year roadmaps and architect foundational services. You will hands-on engineer, re-architect critical systems, and lead initiatives that reduce toil and improve observability at scale. Your work shapes shared infrastructure and guides product teams, enabling rapid, reliable delivery of AI-enabled threat detection. This is a high-impact role with autonomy to influence architecture and delivery across the Falcon Platform. Compensation / Benefits * Define and drive multi-year reliability roadmaps with engineering leadership * Design and implement architectural improvements to services, libraries, and platforms * Develop and maintain scalable, reliable services * Extend libraries for cross-cutting cloud platform concerns * Lead reliability, scalability, performance, and cost-efficiency initiatives in large-scale systems * Establish observability practices and drive automation including continuous delivery * Define and implement SLOs and error budgets to guide prioritization * Lead performance and cost optimization through profiling and capacity planning * Conduct resilience engineering including chaos experiments and failure modelling * Implement automation and infrastructure-as-code to reduce manual toil * Provide technical leadership during incidents and post-incident retrospectives * Identify opportunities to extract common patterns into shared libraries/tools * Continuously re-evaluate architectures for improvement in performance, stability, user experience * Drive strategic technical decisions and influence infrastructure improvements * Mentor engineers and raise technical IQ across the org * Contribute to open source and advocate software engineering best practices (Go) * Collaborate across teams to own deliverables and foster accountable, energetic execution Tasks * 10+ years building and operating distributed systems and service-oriented backends at scale * 5+ years developing microservices for SaaS in a modern backend language (Go, Java, Scala, Kotlin, Python, Node.js) * Expert-level proficiency in at least one programming language, with expert-level Go or willingness to reach expert in Go * Deep understanding of distributed systems, consensus, replication, and scalability patterns * Experience scaling backend systems: sharding, partitioning, horizontal scaling, capacity planning, performance optimization * Strong grasp of multi-threading, concurrency, and parallel processing * Proven track record of architectural decisions at organizational scope * Strong systems thinking and ability to influence across boundaries * Engineering best practices: testing, peer review, resilient architectures * Thrives in fast-paced, test-driven, collaborative environment; team-oriented * Desire to ship code and see it run in production * Degree in Computer Science, or commensurate experience in data structures and algorithms * Experience utilizing AI to enhance decision-making and efficiency Key requirements * Market leader in compensation and equity awards * Wellness programs * Paid parental and adoption leaves * Professional development opportunities * Employee Networks and volunteer opportunities * Great Place to Work Certified

Requirements

Experteer Overview As a Principal SRE, you drive reliability and scalability across CrowdStrike Falcon’s core platform, partnering with product groups to define multi-year roadmaps and architect foundational services. You will hands-on engineer, re-architect critical systems, and lead initiatives that reduce toil and improve observability at scale. Your work shapes shared infrastructure and guides product teams, enabling rapid, reliable delivery of AI-enabled threat detection. This is a high-impact role with autonomy to influence architecture and delivery across the Falcon Platform. Compensation / Benefits * Define and drive multi-year reliability roadmaps with engineering leadership * Design and implement architectural improvements to services, libraries, and platforms * Develop and maintain scalable, reliable services * Extend libraries for cross-cutting cloud platform concerns * Lead reliability, scalability, performance, and cost-efficiency initiatives in large-scale systems * a, distributed systems and service-oriented backends at scale * 5+ years developing microservices for SaaS in a modern backend language (Go, Java, Scala, Kotlin, Python, Node.js) * Expert-level proficiency in at least one programming language, with expert-level Go or willingness to reach expert in Go * Deep understanding of distributed systems, consensus, replication, and scalability patterns * Experience scaling backend systems: sharding, partitioning, horizontal scaling, capacity planning, performance optimization * Strong grasp of multi-threading, concurrency, and parallel processing * Proven track record of architectural decisions at organizational scope * Strong systems thinking and ability to influence across boundaries * Engineering best practices: testing, peer review, resilient architectures * Thrives in fast-paced, test-driven, collaborative environment; team-oriented * Desire to ship code and see it run in production * Degree in Computer Science, or commensurate aaaaa aaaz_ in data structures and algorithms * Experience utilizing AI to enhance decision-making and efficiency Key requirements * Market leader in compensation and equity awards * Wellness programs * Paid parental and adoption leaves * Professional development opportunities * Employee Networks and volunteer opportunities * Great Place to Work Certified

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · WWC 2021

3:09 min

Understanding Kotlin Multiplatform and its compiler targets

Petar Marijanović · LIVE

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · WWC 2023

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:18 min

Recommended resources and frameworks for learning Kotlin natively

Iris Hunkeler · LIVE

Videos

See all

Related articles

See all