Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)

CrowdStrike
Sunnyvale, CA, United States
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Acceptance Test-Driven Development Software as a Service Code Review Shard (Database Architecture) Distributed Systems Python (Programming Language) Node.Js Open Source Technology Reliability Engineering Scala (Programming Language)
+10 more
Multithreading Cloud Platform System Concurrency Parallel Computation Backend Kotlin Falcon Platform Information Technology Production Code Microservices

Job description

Experteer Overview As a Principal SRE, you build foundational reliability for CrowdStrike Falcon by shaping core libraries and services and embedding with product teams to scale reliability. You own architectural decisions and drive multi-cloud, high-scale systems with hands-on engineering. You influence across the Falcon Platform, champion observability, and automate toil to improve performance and resilience. This is a high-impact, engineering-heavy role with autonomy and opportunities to define standards across the organization. Compensation / Benefits * Define and drive multi-year reliability roadmaps with engineering leadership * Architecturally improve services, libraries, and platforms impacting multiple product groups * Develop and maintain scalable, reliable services * Extend libraries for cross-cutting cloud platform concerns * Lead reliability, scalability, performance, and cost-efficiency initiatives in large distributed systems * Establish observability practices and drive automation including CD * Define service-level objectives and error budgets to guide prioritization * Lead optimization efforts (profiling, bottlenecks, capacity, cloud efficiency) * Conduct resilience engineering (chaos testing, failure modeling) * Implement automation and infrastructure-as-code to reduce manual toil * Provide technical leadership during incidents and post-incident retrospectives * Identify opportunities to extract common patterns into shared libraries/tools * Continuously re-evaluate architecture for performance, stability, and developer experience * Drive strategic technical decisions and cross-organizational improvements * Mentor engineers and uplift architectural standards * Contribute to open-source practices and Go-centric engineering conventions * Collaborate across teams to own deliverables and build shared components * Build with a self-starter mindset and accountability Tasks * 10+ years of distributed systems and backend service experience at scale * 5+ years building microservices for SaaS in a modern backend language (Go, Java, Scala, Kotlin, Python, Node.js) * Expert-level proficiency in at least one language, especially Go * Deep understanding of distributed systems concepts and failure modes * Experience scaling backend systems with sharding, partitioning, and capacity planning * Strong multi-threading, concurrency, and parallel processing knowledge * Track record of architectural decisions with production impact * Influence without direct authority across org boundaries * Solid engineering practices: testing, code review, resilient architecture * Thrives in fast-paced, test-driven, collaborative environments * Desire to ship production code and see it run in production * Degree in Computer Science, or commensurate experience * Experience applying AI to improve decisions and workflows Key requirements * market-leading compensation * comprehensive wellness programs * vacation and holidays * paid parental and adoption leaves * professional development opportunities * employee networks and volunteer opportunities

Requirements

_ years building microservices for SaaS in a modern backend language (Go, Java, Scala, Kotlin, Python, Node.js) * Expert-level proficiency in at least one language, especially Go * Deep understanding of distributed systems concepts and failure modes * Experience scaling backend systems with sharding, partitioning, and capacity planning * Strong multi-threading, concurrency, and parallel processing knowledge * Track record of architectural decisions with production impact * Influence without direct authority across org boundaries * Solid engineering practices: testing, code review, resilient architecture * Thrives in fast-paced, test-driven, collaborative environments * Desire to ship production code and see it run in production * Degree in Computer Science, or commensurate experience * Experience applying AI to improve decisions and workflows Key requirements * market-leading compensation * comprehensive wellness programs * vacation and holidays * paid parental and aaaaa org leaves * professional development opportunities * employee networks and volunteer opportunities

About the company

Experteer Overview As a Principal SRE, you build foundational reliability for CrowdStrike Falcon by shaping core libraries and services and embedding with product teams to scale reliability. You own architectural decisions and drive multi-cloud, high-scale systems with hands-on engineering. You influence across the Falcon Platform, champion observability, and automate toil to improve performance and resilience. This is a high-impact, engineering-heavy role with autonomy and opportunities to define standards across the organization. Compensation / Benefits * Define and drive multi-year reliability roadmaps with engineering leadership * Architecturally improve services, libraries, and platforms impacting multiple product groups * Develop and maintain scalable, reliable services * Extend libraries for cross-cutting cloud platform concerns * Lead reliability, scalability, performance, and cost-efficiency initiatives in large distributed systems * Establish observability practices and aa

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:09 min

Understanding Kotlin Multiplatform and its compiler targets

Petar Marijanović · LIVE

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · WWC 2023

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · WWC 2021

2:18 min

Recommended resources and frameworks for learning Kotlin natively

Iris Hunkeler · LIVE

Videos

See all

Related articles

See all