Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)

CrowdStrike
New York, NY, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Acceptance Test-Driven Development Software as a Service Shard (Database Architecture) Distributed Systems Elasticsearch Open Source Technology Software Architecture Reliability Engineering Multithreading Google Cloud
+12 more
Cloud Platform System Concurrency Multi-Cloud Parallel Computation Backend Kubernetes Infrastructure Automation Frameworks Information Technology Cassandra Apache Kafka Oracle Cloud Infrastructure Microservices

Job description

Experteer Overview In this role you will secure and scale CrowdStrike Falcon by shaping reliability across the Core Platform and Embedded Reliability teams. You will work closely with product engineers to implement foundational libraries, services, and tooling that customers rely on at global scale. You’ll tackle complex platform challenges, drive architectural decisions, and improve observability and automation. This is a high-impact, hands-on backend engineering role that values autonomy and continuous improvement in a mission-driven security company. Compensation / Benefits * Partner with engineering leadership to define long-term reliability roadmaps * Design and implement architectural improvements to services, libraries, and platforms * Develop and maintain services meeting stringent reliability and scalability targets * Extend libraries for cross-cutting cloud platform concerns * Lead reliability, scalability, performance, and cost-efficiency initiatives in large-scale systems * Establish observability practices and leverage signals to drive automation and CD * Define service-level objectives and error budgets to guide priorities * Lead performance and cost optimization through profiling and capacity planning * Conduct resilience engineering including chaos experiments and failure injection * Automate infrastructure and pursue infrastructure-as-code to improve reliability * Provide technical leadership during incidents and ensure actionable post-incident improvements * Extract common patterns into shared libraries and collaborate with platform teams * Continuously re-evaluate architectures to improve performance, reliability, and developer experience * Drive strategic technical decisions affecting the organization’s infrastructure * Mentor engineers and raise technical standards across the org * Contribute to and evangelize Go best practices within the open source/community Tasks * 10+ years building and operating distributed systems and service-oriented backends at scale * 5+ years developing microservices for a SaaS product in a modern backend language * Expert-level proficiency in at least one language with expert-level Go is preferred * Deep understanding of distributed systems, consensus, replication, and scalability * Proven experience scaling backend systems (sharding, horizontal scaling, capacity planning) * Strong multi-threading, concurrency, and parallel processing knowledge * Track record of architectural decisions at organizational scope * Strong systems thinking and cross-team influence abilities * Thorough knowledge of engineering best practices and resilient architectures * Thrives in fast-paced, test-driven environments; team-player mentality * Desire to ship code and see it in production * Degree in Computer Science or equivalent experience * Proven experience using AI technologies to enhance decision-making and workflow efficiency * Bonus: Kubernetes, AWS, Cassandra, Kafka, Elasticsearch/OpenSearch, GCP, OCI, multi-cloud, internal platform tooling * Experience delivering or operating services across multiple clouds * Experience with cost optimization, performance engineering, chaos engineering, SLO/SLI, and open source contributions * Prior cybersecurity or intelligence field experience Key requirements * market-leading compensation and equity * wellness programs * vacation and holidays * parential and adoption leaves * professional development opportunities * employee networks and volunteer opportunities

Requirements

  • Establish observability practices and leverage signals to drive automation and CD * Define service-level objectives and error budgets to guide priorities * Lead performance and cost optimization through profiling and capacity planning * Conduct resilience engineering including chaos experiments and failure injection * Automate infrastructure and pursue infrastructure-as-code to improve reliability * Provide technical leadership during incidents and ensure actionable post-incident improvements * Extract common patterns into shared libraries and collaborate with platform teams * Continuously re-evaluate architectures to improve performance, reliability, and developer experience * Drive strategic technical decisions affecting the organization’s infrastructure * Mentor engineers and raise technical standards across the org * Contribute to and evangelize Go best practices within the open source/community Tasks * 10+ years building and operating distributed systems and service-oriented aaaa code at scale * 5+ years developing microservices for a SaaS product in a modern backend language * Expert-level proficiency in at least one language with expert-level Go is preferred * Deep understanding of distributed systems, consensus, replication, and scalability * Proven experience scaling backend systems (sharding, horizontal scaling, capacity planning) * Strong multi-threading, concurrency, and parallel processing knowledge * Track record of architectural decisions at organizational scope * Strong systems thinking and cross-team influence abilities * Thorough knowledge of engineering best practices and resilient architectures * Thrives in fast-paced, test-driven environments; team-player mentality * Desire to ship code and see it in production * Degree in Computer Science or equivalent experience * Proven experience using AI technologies to enhance decision-making and workflow efficiency * Bonus: Kubernetes, AWS, Cassandra, Kafka, Elasticsearch/OpenSearch, GCP, a the multi-cloud, internal platform tooling * Experience delivering or operating services across multiple clouds * Experience with cost optimization, performance engineering, chaos engineering, SLO/SLI, and open source contributions * Prior cybersecurity or intelligence field experience Key requirements * market-leading compensation and equity * wellness programs * vacation and holidays * parential and adoption leaves * professional development opportunities * employee networks and volunteer opportunities

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

3:08 min

Scaling semantic search with Astra DB and Apache Cassandra

David Leconte David Leconte +1 · WWC 2024

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:12 min

Choosing TypeScript for complex backend applications

Maximilian Otto Maximilian Otto · WWC 2024

Videos

See all

Related articles

See all