Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
Experteer Overview As a Principal SRE, you build foundational reliability for CrowdStrike Falcon by shaping core libraries and services and embedding with product teams to scale reliability. You own architectural decisions and drive multi-cloud, high-scale systems with hands-on engineering. You influence across the Falcon Platform, champion observability, and automate toil to improve performance and resilience. This is a high-impact, engineering-heavy role with autonomy and opportunities to define standards across the organization. Compensation / Benefits * Define and drive multi-year reliability roadmaps with engineering leadership * Architecturally improve services, libraries, and platforms impacting multiple product groups * Develop and maintain scalable, reliable services * Extend libraries for cross-cutting cloud platform concerns * Lead reliability, scalability, performance, and cost-efficiency initiatives in large distributed systems * Establish observability practices and drive automation including CD * Define service-level objectives and error budgets to guide prioritization * Lead optimization efforts (profiling, bottlenecks, capacity, cloud efficiency) * Conduct resilience engineering (chaos testing, failure modeling) * Implement automation and infrastructure-as-code to reduce manual toil * Provide technical leadership during incidents and post-incident retrospectives * Identify opportunities to extract common patterns into shared libraries/tools * Continuously re-evaluate architecture for performance, stability, and developer experience * Drive strategic technical decisions and cross-organizational improvements * Mentor engineers and uplift architectural standards * Contribute to open-source practices and Go-centric engineering conventions * Collaborate across teams to own deliverables and build shared components * Build with a self-starter mindset and accountability Tasks * 10+ years of distributed systems and backend service experience at scale * 5+ years building microservices for SaaS in a modern backend language (Go, Java, Scala, Kotlin, Python, Node.js) * Expert-level proficiency in at least one language, especially Go * Deep understanding of distributed systems concepts and failure modes * Experience scaling backend systems with sharding, partitioning, and capacity planning * Strong multi-threading, concurrency, and parallel processing knowledge * Track record of architectural decisions with production impact * Influence without direct authority across org boundaries * Solid engineering practices: testing, code review, resilient architecture * Thrives in fast-paced, test-driven, collaborative environments * Desire to ship production code and see it run in production * Degree in Computer Science, or commensurate experience * Experience applying AI to improve decisions and workflows Key requirements * market-leading compensation * comprehensive wellness programs * vacation and holidays * paid parental and adoption leaves * professional development opportunities * employee networks and volunteer opportunities
Requirements
_ years building microservices for SaaS in a modern backend language (Go, Java, Scala, Kotlin, Python, Node.js) * Expert-level proficiency in at least one language, especially Go * Deep understanding of distributed systems concepts and failure modes * Experience scaling backend systems with sharding, partitioning, and capacity planning * Strong multi-threading, concurrency, and parallel processing knowledge * Track record of architectural decisions with production impact * Influence without direct authority across org boundaries * Solid engineering practices: testing, code review, resilient architecture * Thrives in fast-paced, test-driven, collaborative environments * Desire to ship production code and see it run in production * Degree in Computer Science, or commensurate experience * Experience applying AI to improve decisions and workflows Key requirements * market-leading compensation * comprehensive wellness programs * vacation and holidays * paid parental and aaaaa org leaves * professional development opportunities * employee networks and volunteer opportunities
About the company
Experteer Overview As a Principal SRE, you build foundational reliability for CrowdStrike Falcon by shaping core libraries and services and embedding with product teams to scale reliability. You own architectural decisions and drive multi-cloud, high-scale systems with hands-on engineering. You influence across the Falcon Platform, champion observability, and automate toil to improve performance and resilience. This is a high-impact, engineering-heavy role with autonomy and opportunities to define standards across the organization. Compensation / Benefits * Define and drive multi-year reliability roadmaps with engineering leadership * Architecturally improve services, libraries, and platforms impacting multiple product groups * Develop and maintain scalable, reliable services * Extend libraries for cross-cutting cloud platform concerns * Lead reliability, scalability, performance, and cost-efficiency initiatives in large distributed systems * Establish observability practices and aa
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs
Dev Digest 121 - AI goes offline
Where To Find Software Engineering Jobs