Sr Engineer, SRE TechOps CICD (Remote)

CrowdStrike
United States
2 days ago
Apply on crowdstrike.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$140,000.0 - $215,000.0
Working hours
Regular working hours

Tech stack

Active Directory Artificial Intelligence Airflow Amazon Web Services Microsoft Azure Bash Shell Big Data Code Review Continuous Integration Data Infrastructure Extract Transform Load (ETL) Distributed Systems
+46 more
Domain Name System (DNS) Github Python (Programming Language) Network Troubleshooting PostgreSQL Machine Learning Automation of Marketing Windows Servers MongoDB MySQL Routing Oracle (Applications) Performance Tuning Windows PowerShell Queueing Systems RabbitMQ Redis Reliability Engineering Ansible Prometheus Software Engineering Datadog Scripting Google Cloud Load Balancing Cloud Platform System Grafana Apache Spark Gitlab Build Management Gitlab-ci Kubernetes Storage Technologies Information Technology Cassandra Apache Kafka Non-relational Database Bitbucket Puppet Firewall Services Module Terraform Splunk New Relic (SaaS) Software Version Control Jenkins Golang

Job description

As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed - we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We’re proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We’re always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you.

About the Role: CrowdStrike’s internal SRE team owns the automation, reliability, and observability of the internal developer platform that thousands of CrowdStrike engineers rely on to build and deploy software rapidly, efficiently, and at scale. We provide the resilient infrastructure and operational rigor that let product, platform, and application teams ship with confidence, without needing to think about the systems underneath them.

We’re looking for a Senior Site Reliability Engineer on the Technical Operations team to bring deep, hands-on expertise across load balancers, relational and non-relational databases, message queues (Kafka, Pulsar, RabbitMQ, RedPanda), and caching layers (Redis/Valkey, Varnish), paired with strong SLI/SLO instincts. This is a technical-anchor role: you’re the person the team routes hard problems to, and you set the bar for operational excellence through the quality of your own work and the judgment you bring to design and incident reviews. You’ll build relationships with technical leaders across the organization, contribute to architectural direction for the services you own, and help position those services for the company’s next level of scale.

What You’ll Do:

  • Own the availability and health of key services within the CICD environment, maintaining a holistic view of system health across the platform.
  • Build software and systems to manage platform infrastructure and applications, and drive automation for service deployment and operational workflows.
  • Carry on-call responsibility for owned services; drive incident response and blameless postmortems to root cause.
  • Gather and analyze metrics from operating systems and applications to support performance tuning and root cause analysis.
  • Lead system design discussions, production readiness reviews, and capacity planning exercises.
  • Evaluate and integrate agentic and AI-assisted workflows into existing team processes, and help teammates adopt them.
  • Mentor mid-level and junior engineers through code review, design pairing, and incident retrospectives.
  • Investigate and evaluate emerging technologies, and provide recommendations that support future roadmap goals.
  • Build and maintain automated reporting on service health and compliance.
  • Provide technical feedback and guidance on projects outside your core area of ownership, helping raise the bar across the broader engineering organization.
  • Partner with peer senior engineers and engineering leaders to drive cross-team reliability improvements.
  • Contribute to the Embedded SRE model, helping strengthen partnerships between SRE and the services teams.

Requirements

  • Must be eligible for CJIS clearance (requires U.S. citizenship or Green Card/permanent resident status).
  • 10+ years of experience working in large-scale production SRE or infrastructure environments.
  • 3+ years of experience leveraging and integrating AI-assisted workflows to increase engineering efficiency.
  • Bachelor’s degree in computer science or another highly technical, scientific discipline, or equivalent work experience.
  • On-premise and cloud expertise deploying and operating CI/CD tools (Bazel, Jenkins, GitLab CI, GitHub Actions), IaC provisioning (Ansible, Chef, Puppet, Salt, Terraform), source code management (Bitbucket, GitHub, GitLab), and monitoring/observability platforms (Datadog, Grafana, Humio/LogScale, Honeycomb, New Relic, Prometheus, Splunk).
  • Experience creating, deploying, operating, and scaling applications on Kubernetes.
  • Extensive experience deploying and managing data infrastructure at scale (Cassandra, Postgres, MySQL, MongoDB, OpenSearch, Kafka, Redis/Valkey).
  • Proficiency in common scripting languages (Python, Go, Bash, PowerShell).
  • Experience with storage technologies (SAN, NAS, NFS, Object Storage).
  • Experience architecting and deploying big data systems.
  • Security-first mindset with a working understanding of cybersecurity principles.
  • Proven ability to make well-informed, timely decisions under ambiguity.
  • Ability to balance short-term operational needs against long-term strategic goals.
  • Self-directed learner who takes initiative in fast-moving environments.
  • Must be able to work with a distributed team across multiple time zones.
  • Meticulous attention to detail.
  • Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.

Bonus Points:

  • Knowledge of networking patterns and general network troubleshooting (Load balancers, DNS, VIPS, Routing, Firewall rules)
  • Knowledge and proven operation ability across multiple cloud hyperscalers such as AWS, Azure, GCP, Oracle
  • Experience building self-service / provisioning-automation platforms that reduce operational toil.
  • Experience with Active Directory / Windows Server and hybrid on-prem + cloud environments.
  • Experience with data science, machine learning, and ETL, principles and tooling such as Apache Airflow, Apache Spark, ect

Benefits & conditions

CrowdStrike, Inc. is committed to fair and equitable compensation practices. Placement within the pay range is dependent on a variety of factors including, but not limited to, relevant work experience, skills, certifications, job level, supervisory status, and location. The base salary range for this position for all U.S. candidates is $140,000 - $215,000 per year, with eligibility for bonuses, equity grants and a comprehensive benefits package that includes health insurance, 401k and paid time off.

About the company

This role will require the candidate to periodically undergo and pass additional background and fingerprint check(s) consistent with government customer requirements.

Benefits of Working at CrowdStrike:

  • Market leader in compensation and equity awards
  • Comprehensive physical and mental wellness programs
  • Competitive vacation and holidays for recharge
  • Paid parental and adoption leaves
  • Professional development opportunities for all employees regardless of level or role
  • Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections
  • Vibrant office culture with world class amenities
  • Great Place to Work Certified across the globe

CrowdStrike is proud to be an equal opportunity employer. We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed. We support veterans and individuals with disabilities through our affirmative action program.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on crowdstrike.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · World Congress 2022

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

Videos

See all

Related articles

See all