Site Reliability Engineer - Monitoring and Anomaly Detection (Monetization)

GitLab
United States
2 days ago
Apply on job-boards.greenhouse.io
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Languages
English

Tech stack

Artificial Intelligence Data Integrity Data Stores Python (Programming Language) Machine Learning Ruby on Rails Reliability Engineering Prometheus Salesforce.Com Systems Integration Grafana Change Data Capture
+4 more
Backend Gitlab Vertica Zuora

Job description

The Observability, Monitoring, and Integrations team sits within GitLab’s Monetization section and owns end-to-end observability, monitoring, and detection across the systems that power how customers buy and use GitLab. As a Backend Engineer on the Fulfillment Workflow Monitoring (Catch All) team, you’ll build and operate telemetry, detection, and reconciliation tooling that identifies billing, data, and event anomalies across CustomersDot, Salesforce, and Zuora before they affect revenue or the customer experience., This is a greenfield team, and you’ll help build it from the ground up. You’ll shape its operating rhythm, incident handling model, and quality bar while using telemetry, artificial intelligence, and machine learning to proactively detect and resolve issues. Because Monetization changes can directly affect revenue and often involve confidential or financial information, you’ll use strong judgment, careful rollout practices, and a high bar for reliability. We’re opening this role at the Senior Software Engineer and Staff Software Engineer levels.

  • Automated billing anomaly detection for event drops, abuse spikes, and transactional discrepancies
  • End-to-end telemetry, tracing, and reconciliation across CustomersDot, Salesforce, Zuora, and usage and billing pipelines

What you’ll do

  • Design, build, and operate metrics, logs, and traces across the Monetization stack using tools such as Prometheus and Grafana.
  • Implement automated detection for billing, data, and event anomalies, and route alerts to the designated feature teams.
  • Develop reconciliation and data integrity checks across usage and billing pipelines.
  • Define and track service reliability targets (SLOs) and service level indicators (SLIs) to measure reliability, write runbooks, and take part in incident handling.
  • Explore artificial intelligence and machine learning techniques to predict system anomalies and accelerate resolution.
  • Review merge requests and offer feedback to other Monetization engineers.
  • Collaborate with Product, Finance, Support, and other partners to turn operational needs into reliable tooling., You’ll report to the Engineering Manager, Observability, Monitoring, and Integrations. Our team works asynchronously, with regular planning and retrospectives, and we jointly manage deployments, incident management, testing, and safe rollouts. For more on how Monetization works, see [Link: Monetization Sub-department Handbook]., What is your current country of residence?* Select… Are you subject to any employment agreements and/or post-employment restrictions with your current employer or a past employer?* Select… It is important to us to create an accessible and inclusive interview experience. Please let us know if there are any adjustments we can make to assist you during the hiring and interview process. What is your GitLab username? Will you now or in the future require sponsorship for a visa to remain in your current location?* Select… Have you previously worked at or consulted for GitLab?* Select… How many years of professional, hands-on experience do you have building production applications in Ruby on Rails or Python?* Select… Is your current location Bangalore?* Select… Which of the following do you have professional experience with?* Select…

Voluntary Self-Identification

For government reporting purposes, we ask candidates to respond to the below self-identification survey. Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiring process or thereafter. Any information that you do provide will be recorded and maintained in a confidential file.

As set forth in GitLab’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law. Gender Select… Are you Hispanic/Latino? Select… Race & Ethnicity Definitions

If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection. As a government contractor subject to the Vietnam Era Veterans Readjustment Assistance Act (VEVRAA), we request this information in order to measure the effectiveness of the outreach and positive recruitment efforts we undertake pursuant to VEVRAA. Classification of protected categories is as follows:

A “disabled veteran” is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service-connected disability.

A “recently separated veteran” means any veteran during the three-year period beginning on the date of such veteran’s discharge or release from active duty in the U.S. military, ground, naval, or air service.

An “active duty wartime or campaign badge veteran” means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense.

An “Armed forces service medal veteran” means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985. Veteran Status Select…

Requirements

  • Professional experience with Ruby on Rails.
  • A background in site reliability or observability engineering, including monitoring, alerting, SLOs, SLIs, runbooks, incident handling, and tools such as Prometheus, Grafana, and OpenTelemetry.
  • Experience building anomaly detection, monitoring, or risk management tooling.
  • Exposure to data stores for reporting and insights, especially ClickHouse, and the change data capture and event streaming pipelines that feed them, such as NATS JetStream.
  • Experience with billing, financial, or other business-critical systems.
  • Experience owning a project from concept to production, including proposal, discussion, execution, and monitoring.
  • Clear, concise communication about complex technical, architectural, and organizational problems in English, with the ability to propose thorough, iterative solutions in a remote and largely asynchronous work environment.
  • Additional relevant experience includes working knowledge of Python for anomaly detection and data work, and experience with Zuora or Salesforce.

About the company

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster.

The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software.

*Fortune 500ÂŽ is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on job-boards.greenhouse.io
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:17 min

Overcoming limitations in current machine learning monitoring tools

David Mosen ¡ World Congress 2021

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet ¡ LIVE

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler ¡ LIVE

2:24 min

Selecting monitoring infrastructure stacks for machine learning operations

Lina Weichbrodt ¡ LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz ¡ World Congress 2025

4:54 min

Implementing geographic salary tiers for compensation equity and fairness

Rudi Bauer Rudi Bauer +1 ¡ Cappuccino with HR

Videos

See all

Related articles

See all