SRE

ALL JOBS 8,LLC
Alpharetta, GA, United States
16 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Build Automation Bash Shell Cloud Computing Continuous Integration Issue Tracking Systems Python (Programming Language) Nagios Reliability Engineering Datadog Scripting Grafana Mttr
+6 more
Containerization Splunk Dynatrace Servicenow Golang Programming Languages

Job description

Own the reliability, resiliency and availability of the Embedded Finance platform, proactively identifying and mitigating risks to service continuity. Design, implement and maintain comprehensive monitoring and alerting frameworks leveraging Splunk, Dynatrace, Grafana and Datadog to provide end-to-end observability across the platform. Define and track service level objectives (SLOs), service level indicators (SLIs) and error budgets to measure and improve platform health. Lead and participate in incident response, serving as a technical driver during remediation calls and coordinating with impacted and impacting technical and product teams. Own and advance the root cause analysis (RCA) process - investigating incidents, documenting the sequence of events and remediating actions, and clearly identifying underlying root causes to prevent recurrence. Ensure timely creation and management of incident tickets (e.g., ServiceNow) and accurate incident tracking, aging and reporting. Build automation and tooling to reduce toil, improve mean time to detection (MTTD) and mean time to resolution (MTTR), and increase operational efficiency. Collaborate with engineering, product and risk stakeholders to embed reliability best practices into the platform lifecycle.

Requirements

Hands-on experience with monitoring, observability and alerting tools, specifically Splunk, Dynatrace, Grafana and Datadog. Proven experience operating and supporting a large-scale enterprise platform environment. Demonstrated experience with incident response and leading or contributing to root cause analysis (RCA) processes. Strong understanding of reliability engineering principles, including availability, resiliency, monitoring and alerting best practices. Experience with ticketing and incident management workflows (e.g., ServiceNow). Excellent communication skills, with the ability to drive remediation efforts and collaborate across technical, product and risk teams. What would be great to have: Experience in financial services, payments or embedded finance environments. Proficiency with scripting or programming languages for automation (e.g., Python, Go, Bash). Familiarity with cloud platforms, containerization and CI/CD pipelines. Experience defining and managing SLOs, SLIs and error budgets

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Loading talks and stories from around this role…