Lead Site Reliability Engineer

TEKCHRONICLES, INC.
Jersey City, NJ, United States
about 2 months ago
Apply on indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$145,600.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Applications Architecture Build Automation Software Quality Code Review Databases Software Design Patterns Disaster Recovery Distributed Systems Middleware Failover
+14 more
Python (Programming Language) OpenShift Reliability Engineering Site Reliability Engineering Practices Ansible Prometheus Runbook Datadog Grafana Kubernetes Terraform Splunk New Relic (SaaS) Dynatrace

Job description

The ideal candidate will be responsible for defining and enabling SRE goals, establishing reliability requirements, evaluating SLAs, SLOs, SLIs, and error budgets, identifying critical user journeys, and ensuring that business-critical applications are supported with the right monitoring, alerting, automation, and operational practices., * Lead SRE enablement for business-critical applications across Risk Technology.

  • Partner closely with application support, development, infrastructure, and observability teams to improve reliability and resiliency.
  • Define SRE priorities, goals, standards, and measurable outcomes for application teams.
  • Establish and evaluate SLAs, SLOs, SLIs, error budgets, and service health indicators.
  • Identify and document critical user journeys, application dependencies, failure points, and recovery expectations.
  • Drive observability improvements by ensuring the right monitors, alerts, dashboards, logs, traces, and metrics are in place.
  • Review application architecture, workflows, design patterns, and production support processes to identify reliability gaps.
  • Support code-level analysis, code review discussions, and engineering recommendations from an SRE perspective.
  • Improve incident management, post-incident reviews, root cause analysis, runbooks, and operational readiness.
  • Build automation using Python, Ansible, and Terraform to reduce manual effort and improve operational efficiency.
  • Leverage Amazon/AWS products, including AI-based solutions, to improve SRE efficiency, automation, monitoring, and operational outcomes.
  • Help application teams adopt industry-standard SRE practices inspired by mature engineering organizations.
  • Work hands-on with teams to improve production stability, resiliency, scalability, and supportability.

Requirements

Do you have experience in Terraform?, We are looking for a highly experienced Lead Site Reliability Engineer to drive SRE outcomes for business-critical applications in the Risk Technology space. This role requires a strong application, infrastructure, and engineering mindset, with the ability to work closely with application support, development, observability, and technology teams to improve reliability, resiliency, operational readiness, and automation maturity., * Minimum 10 years of experience in SRE, production engineering, application reliability, infrastructure engineering, or related technology roles.

  • Strong understanding of SRE principles, including SLIs, SLOs, SLAs, error budgets, toil reduction, incident management, and reliability engineering.
  • Deep experience supporting business-critical applications in production environments.
  • Strong application architecture knowledge with the ability to understand design, workflows, dependencies, and failure scenarios.
  • Hands-on experience with Python automation.
  • Hands-on experience with Ansible and Terraform automation.
  • Strong knowledge of observability practices, including metrics, logs, traces, dashboards, alerting, and service health monitoring.
  • Ability to partner with application support and development teams to improve reliability from both operational and engineering perspectives.
  • Strong understanding of cloud, infrastructure, networking, databases, middleware, and application runtime environments.
  • Experience reviewing code, supporting code quality discussions, and identifying reliability risks in application changes.
  • Strong problem-solving skills with the ability to deep dive into complex technical issues.
  • Excellent communication skills with the ability to translate technical risks into business-impacting outcomes., * Experience in financial services, banking, risk technology, regulatory platforms, or other high-criticality environments.
  • Exposure to AWS/Amazon services and AI-enabled automation or operational intelligence capabilities.
  • Experience with Prometheus, Grafana, OpenTelemetry, Splunk, Datadog, Dynatrace, New Relic, or similar observability platforms.
  • Knowledge of Kubernetes, OpenShift, containers, CI/CD pipelines, and modern distributed systems.
  • Experience building reliability scorecards, operational readiness reviews, service maturity assessments, and production support standards.
  • Strong understanding of resiliency patterns, failover, disaster recovery, capacity planning, and performance engineering., The ideal candidate is a hands-on SRE leader who can think like an engineer, operate like a production owner, and partner like a trusted advisor to application teams. They should be comfortable going deep into application behavior, understanding business workflows, challenging reliability gaps, and enabling practical SRE outcomes that improve stability, resiliency, and operational excellence.

Benefits & conditions

$70 an hour - Contract

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · World Congress 2025

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all