> Markdown version of [/jobs/ext/731727-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/731727-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Site Reliability Engineer - **Company:** TEKCHRONICLES, INC. - **Location:** Jersey City, NJ, United States - **Experience:** Expert - **Salary:** $145,600.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Applications Architecture, Build Automation, Software Quality, Code Review, Databases, Software Design Patterns, Disaster Recovery, Distributed Systems, Middleware, Failover, Python (Programming Language), OpenShift, Reliability Engineering, Site Reliability Engineering Practices, Ansible, Prometheus, Runbook, Datadog, Grafana, Kubernetes, Terraform, Splunk, New Relic (SaaS), Dynatrace - **Published:** June 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=c6686ebbcc37787e ## About the Role Do you have experience in Terraform?, We are looking for a highly experienced Lead Site Reliability Engineer to drive SRE outcomes for business-critical applications in the Risk Technology space. This role requires a strong application, infrastructure, and engineering mindset, with the ability to work closely with application support, development, observability, and technology teams to improve reliability, resiliency, operational readiness, and automation maturity., * Minimum 10 years of experience in SRE, production engineering, application reliability, infrastructure engineering, or related technology roles. * Strong understanding of SRE principles, including SLIs, SLOs, SLAs, error budgets, toil reduction, incident management, and reliability engineering. * Deep experience supporting business-critical applications in production environments. * Strong application architecture knowledge with the ability to understand design, workflows, dependencies, and failure scenarios. * Hands-on experience with Python automation. * Hands-on experience with Ansible and Terraform automation. * Strong knowledge of observability practices, including metrics, logs, traces, dashboards, alerting, and service health monitoring. * Ability to partner with application support and development teams to improve reliability from both operational and engineering perspectives. * Strong understanding of cloud, infrastructure, networking, databases, middleware, and application runtime environments. * Experience reviewing code, supporting code quality discussions, and identifying reliability risks in application changes. * Strong problem-solving skills with the ability to deep dive into complex technical issues. * Excellent communication skills with the ability to translate technical risks into business-impacting outcomes., * Experience in financial services, banking, risk technology, regulatory platforms, or other high-criticality environments. * Exposure to AWS/Amazon services and AI-enabled automation or operational intelligence capabilities. * Experience with Prometheus, Grafana, OpenTelemetry, Splunk, Datadog, Dynatrace, New Relic, or similar observability platforms. * Knowledge of Kubernetes, OpenShift, containers, CI/CD pipelines, and modern distributed systems. * Experience building reliability scorecards, operational readiness reviews, service maturity assessments, and production support standards. * Strong understanding of resiliency patterns, failover, disaster recovery, capacity planning, and performance engineering., The ideal candidate is a hands-on SRE leader who can think like an engineer, operate like a production owner, and partner like a trusted advisor to application teams. They should be comfortable going deep into application behavior, understanding business workflows, challenging reliability gaps, and enabling practical SRE outcomes that improve stability, resiliency, and operational excellence. ## Description The ideal candidate will be responsible for defining and enabling SRE goals, establishing reliability requirements, evaluating SLAs, SLOs, SLIs, and error budgets, identifying critical user journeys, and ensuring that business-critical applications are supported with the right monitoring, alerting, automation, and operational practices., * Lead SRE enablement for business-critical applications across Risk Technology. * Partner closely with application support, development, infrastructure, and observability teams to improve reliability and resiliency. * Define SRE priorities, goals, standards, and measurable outcomes for application teams. * Establish and evaluate SLAs, SLOs, SLIs, error budgets, and service health indicators. * Identify and document critical user journeys, application dependencies, failure points, and recovery expectations. * Drive observability improvements by ensuring the right monitors, alerts, dashboards, logs, traces, and metrics are in place. * Review application architecture, workflows, design patterns, and production support processes to identify reliability gaps. * Support code-level analysis, code review discussions, and engineering recommendations from an SRE perspective. * Improve incident management, post-incident reviews, root cause analysis, runbooks, and operational readiness. * Build automation using Python, Ansible, and Terraform to reduce manual effort and improve operational efficiency. * Leverage Amazon/AWS products, including AI-based solutions, to improve SRE efficiency, automation, monitoring, and operational outcomes. * Help application teams adopt industry-standard SRE practices inspired by mature engineering organizations. * Work hands-on with teams to improve production stability, resiliency, scalability, and supportability. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers)