Site Reliability Engineer

LexisNexis Risk Solutions Group
Boca Raton, FL, United States
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$104,900.0 - $174,700.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Confluence JIRA Microsoft Azure Linux DevOps Github Reliability Engineering Grafana Git Flow Kubernetes Terraform
+2 more
Docker Servicenow

Job description

We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory role you will be directly involved in designing infrastructure, writing Terraform, improving observability, and responding to real production incidents., * Design, build, and operate highly available, scalable systems in AWS

  • Write, maintain, and review Terraform to provision and manage infrastructure
  • Own and improve monitoring, alerting, and observability using Grafana, Pingdom, and Uptrends
  • Participate in a rotating on-call schedule, responding to production incidents and driving issues to resolution
  • Lead incident response, root cause analysis, and post-incident reviews with a focus on prevention and automation
  • Define and manage SLOs, SLIs, and error budgets
  • Build and improve CI/CD pipelines and operational workflows using Azure DevOps and GitHub
  • Work directly with application teams to improve reliability, performance, and deployability
  • Automate manual operational tasks to reduce toil
  • Maintain clear, actionable runbooks and documentation in Confluence
  • Track work, incidents, and operational improvements using Jira and ServiceNow
  • Mentor other engineers and help set SRE standards and best practices

Requirements

  • 5+ years of hands-on experience in SRE, DevOps, or Infrastructure Engineering roles
  • Strong production experience in AWS
  • Required: Significant hands-on experience with Terraform in real-world environments
  • Experience operating monitoring and uptime platforms such as Grafana, Pingdom, and Uptrends
  • Strong Linux systems, networking, and troubleshooting skills
  • Experience supporting production systems through incident response and on-call rotations
  • Proficiency with GitHub and modern Git workflows
  • Experience building or maintaining CI/CD pipelines with Azure DevOps
  • Familiarity with ITSM and incident workflows using ServiceNow
  • Strong written communication skills with experience documenting systems and processes in Confluence
  • Ability to work independently in a remote or hybrid environment

Preferred Qualifications

  • Experience defining and operating against SLOs and error budgets
  • Infrastructure-as-Code best practices beyond Terraform (modules, testing, CI integration)
  • Experience with containers and orchestration (Docker, Kubernetes)
  • Experience supporting large-scale, high-availability production systems
  • Prior experience mentoring engineers or serving as a technical lead

About the company

LexisNexis Risk Solutions is the essential partner in the assessment of risk. Within our Business Services vertical, we offer a multitude of solutions focused on helping businesses of all sizes drive higher revenue growth, maximize operational efficiencies, and improve customer experience. Our solutions help our customers solve difficult problems in the areas of Anti-Money Laundering/Counter Terrorist Financing, Identity Authentication & Verification, Fraud and Credit Risk mitigation and Customer Data Management. You can learn more about LexisNexis Risk at the link below, https://risk.lexisnexis.com

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

5:47 min

Integrating user stories and test automation via Jira tools

Christoph Ruggenthaler · LIVE

Videos

See all

Related articles

See all