Site Reliability Engineer - Big Data Platform

The Hartford
Hartford, CT, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$136,000.0 - $204,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Business Analytics Applications Data Analysis Bash Shell Big Data Software as a Service Cloud Computing Configuration Management Computer Programming Databases
+47 more
Computer Engineering Continuous Integration Data Cleansing Information Engineering Data Infrastructure Extract Transform Load (ETL) Data Mining Data Normalization Data Systems Data Visualization Linux Digital Assets Apache Hadoop Identity and Access Management Job Scheduling Python (Programming Language) Kerberos (Protocol) Machine Learning Meta-Data Management Cisco Nexus Switches Platform as a Service (PAAS) Performance Tuning Reliability Engineering Ansible Cloudera SQL Databases Data Processing Scripting Data Ingestion Delivery Pipeline Apache Spark Reliability of Systems Infrastructure as Code (IaC) Cloudformation Amazon Relational Database Service Containerization Information Technology Data Analytics Data Management Route53 Cloudwatch Terraform Splunk Dynatrace Serverless Computing Amazon Elastic Mapreduce (EMR) Jenkins

Job description

The Hartford is seeking a seasoned SRE with Big Data experience to join our Cloud Big Data Platform Engineering & Operations team. This role is instrumental in fostering a customer-first mindset and ensuring the stability, scalability, and reliability of our data platforms to support the evolving needs of Data & Analytics applications across the enterprise.

As a technical lead, you will apply your deep expertise in AWS Big Data/EMR infrastructure, Infrastructure as Code (IaC), security, automation and observability to engineer, optimize, and maintain robust, scalable solutions. You will collaborate closely with data engineers to analyze requirements & challenges to recommend platform-driven solutions that maximize performance and efficiency.

We’re looking for a passionate technologist who thrives in a dynamic, fast-paced environment and is committed to building resilient, future-ready data platforms. You will mentor and guide fellow platform and reliability engineers, promoting a culture of technical excellence, innovation, and collaboration.

RESPONSIBILITIES

* Administer and engineer Big Data platforms across multiple Hadoop clusters in the cloud (AWS EMR), including serverless and containerized environments, to ensure scalability, performance, and reliability.

  • Design, implement and maintain highly scalable and resilient multi-tenant Data Platforms through Infrastructure as Code (IAC) aligning with The Hartford’s engineering, security and governance principles

  • Ensure operational excellence, independently drive the triaging and service restoration of all high impact incidents to minimize the mean time to service restoration and impact to the business. Demonstrate end to end ownership

  • Apply Site Reliability Engineering (SRE) principles to design and implement robust tooling, proactive alerting, and automated response mechanisms that identify, mitigate, and resolve reliability risks-focusing on prevention, early detection, and continuous improvement through automation.

  • Own and evolve the architecture of Platform, PaaS, and SaaS solutions to meet current and future business need-driving innovation, scalability, and operational excellence.

  • Participate in an on-call rotation, providing hands-on technical expertise during service-impacting incidents to ensure rapid diagnosis, effective resolution, and continuous improvement of system reliability.

  • Evaluate, implement, and manage emerging data technologies with a focus on big data, analytics, data wrangling, business intelligence, and data visualization to drive innovation and efficiency.

  • Serve as a subject matter expert and technical lead for data platforms, tools, and application interfaces - driving root cause analysis, resolving complex technical issues, and ensuring platform reliability and performance.

  • Provide technical leadership and mentorship to junior and mid-level data engineers, fostering skill development and promoting engineering best practices.

  • Collaborate with and empower data engineers, data scientists, and business analysts by enabling self-service capabilities for data wrangling, exploration, and analysis

  • Develop training materials and deliver end-user training sessions to drive adoption, ensure effective use of data solutions, and enhance customer engagement.

  • Support the documentation, metadata management, and visualization of data assets to promote data discoverability, transparency, and self-service analytics.

  • Foster a culture of accountability and collaboration by building strong team commitment to shared priorities and strategic goals.

Requirements

  • Bachelor’s degree in computer science, Computer Engineering, or a closely related field (or foreign equivalent). Relevant experience may be obtained through a qualifying post-baccalaureate academic program.

  • Extensive expertise in platform administration, big data technologies, data engineering, analytics, and operations.

  • Proficiency in a broad range of tools and concepts including Hadoop, Linux, Python, SQL, Spark, Kerberos, cloud platforms, security protocols, performance tuning, machine learning algorithms, production engineering, job scheduling, and operational support

  • At least 7 years of progressive and diverse experience in IT, platform administration, database management, analytics or an equivalent combination of education and work experience

  • At least 3 years of hands-on engineering experience in developing platform solutions on AWS using Infrastructure as Code (CloudFormation, Terraform, Ansible) and CICD pipelines (Jenkins, Nexus, AWS CodeBuild, CodeDeploy or CodePipeline).

  • At least 3 years of experience providing architectural guidance and technical direction to developers, or platform administrators

  • Knowledge in data engineering, with exposure to designing and building data applications using database platforms.

  • Responsibilities should include data ingestion, data preparation, ETL processes, data aggregation, data mining, and the development of database and analytics applications

  • Experience in managing change through Change Management and Incident Management processes

  • Experience in implementing Reliability Engineering practices and Observability dashboards leveraging Splunk, Dynatrace, CloudWatch etc. will be a plus

  • Ability to learn new technologies quickly, and perform major job responsibilities proficiently within 6-12 months

  • Strong analytical ability, problem analysis techniques, and broad knowledge of alternatives technology

  • Strong communication skills, and the ability to work effectively with business and IT resources

PREFERRED SKILLS

  • Advanced experience in primary AWS services (EMR, EKS, EC2, IAM, RDS, Route53 & S3, etc)

  • Experience in Configuration management using CloudFormation / Terraform.

  • Advanced experience with programming and/or scripting languages (Python, bash)

  • Preferred to have “AWS Solution Architect Certification”.

  • Preferred to have “Cloudera Admin Certification”.

This role will have a Hybrid work schedule, with the expectation of working in an office (Columbus, OH, Chicago, IL, Hartford, CT or Charlotte, NC) 3 days a week (Tuesday through Thursday).

Candidates must be authorized to work in the US without company sponsorship. The company will not support the STEM OPT I-983 Training Plan endorsement for this position.

Benefits & conditions

The listed annualized base pay range is primarily based on analysis of similar positions in the external market. Actual base pay could vary and may be above or below the listed range based on factors including but not limited to performance, proficiency and demonstration of competencies required for the role. The base pay is just one component of The Hartford’s total compensation package for employees. Other rewards may include short-term or annual bonuses, long-term incentives, and on-the-spot recognition. The annualized base pay range for this role is: $136,000 - $204,000

About the company

We’re determined to make a difference and are proud to be an insurance company that goes well beyond coverages and policies. Working here means having every opportunity to achieve your goals - and to help others accomplish theirs, too. Join our team as we help shape the future.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt ¡ LIVE

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto ¡ WWC 2024

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 ¡ LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard ¡ WWC 2025

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger ¡ WWC 2025

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 ¡ LIVE

Videos

See all

Related articles

See all