Data Infrastructure Site Reliability Engineer
Role details
Job location
Tech stack
Job description
- Maintain and support highly available, scalable, and secure data infrastructure platforms across AWS and on-premises environments.
- Drive operational excellence through automation of repetitive tasks, incident reduction, and proactive reliability improvements.
- Monitor platform health, troubleshoot complex issues, and lead root cause analysis efforts to minimize downtime and improve system resiliency.
- Collaborate closely with engineering, platform, and data teams to support deployments, integrations, upgrades, and performance optimization initiatives.
- Participate in on-call rotations and partner with teams across the US and India to provide 24x7 operational support. US support is aligned to Pacific Time zone.
- Continuously improve observability, alerting, and incident response processes to enhance platform reliability and reduce operational noise.
Requirements
Skills - Site Reliability Engineering, Big Data, AWS, AWS EMR, AWS EKS, AWS MSK, AWS Athena, AWS Glue, Spark, Iceberg, Python, Terraform, Cloudera Hadoop, Agentic Ai, Monitoring Tools, We are seeking a highly skilled L2+/Senior Data Infrastructure Site Reliability Engineer (SRE) to operate, optimize, and automate large-scale data infrastructure environments across AWS and on-premises Hadoop platforms. This role requires a strong reliability engineering mindset focused on platform stability, performance, observability, automation, and incident response.
As a key member of the Data Infrastructure SRE team, you will ensure the availability, scalability, security, and operational excellence of mission-critical data platforms while driving continuous improvements through Infrastructure as Code (IaC), AI-enabled automation, and modern SRE practices., Site Reliability Engineering (SRE)
-
Strong SRE mindset with a proven focus on reliability, availability, performance optimization, incident management, and operational excellence.
-
Experience delivering services within defined SLAs and ensuring timely resolution of production issues.
-
Expertise in troubleshooting complex distributed systems and identifying root causes quickly and effectively.
-
AWS & Cloud Infrastructure
-
Deep hands-on experience with AWS services, including:
-
EMR
-
EKS
-
MSK
-
Athena
-
Glue
-
IAM
-
Amazon S3
-
VPC
-
AWS networking and security services
-
Strong understanding of cloud-native architectures, scalability, and infrastructure resilience.
Big Data Platforms
- Extensive operational experience managing Hadoop clusters, with a strong focus on administration, platform maintenance, and automation of day-to-day operational activities.
- Experience supporting both AWS-based data platforms and on-premises Cloudera CDH/CDP environments.
- Solid understanding of Kerberos authentication and security implementation within Hadoop ecosystems.
Hands-on experience with:
- Apache Spark
- Apache Iceberg
- Big Data platform architecture
- Performance tuning and optimization
Linux & System Administration
- Strong Linux administration and operational support experience.
- Expertise in user and access management, system configuration, customization, and platform administration.
Observability & Incident Management
- Hands-on experience with monitoring and observability platforms such as:
- AWS CloudWatch
- Datadog
- PagerDuty
- Similar enterprise monitoring solutions
- Proven success improving alert quality, reducing false positives, and minimizing alert fatigue.
- Excellent debugging and troubleshooting skills across infrastructure, applications, Spark workloads, and Iceberg environments.
Java Platform Operations
- Strong understanding of Java application administration, including:
- JVM tuning
- Thread dump analysis
- Heap dump analysis
- JVM parameters
- Application log analysis and troubleshooting
Automation & Infrastructure as Code
- Expert-level Python scripting for operational automation, monitoring, and reliability engineering use cases.
Strong experience with:
- Terraform
- Infrastructure as Code (IaC)
- CI/CD pipeline implementation and automation
- Platform engineering best practices, * Current hands-on experience applying AI technologies to Data Infrastructure and SRE operations.
- Demonstrated ability to design and implement:
- Agentic AI solutions
- AI-assisted operational workflows
- Intelligent automation for routine SRE activities
- Strong creativity and problem-solving skills in leveraging AI to improve operational efficiency, reliability, and productivity.