> Markdown version of [/jobs/ext/1410932-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1410932-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Software Technology, Inc. - **Location:** Princeton, United States (Remote available) - **Experience:** Expert - **Contract:** Temporary to permanent - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Apache HTTP Server, Big Data, Cloud Computing, Cloud Engineering, System Configuration, Continuous Integration, Data Infrastructure, Data Security, Software Debugging, Linux, Distributed Systems, Apache Hadoop, Monitoring of Systems, Identity and Access Management, Java Virtual Machine (JVM), Python (Programming Language), Kerberos (Protocol), Linux System Administration, Log Analysis, Performance Tuning, Reliability Engineering, Site Reliability Engineering Practices, Cloudera, Systems Integration, Scripting, Apache Spark, Infrastructure as Code (IaC), AWS Glue, Data Management, Virtual Agents, Cloudwatch, Terraform, Heap (Data Structure), AWS EKS, Amazon Elastic Mapreduce (EMR) - **Published:** July 23, 2026 - **Apply:** https://www.dice.com/job-detail/577e61cb-3a8d-488a-baef-8d0d19c32370 ## About the Role Skills - Site Reliability Engineering, Big Data, AWS, AWS EMR, AWS EKS, AWS MSK, AWS Athena, AWS Glue, Spark, Iceberg, Python, Terraform, Cloudera Hadoop, Agentic Ai, Monitoring Tools, We are seeking a highly skilled L2+/Senior Data Infrastructure Site Reliability Engineer (SRE) to operate, optimize, and automate large-scale data infrastructure environments across AWS and on-premises Hadoop platforms. This role requires a strong reliability engineering mindset focused on platform stability, performance, observability, automation, and incident response. As a key member of the Data Infrastructure SRE team, you will ensure the availability, scalability, security, and operational excellence of mission-critical data platforms while driving continuous improvements through Infrastructure as Code (IaC), AI-enabled automation, and modern SRE practices., Site Reliability Engineering (SRE) Strong SRE mindset with a proven focus on reliability, availability, performance optimization, incident management, and operational excellence. Experience delivering services within defined SLAs and ensuring timely resolution of production issues. Expertise in troubleshooting complex distributed systems and identifying root causes quickly and effectively. AWS & Cloud Infrastructure Deep hands-on experience with AWS services, including: EMR, AWS networking and security services Strong understanding of cloud-native architectures, scalability, and infrastructure resilience. Big Data Platforms * Extensive operational experience managing Hadoop clusters, with a strong focus on administration, platform maintenance, and automation of day-to-day operational activities. * Experience supporting both AWS-based data platforms and on-premises Cloudera CDH/CDP environments. * Solid understanding of Kerberos authentication and security implementation within Hadoop ecosystems. Hands-on experience with: * Apache Spark * Apache Iceberg * Big Data platform architecture * Performance tuning and optimization Linux & System Administration Strong Linux administration and operational support experience. Expertise in user and access management, system configuration, customization, and platform administration. Observability & Incident Management Hands-on experience with monitoring and observability platforms such as: AWS CloudWatch, Proven success improving alert quality, reducing false positives, and minimizing alert fatigue. Excellent debugging and troubleshooting skills across infrastructure, applications, Spark workloads, and Iceberg environments. Java Platform Operations Strong understanding of Java application administration, including: JVM tuning Thread dump analysis Heap dump analysis JVM parameters Application log analysis and troubleshooting Automation & Infrastructure as Code Expert-level Python scripting for operational automation, monitoring, and reliability engineering use cases. Strong experience with: Terraform Infrastructure as Code (IaC) CI/CD pipeline implementation and automation Platform engineering best practices AI-Driven Operations Current hands-on experience applying AI technologies to Data Infrastructure and SRE operations. Demonstrated ability to design and implement: Agentic AI solutions AI-assisted operational workflows Intelligent automation for routine SRE activities Strong creativity and problem-solving skills in leveraging AI to improve operational efficiency, reliability, and productivity. Preferred Candidate Profile The ideal candidate combines deep expertise in AWS cloud platforms, Hadoop ecosystems, SRE practices, observability, automation, and AI-driven operations, with a passion for supporting resilient data platforms at scale and eliminating operational toil through engineering excellence. ## Description * Maintain and support highly available, scalable, and secure data infrastructure platforms across AWS and on-premises environments. * Drive operational excellence through automation of repetitive tasks, incident reduction, and proactive reliability improvements. * Monitor platform health, troubleshoot complex issues, and lead root cause analysis efforts to minimize downtime and improve system resiliency. * Collaborate closely with engineering, platform, and data teams to support deployments, integrations, upgrades, and performance optimization initiatives. * Participate in on-call rotations and partner with teams across the US and India to provide 24x7 operational support. US support is aligned to Pacific Time zone. * Continuously improve observability, alerting, and incident response processes to enhance platform reliability and reduce operational noise. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [How we built an AI-powered code reviewer in 80 hours](https://www.wearedevelopers.com/videos/1511-how-we-built-an-ai-powered-code-reviewer-in-80-hours) - [Hosting a modern justice system](https://www.wearedevelopers.com/videos/332-hosting-a-modern-justice-system) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Tips, Techniques, and Common Pitfalls Debugging Kafka](https://www.wearedevelopers.com/videos/838-tips-techniques-and-common-pitfalls-debugging-kafka) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)