> Markdown version of [/jobs/ext/116883-site-reliability-engineer-sre-ii](https://www.wearedevelopers.com/jobs/ext/116883-site-reliability-engineer-sre-ii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (SRE) - II - **Company:** Huntington Bancshares - **Location:** Columbus, OH, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Engineering, Continuous Integration, Information Engineering, DevOps, Distributed Systems, Domain Name System (DNS), Fault Tolerance, Monitoring of Systems, Hypertext Transfer Protocols (HTTP), Python (Programming Language), Networking Basics, Performance Tuning, Reliability Engineering, Ansible, Prometheus, Software Deployment, Software Engineering, TCP/IP, Software Vulnerability Management, Datadog, Circleci, Scripting, Google Cloud, Load Balancing, Grafana, Reliability of Systems, Infrastructure as Code (IaC), Cloudformation, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Machine Learning Operations, Functional Programming, Puppet, Terraform, Dynatrace, Docker, Elk Stack, Jenkins, Golang, Microservices - **Published:** May 16, 2026 - **Apply:** https://huntington-careers.com/search/jobdetails/site-reliability-engineer-sre--ii/537e212a-dbc6-4681-97e6-7ef7042b8036 ## About the Role * Minimum 5 years of experience in site reliability engineering, DevOps, systems administration, or related roles. * Strong experience with Linux/Unix administration and proficiency in scripting (e.g., Python, Bash, Go). * Deep understanding of cloud platforms (AWS, GCP, Azure) and related services (EC2, S3, Lambda, Kubernetes, etc.). + Experience with containerization and orchestration technologies like Docker and Kubernetes. + Proficiency with monitoring and observability tools such as dynatrace, Prometheus, Grafana, Datadog, ELK Stack, or similar platforms. + Strong understanding of networking fundamentals (DNS, HTTP, TCP/IP), load balancing, and CDNs. + Experience with CI/CD tools (Jenkins, GitLab CI, CircleCI) and infrastructure automation (Terraform, Ansible, Puppet). + Familiarity with distributed systems and microservices architecture. + Excellent problem-solving and troubleshooting skills, especially in diagnosing production issues in high-scale environments. Preferred: * Background in MLOps, data engineering, and/or cloud-native AI deployment. * Strong communication and documentation abilities * Knowledge of security best practices for AI and cloud infrastructure. * Contributions to open source AI/SRE projects or relevant technical communities * Proven track record of managing complex infrastructure, troubleshooting production issues, and optimizing system performance ## Description As a Site Reliability Engineer (SRE) Level II, you will play a key role in maintaining the availability, scalability, and performance of critical infrastructure and services. You will be responsible for building and automating solutions that enhance system reliability and support continuous delivery. In this role, you will handle more complex operational tasks and incidents, provide mentorship to junior SREs, and collaborate with development teams to ensure systems are designed for reliability from the ground up. * Incident Management : * complex incidents, and ensure service uptime. * Lead troubleshooting efforts for high-impact production issues, providing detailed root cause analysis (RCA) and preventative measures. * Participate in on-call rotations, acting as an escalation point for Level 1 SREs during major incidents. * Automation & Infrastructure as Code (IaC): * Develop and maintain automation scripts and infrastructure using tools like Terraform, Ansible, or CloudFormation. * Implement automation solutions to eliminate manual tasks and improve system reliability, scalability, and performance. * Performance & Scalability: * Analyze system performance and recommend optimizations for scalability and reliability. * Support capacity planning efforts by monitoring system metrics, traffic * patterns, and usage trends to predict future resource needs. * System Design & Architecture: * Collaborate with software engineering teams to influence the design of new services and applications, ensuring they are scalable, reliable, and resilient from the start. * Contribute to architectural decisions, ensuring alignment with best practices in fault tolerance, redundancy, and recovery. * Monitoring & Observability: * Build and maintain robust monitoring, alerting, and observability solutions to proactively detect and resolve issues before they impact end users. * Optimize existing monitoring tools (e.g., Prometheus, Grafana, Datadog, Dynatrace) and build custom dashboards for better visibility into system health. * Security & Compliance: * Ensure systems and infrastructure are secure, compliant, and aligned with organizational policies and industry best practices. * Assist with vulnerability management, system patching, and implementing security measures to protect the integrity and availability of services. * Continuous Improvement: * Lead efforts to continuously improve operational processes, tools, and workflows. * Implement and enforce best practices in deployment, monitoring, and incident management to improve overall system reliability and reduce downtime., Certain positions outside our branch network may be eligible for a flexible work arrangement. We're combining the best of both worlds: in-office and work from home. Our approach enables our teams to deepen connections, maintain a strong community, and do their best work. Remote roles will also have the opportunity to come together in our offices for moments that matter. Specific work arrangements will be provided by the hiring team. ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025)