> Markdown version of [/jobs/ext/2167011-site-reliability-engineer-ii](https://www.wearedevelopers.com/jobs/ext/2167011-site-reliability-engineer-ii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer II - **Company:** Bank of America - **Location:** Jersey City, NJ, United States - **Experience:** Expert - **Salary:** $108,000.0 - $161,900.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Oracle WebLogic Server, Unix, Cloud Computing, Cyber Security, Computer Programming, Databases, Continuous Delivery, Continuous Integration, Database Theory, Linux, DevOps, Domain Name System (DNS), Java Platform Enterprise Edition (J2EE), Monitoring of Systems, HP SiteScope, WildFly (JBoss AS), Python (Programming Language), Network Troubleshooting, Unix Shell, Network Administration, Oracle Databases, Oracle (Applications), Reliability Engineering, Standard Sql, Runbook, Software Engineering, PL-SQL, SQL Databases, Time Tracking Software, Web Services, Network Routing, Scripting, Load Balancing, Cloud Platform System, Spring-boot, Reliability of Systems, Splunk, Dynatrace - **Published:** August 21, 2026 - **Apply:** https://www.careerbuilder.com/job-details/site-reliability-engineer-ii-jersey-city-nj--d550b52d-fc04-4326-8784-008ec3cb4710 ## About the Role * Minimum 7+ years of experience in application production support role * Production support experience supporting Java/J2EE applications in an enterprise environment, including WebLogic, web services, Spring Boot, and strong SQL/PL/SQL skills for troubleshooting * Strong working knowledge of Linux/Unix environments and scripting languages such as Shell/Python, including applications deployed on JBoss * Experience using monitoring and observability tools (e.g., Splunk, Dynatrace, Nastel, SiteScope) in a production support environment * Hands on experience troubleshooting network related production incidents, including load balancing, traffic routing, and DNS issues, across on prem and cloud environments, with a focus on rapid service restoration * Experience and understanding database concepts (SQL / Oracle) and writing basic queries * Banking or Capital Markets domain knowledge preferred * Ability to manage multiple tasks simultaneously and adapt quickly to changing priorities and production demands * Self-starter with the ability to work independently as well as collaboratively within cross-functional teams * Excellent analytical skills with the ability to identify root causes of complex production issues * Familiarity with SRE principles including SLO/SLI definition, error budgets, and toil reduction strategies * Experience identifying and automating repetitive operational tasks to reduce toil and improve team efficiency * Understanding of CI/CD pipelines and ability to support post-deployment validation in an automated release environment * Experience defining and owning application availability targets, contributing to reliability improvement plans, and driving proactive measures to prevent recurrence of production degradation * Provide on-call rotational support, including off-hours support, during weeknights, Saturdays, and Sundays, * Proven experience as a proactive problem solver in a production support environment * Strong verbal and written communication skills, with the ability to convey technical concepts to both technical and non technical audiences * Ability to review system logs and monitor data to identify subtle performance anomalies * Strong prioritization skills with the ability to manage multiple incidents under tight deadlines * Collaborative mindset with experience working across cross functional teams Skills: * Analytical Thinking * Automation * Collaboration * Production Support * Result Orientation * Application Development * Architecture * Influence * Project Management * Solution Design * Adaptability * DevOps Practices * Risk Management * Solution Delivery Process * Stakeholder Management, Analysis Skills, Automation, Banking Services, Budgeting, Capital Markets, Career Development, Cloud Computing, Communication Skills, Computer Security, Continuous Deployment/Delivery, Continuous Integration, Corrective Action, Cross-Functional, Customer Support/Service, DNS (Domain Name System), Database Technology, DevOps, Documentation, Establish Priorities, Help Desk, Identify Issues, Incident Management, Incident Response, Instrumentation, Investment Services, JBoss Application Server, Java, Java Platform Enterprise Edition (Java EE/J2EE), Leadership, Linux Operating System, Load Balancing, Machine Tool, Maintain Compliance, Mentoring, Military, Multitasking, Network Administration/Management, Network Routing, On Call, Oracle Database, Oracle PL-SQL, Oracle WebLogic Server, Performance Analysis, Performance Metrics, Presentation/Verbal Skills, Problem Solving Skills, Process Improvement, Production Control, Production Support, Production Systems, Project/Program Management, Python Programming/Scripting Language, Reliability Engineering, Reporting Dashboards, Risk Management, Root Cause Analysis, SQL (Structured Query Language), Scripting (Scripting Languages), Service Level Agreement (SLA), Software Development, Splunk, Systems Reliability, Talent Management, Team Player, Technical Operations, Time Management, Time Tracking, Unix Operating Systems, Unix Shell Programming, Web Services, Writing Skills ## Description * Develops and maintains reliability scripts, tools and libraries and leverages them for common instrumentation, automation, and operational needs, and when mentoring Site Reliability Engineer (SRE) resources on reliability practices and established tools/capabilities * Collaborates with Development and Infrastructure teams to understand technical solutions and implement monitoring capabilities outlined in the application and system monitoring designs put forward by the SRE Lead * Partners to implement code changes to make use of common reliability libraries and tools and helps Application Production Services and Application Development teammates understand how to use them * Identifies vulnerabilities and opportunities for reliability improvement, such as investigating low level error rates and 'noise' in monitoring, and defines solutions to reduce manual support effort and/or improve system reliability * Engages as a subject matter expert in major incident triage efforts and failure scenario modelling and diagnosis with Problem Manager root causes for major incident/problem management investigations * Participates regularly in an on-call rotation with Production Support teammates to learn more about reliability issues affecting their portfolio * Provide front line production support and monitoring to ensure application stability and availability * Triage, troubleshoot, and resolve production incidents, including business impacting issues * Lead incident response and bridge calls, coordinating troubleshooting and escalation as needed * Perform root cause analysis and drive remediation and preventative actions * Monitor system alerts, logs, dashboards, and performance metrics to assess impact and restore service * Support batch jobs and data feeds, including time sensitive failures * Conduct post release validation and routine application health checks * Resolve user requests related to access, technical issues, and data discrepancies * Maintain accurate incident documentation, runbooks, and knowledge articles * Partner with technology, operations, vendors, and business teams to improve reliability * Participate in a rotational weekend and after hours support schedule as required, This role is eligible to participate in the annual discretionary plan. Employees are eligible for an annual discretionary award based on their overall individual performance results and behaviors, the performance and contributions of their line of business and/or group; and the overall success of the Company. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [WeAreDevelopers LIVE - Node and Package Security](https://www.wearedevelopers.com/videos/2138-wearedevelopers-live-node-and-package-security) - [Enterprise-Cloud-Native - Fast-Paced Development & Deployment in a Highly Secure Banking Environment](https://www.wearedevelopers.com/videos/671-enterprise-cloud-native-fast-paced-development-deployment-in-a-highly-secure-banking-environment) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [93 Java Interview Questions You Should Prepare For](https://www.wearedevelopers.com/magazine/14-93-java-interview-questions-you-should-prepare-for) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology) - [Software Developer Salary in Switzerland [2023]](https://www.wearedevelopers.com/magazine/215-software-developer-salary-in-switzerland-2023) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)