Wealth Management Production Management Site Reliability Engineer

PamTen
Atlanta, GA, United States
3 days ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Starter
Working hours
Regular working hours

Tech stack

Java (Programming Language) Agile Methodology CA Workload Automation Ae Unix Cloud Computing Information Systems Computer Programming Continuous Delivery Continuous Integration Cron IBM DB2 Software Debugging
+21 more
Linux DevOps Disaster Recovery Spring Framework Unix Shell Pattern Recognition Reliability Engineering Prometheus Server Administration Shell Script Software Engineering SQL Databases ReactJS Grafana Software Troubleshooting Build Management AngularJS Information Technology Deployment Automation Kibana Splunk

Job description

Position Overview: The Wealth Management Production Management Site Reliability Engineer position is a highly visible/critical role, which will be a team member of technical SMEs managing the stability and optimization of the Wealth Management systems. Scope includes but not limited to, the day-to-day support of the organization’s technology related outages, collaboration on technology projects focused on stability, optimization, business impact analysis, and associated risk-related methodologies. This role will be responsible for overall stability of the Wealth Management Investment Management application platforms, participation on key optimization initiatives, and collaboration with multiple technical teams within . Additionally, partner with WM business units, various levels of management and staff to collect, analyze and make recommendations on optimizing the platform. This position will mainly perform DevOps/SRE role in Java, Unix & SQL technologies technology.

Responsibilities include:

  • Incident Management -Create and manage necessary process involving incidents
  • Partner with Ops Control to ensure IT and/or End User communications are handled appropriately
  • Engage with the development team throughout the life cycle to support Application build for Reliability
  • Develop software to automate manual operational work
  • Run, maintain and improve the service against established Service Level Objectives by applying software engineering principles
  • Responsible for the availability, performance, change (CP) management, monitoring, and capacity management of their services
  • Troubleshoot priority incidents, conduct blameless post-mortems, and ensure permanent closure of the incidents
  • Analyze patterns of production incidents, develop permanent remediation plans, and implement automation to prevent future incidents from occurring through software engineering
  • Manage process related functions around large-scale events such as disaster recovery. Communicate closely with impacted groups to ensure all events are properly managed.

Requirements

Primary Skills / Must have:

  • Site Reliability Engineer (SRE) in which 80% will be support [React/Protect], 10% will be in Dev Ops[Enable] space.
  • Proven track record supporting large scale multi-tiered cloud-based applications.
  • Analyze ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns
  • Hands on experience with Java, Angular, Spring, DB2, Unix scripting and experienced in scheduler tools such as TWS, autosys

  • L2-L3 Production Support, debugging skills, problem solving

  • Experience working in an Agile Development environment
  • Proven ability to understand and troubleshoot complex problems under pressure
  • Excellent communication skills (both written and oral), listening skills, influencing and negotiation skills
  • Experience with performance troubleshooting and remediation
  • Experience with observability tools such as Splunk, Kibana, Grafana, Prometheus
  • Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.

Secondary Skills / Desired skills

  • Having good expertise on Linux and shell scripting. Need to be very comfortable with Linux
  • Grafana/Kibana dashboarding experience
  • Good problem-solving skills
  • Good communicator
  • Good understanding of brokerage business
  • Jobs (controlM/CBSS/CRON) experience
  • Bachelor’s/master’s degree in computer science, Information Systems or related field

Skills: Agile Programming Methodologies, Analysis Skills, AngularJS, Automation, Best Practices, Brokerage, Business impact analysis (BIA), CA Workload Automation AE (AutoSys Edition), Capacity Management, Cloud Applications, Communication Skills, Computer Science, Continuous Deployment/Delivery, Continuous Integration, Cron Job Scheduling, Debugging Skills, DevOps, Disaster Recovery, IBM DB2, IT Service Management (ITSM), Identify Issues, Incident Management, Information Technology & Information Systems, Investment Management, Java, Linux Operating System, Negotiation Skills, Operations Control, Operations Planning, Pattern Analysis, Presentation/Verbal Skills, Problem Solving Skills, Process Improvement, Process Management, Production Management, Production Support, Reliability Engineering, Reporting Dashboards, Risk, SQL (Structured Query Language), Schedule Development, Software Administration, Software Development, Software Engineering, Splunk, Technical Support, Unix Operating Systems, Unix Shell Programming, Wealth Management, Writing Skills

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:03 min

Microsoft integrating native Unix coreutils into Windows environments

Chris Heilmann +2 · LIVE

1:02 min

Scheduling recurring automated operations through standard periodic cron jobs

Aurélie Vache Aurélie Vache · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:04 min

Defining timestamps and the international standard format

Denny Biasiolli Denny Biasiolli · Europe 2026 Virtual

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all