SRE - Site Reliability Engineer

TD Ameritrade
Austin, TX, United States
7 days ago
Apply on www.schwabjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$150,000.0 - $162,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) .NET Framework Application Services Bash Shell Software as a Service Cloud Foundry Databases System Configuration Dynamic Host Configuration Protocol Disaster Recovery Distributed Systems Domain Name System (DNS)
+26 more
IBM WebSphere MQ Internet Protocol Python (Programming Language) Linux System Administration Microsoft SQL Server Windows Servers MongoDB Routing Oracle (Applications) Performance Tuning Windows PowerShell Systems Development Life Cycle RabbitMQ Reliability Engineering Software Engineering Software Vulnerability Management Scripting Google Cloud System Availability Firewalls (Computer Science) Tanzu Information Technology Apache Kafka Actimize Splunk Appdynamics

Job description

As a Senior Site Reliability Engineer, you will serve as a technical leader responsible for advancing platform reliability, resiliency, and operational excellence across complex distributed systems. This role combines engineering expertise, problem-solving, and operational leadership to deliver scalable solutions that improve system stability, accelerate issue resolution, and reduce operational risk. You will collaborate closely with software engineering, infrastructure, security, architecture, and business teams to modernize platforms, strengthen observability, automate operational processes, and support critical business outcomes.

Success in this role requires balancing strategic thinking with hands-on technical execution while influencing cross-functional partners, driving continuous improvement initiatives, and ensuring systems remain secure, compliant, and highly available. You will play a key leadership role in incident management, recovery planning, infrastructure coordination, and the adoption of Site Reliability Engineering (SRE) practices that enhance client experiences and business resilience.

As part of the operational support model, you will participate in an on-call rotation approximately once every 5 to 6 weeks, providing support for critical production systems and helping ensure timely response and resolution of service-impacting incidents.

Requirements

  • 6+ years of experience supporting and administering enterprise technology platforms in large-scale environments.
  • 6+ years of experience with automation, scripting, monitoring solutions, alert management, and operational process improvement.
  • 6+ years of experience working within Software Development Lifecycle (SDLC) practices and supporting continuous improvement initiatives.
  • Experience supporting high-availability distributed systems, production operations, and platform reliability initiatives.
  • Experience leading incident response, root cause analysis, and service recovery efforts for mission-critical applications.
  • Experience with Windows Server (2019/2022) and Linux system administration, troubleshooting, performance tuning, and operational support.
  • Experience deploying, configuring, supporting, or migrating cloud-based applications and infrastructure.
  • Knowledge of IP networking concepts including DNS, DHCP, firewalls, and routing.
  • Development or scripting experience using one or more technologies such as PowerShell, Python, Java, .NET, or Bash.
  • Experience working with database technologies such as SQL Server, Oracle, MongoDB, or similar platforms.
  • Experience supporting messaging and event-driven technologies such as Kafka, RabbitMQ, IBM MQ, or Solace.
  • Experience using observability and monitoring platforms such as Splunk, AppDynamics, or equivalent tools.
  • Ability to analyze complex technical issues, make sound operational decisions, and communicate recommendations effectively to technical and non-technical audiences.
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field., * 8+ years of experience supporting large-scale, mission-critical platforms within financial services or other highly regulated industries.
  • Experience implementing and scaling Site Reliability Engineering (SRE) practices, including Service Level Objectives (SLOs), post-incident reviews, observability, and reliability metrics.
  • Strong background in production operations, availability engineering, and operational risk management.
  • Experience leading infrastructure modernization initiatives, disaster recovery planning, vulnerability remediation, and security-focused operational programs.
  • Experience partnering across engineering, infrastructure, security, vendor, and business teams to deliver complex technology solutions.
  • Familiarity with audit, regulatory, PCI, security, and compliance requirements within banking or financial services environments.
  • Experience designing and implementing automation solutions that reduce operational overhead and improve reliability outcomes.
  • Demonstrated ability to mentor engineers, establish operational standards, and promote a culture of accountability and continuous improvement.
  • Experience with Google Cloud Platform (GCP), Tanzu Application Service/Cloud Foundry (PCF), or similar cloud platforms.
  • Working knowledge of Actimize and related financial crime or regulatory technology platforms.

Applicants must be currently authorized to work in the United States on a full-time basis without employer sponsorship.

Benefits & conditions

We offer a competitive benefits package that takes care of the whole you - both today and in the future:

  • 401(k) with company match and Employee stock purchase plan
  • Paid time for vacation, volunteering, and 28-day sabbatical after every 5 years of service for eligible positions
  • Paid parental leave and family building benefits
  • Tuition reimbursement
  • Health, dental, and vision insurance

About the company

At Schwab, you’re empowered to make an impact on your career. Here, innovative thought meets creative problem solving, helping us challenge the status quo and transform the finance industry together. We believe in the importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).

Schwab Technology Services (STS) enables the future of how clients manage their money by delivering innovative and reliable technology solutions that support investing, banking, and financial planning. Within the Bank Platform Operations and Engineering organization, you will help ensure the reliability, availability, and operational readiness of more than 300 mission-critical banking applications that support client transactions and highly regulated business functions., At Schwab, you’re empowered to shape your future. We champion your growth through meaningful work, continuous learning, and a culture of trust and collaboration-so you can build the skills to make a lasting impact. Our Hybrid Work and Flexibility approach balances our ongoing commitment to workplace flexibility, serving our clients, and our strong belief in the value of being together in person on a regular basis.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.schwabjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:58 min

Configuring a baseline isolated agent on Tanzu platform

Oren Penso Oren Penso · World Congress 2026 Europe

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

1:51 min

Overview of the three Google Maps routing applications

Germán Álvarez · LIVE

Videos

See all

Related articles

See all