> Markdown version of [/jobs/ext/2103903-sre-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2103903-sre-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE - Site Reliability Engineer - **Company:** TD Ameritrade - **Location:** Austin, TX, United States - **Experience:** Expert - **Salary:** $150,000.0 - $162,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), .NET Framework, Application Services, Bash Shell, Software as a Service, Cloud Foundry, Databases, System Configuration, Dynamic Host Configuration Protocol, Disaster Recovery, Distributed Systems, Domain Name System (DNS), IBM WebSphere MQ, Internet Protocol, Python (Programming Language), Linux System Administration, Microsoft SQL Server, Windows Servers, MongoDB, Routing, Oracle (Applications), Performance Tuning, Windows PowerShell, Systems Development Life Cycle, RabbitMQ, Reliability Engineering, Software Engineering, Software Vulnerability Management, Scripting, Google Cloud, System Availability, Firewalls (Computer Science), Tanzu, Information Technology, Apache Kafka, Actimize, Splunk, Appdynamics - **Published:** August 18, 2026 - **Apply:** https://www.schwabjobs.com/financial-consultant-academy-opportunity-branch-network ## About the Role * 6+ years of experience supporting and administering enterprise technology platforms in large-scale environments. * 6+ years of experience with automation, scripting, monitoring solutions, alert management, and operational process improvement. * 6+ years of experience working within Software Development Lifecycle (SDLC) practices and supporting continuous improvement initiatives. * Experience supporting high-availability distributed systems, production operations, and platform reliability initiatives. * Experience leading incident response, root cause analysis, and service recovery efforts for mission-critical applications. * Experience with Windows Server (2019/2022) and Linux system administration, troubleshooting, performance tuning, and operational support. * Experience deploying, configuring, supporting, or migrating cloud-based applications and infrastructure. * Knowledge of IP networking concepts including DNS, DHCP, firewalls, and routing. * Development or scripting experience using one or more technologies such as PowerShell, Python, Java, .NET, or Bash. * Experience working with database technologies such as SQL Server, Oracle, MongoDB, or similar platforms. * Experience supporting messaging and event-driven technologies such as Kafka, RabbitMQ, IBM MQ, or Solace. * Experience using observability and monitoring platforms such as Splunk, AppDynamics, or equivalent tools. * Ability to analyze complex technical issues, make sound operational decisions, and communicate recommendations effectively to technical and non-technical audiences. * Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field., * 8+ years of experience supporting large-scale, mission-critical platforms within financial services or other highly regulated industries. * Experience implementing and scaling Site Reliability Engineering (SRE) practices, including Service Level Objectives (SLOs), post-incident reviews, observability, and reliability metrics. * Strong background in production operations, availability engineering, and operational risk management. * Experience leading infrastructure modernization initiatives, disaster recovery planning, vulnerability remediation, and security-focused operational programs. * Experience partnering across engineering, infrastructure, security, vendor, and business teams to deliver complex technology solutions. * Familiarity with audit, regulatory, PCI, security, and compliance requirements within banking or financial services environments. * Experience designing and implementing automation solutions that reduce operational overhead and improve reliability outcomes. * Demonstrated ability to mentor engineers, establish operational standards, and promote a culture of accountability and continuous improvement. * Experience with Google Cloud Platform (GCP), Tanzu Application Service/Cloud Foundry (PCF), or similar cloud platforms. * Working knowledge of Actimize and related financial crime or regulatory technology platforms. Applicants must be currently authorized to work in the United States on a full-time basis without employer sponsorship. ## Description As a Senior Site Reliability Engineer, you will serve as a technical leader responsible for advancing platform reliability, resiliency, and operational excellence across complex distributed systems. This role combines engineering expertise, problem-solving, and operational leadership to deliver scalable solutions that improve system stability, accelerate issue resolution, and reduce operational risk. You will collaborate closely with software engineering, infrastructure, security, architecture, and business teams to modernize platforms, strengthen observability, automate operational processes, and support critical business outcomes. Success in this role requires balancing strategic thinking with hands-on technical execution while influencing cross-functional partners, driving continuous improvement initiatives, and ensuring systems remain secure, compliant, and highly available. You will play a key leadership role in incident management, recovery planning, infrastructure coordination, and the adoption of Site Reliability Engineering (SRE) practices that enhance client experiences and business resilience. As part of the operational support model, you will participate in an on-call rotation approximately once every 5 to 6 weeks, providing support for critical production systems and helping ensure timely response and resolution of service-impacting incidents. ## Related Videos - [The Private AI Platform: Why Agentic Apps Need a Private Application Platform](https://www.wearedevelopers.com/videos/100162-the-private-ai-platform-why-agentic-apps-need-a-private-application-platform) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [A Technical Introduction to Bitcoin's 2nd Layer- The Lightning Network](https://www.wearedevelopers.com/videos/15-a-technical-introduction-to-bitcoin-s-2nd-layer-the-lightning-network) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)