> Markdown version of [/jobs/ext/2153569-site-reliability-engineering-sre-the-core-engineering-vice-president-dallas](https://www.wearedevelopers.com/jobs/ext/2153569-site-reliability-engineering-sre-the-core-engineering-vice-president-dallas). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineering (SRE), The Core Engineering, Vice President, Dallas - **Company:** The Goldman Sachs Group Inc - **Location:** Dallas, TX, United States - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Java (Programming Language), Amazon Web Services, Systems Engineering, Automation of Tests, Microsoft Azure, Cloud Engineering, Data Structures, Linux, Distributed Systems, Fault Tolerance, Python (Programming Language), Load Testing, Network Protocols, Node.Js, Performance Tuning, Systems Development Life Cycle, Reliability Engineering, Ansible, Prometheus, Software Engineering, Datadog, Data Logging, Load Balancing, Grafana, Software Application Programming, Rate Limiting, Cloudformation, Kubernetes, Information Technology, Cloudwatch, Terraform, Splunk, Dynatrace, Docker, Programming Languages - **Published:** August 20, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/87179848/1 ## About the Role * Strong proficiency in at least one major programming language (e.g., Java, Python, or Node.js) with a focus on writing clean, maintainable code for tooling and automation. * Hands-on experience with Infrastructure as Code (IaC) frameworks such as Terraform, Ansible, or CloudFormation. * Deep understanding of containerization and orchestration technologies, specifically Docker and Kubernetes (K8s), including service meshes and ingress controllers. * Advanced experience with major cloud providers (AWS, GCP, or Azure), specifically building and operating highly resilient cloud-native architectures. * Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch) * Experience with automated testing and SDLC concepts, developing applications in a Linux environment, and sound knowledge of algorithms, data structures and software design. * Knowledge of networking protocols and load balancing strategies in a distributed systems environment. Core Competencies & Soft Skills * Ability to analyze complex, distributed systems holistically and understand how individual components interact under load. * Strong interpersonal skills to collaborate with product developers, influence architectural decisions, prioritize toil reduction, and drive SRE adoption without direct authority. * Ability to translate complex technical issues into clear, actionable insights for both technical and non-technical stakeholders. * Highly motivated, pro-active and capable of multi-tasking under pressure in a fast-paced environment without compromising quality. * Commitment to fostering a blameless culture where failures are treated as opportunities to learn and improve systems. * Interest in financial markets and technology. Preferred Qualifications * Bachelor's degree in Computer Science, System Engineering, or a related technical field that involves programming. * 7 to 10 years of experience ## Description * Partner with engineering leadership to establish service level objectives (SLOs), service level indicators (SLIs), and error budgets. * Collaborate with product developers to architect highly available, fault-tolerant, and self-healing systems. Conduct architectural reviews and introduce patterns like circuit breakers, graceful degradation, and rate limiting. * Reduce operational toil by building automation, tooling, and self-service capabilities that remove repetitive manual work. * Improve production readiness through load testing, performance tuning, capacity forecasting, and reliability reviews. * Lead the response to complex, multi-system production incidents. Facilitate blameless post-mortems to identify root causes and drive long-term preventative actions. * Promote sustainable operations by helping design healthy on-call models, clear escalation paths, and balanced pager responsibilities. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [How Much Does a Software Engineer Make? Realistic Software Engineering Salaries](https://www.wearedevelopers.com/magazine/425-how-much-does-a-software-engineer-make-realistic-software-engineering-salaries)