Site Reliability Engineering in Dallas

Energy Jobline
Dallas, TX, United States
4 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Agile Methodology Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Code Review Distributed Systems Perl (Programming Language) Python (Programming Language) Node.Js Reliability Engineering Ruby
+11 more
Service-Oriented Architecture Software Engineering Systems Integration Google Cloud Enterprise Software Applications Computer Network Technologies Kubernetes Software Performance Api Management Golang Microservices

Job description

  • You’ll have the opportunity to design and implement major infrastructure components, systems, and developer-friendly capabilities to improve the availability, scalability, latency, and efficiency of our services
  • You will provide technical leadership to cross-functional engineering, infrastructure, and product teams, and evangelize cloud best practices while building a culture of reliability and observability
  • Engage in and improve the end to end lifecycle of software development–from inception and design, through deployment, operation and refinement of a highly distributed system running in public cloud
  • Serve as subject matter expert in an SRE mindset, best practices, and cloud- principles
  • Scale systems sustainably through automation to improve reliability and velocity
  • Assist with all aspects of operational security and compliance
  • Run software performance analysis and system tuning
  • Design and implement tools to collect data from various sources and provide actionable insights
  • Participate in critical incident management and timely post-mortems of production incidents to drive practices around blameless analysis, resolution, and continuous improvement work with cross-functional teams Develop the rest of the team by conducting code reviews, providing mentorship, pairing, and training opportunities

Requirements

  • We are looking for Principal SRE with proven experience in running distributed systems at scale, in production
  • You have 15+ years of experience in relevant skills gained and developed in the same or similar role
  • Strong knowledge of container orchestration, preferably Kubernetes and networking technology
  • Hands-on experience in one or more , such as Node JS, Python, Go, Perl, Ruby, and Bash
  • Experience with SOA, Microservices architecture, API Management & Enterprise system Integrations
  • Strong production experience with cloud infrastructure, AWS, Azure & Google Cloud
  • Strong sense of ownership, and an ability to drive tasks to completion
  • Experience developing and monitoring distributed systems
  • Experience working in an Agile Environment with great collaboration skills

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.energyjobline.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

50 sec

Why developer happiness matters in web frameworks

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · WWC 2023

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

3:30 min

Falling in love with Ruby and creating Basecamp

David Heinemeier Hansson David Heinemeier Hansson +1 · Coffee With Developers

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

Videos

See all

Related articles

See all