Site Reliability Engineer

Fmr LLC
Westlake, TX, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

JavaScript (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Data Analysis Microsoft Azure Code Generation DevOps Distributed Systems Python (Programming Language) Microsoft Visual Studio Node.Js
+11 more
Windows PowerShell Reliability Engineering Software Engineering SQL Databases TypeScript Data Logging GitHub Copilot Large Language Models Reliability of Systems Api Design Software Coding

Job description

  • Build software solutions to improve reliability, reduce operational toil, and scale production systems-not just respond to issues.
  • Develop production-quality code using Node.js / JavaScript / TypeScript and Python/PowerShell, including testing and documentation.
  • Leverage modern development tooling such as VS Code and AI-assisted tools (e.g., GitHub Copilot) to accelerate delivery and problem-solving.
  • Independently own well-scoped features, fixes, or improvements end-to-end-from design through deployment and operational validation.
  • Participate in on-call rotations, respond to incidents, execute runbooks, and ensure clear communication and handoffs.
  • Analyze incidents and recurring issues to identify patterns, reduce alert noise, and implement durable fixes.
  • Implement and improve observability (logging, metrics, dashboards, alerts) for owned services.
  • Build automations, scripts, and lightweight tools to eliminate repetitive manual work and improve operational efficiency.
  • Identify and act on opportunities to improve system reliability, performance, and maintainability.
  • Develop an understanding of how systems impact customer experience and business outcomes.

Requirements

Do you have experience in Visual Studio Code?, * ~2 plus years of experience in SRE, software engineering, DevOps, or production engineering.

  • Strong hands-on coding skills with emphasis on:
  • Node.js / JavaScript / TypeScript
  • Python for scripting and automation
  • Experience building tools, APIs, or automations to solve engineering or operational problems.
  • Familiarity with AI-assisted development workflows (e.g., GitHub Copilot, code generation tools) and interest in applying AI/LLMs to improve engineering productivity.
  • Foundational knowledge of:
  • Monitoring, logging, and observability concepts
  • Distributed systems and API-based architectures
  • SQL and data analysis for troubleshooting
  • Exposure to cloud platforms (AWS or Azure), CI/CD pipelines, and modern development practices.
  • Basic understanding of incident management, problem management, and production support processes., * A self-starter who proactively identifies key issues and trends, performs thoughtful analysis, and develops creative, high-impact solutions that deliver measurable value.
  • A builder mindset, you instinctively solve operational problems by writing code and creating automation.
  • A modern engineer who leverages AI tools to increase speed, quality, and effectiveness.
  • Strong problem-solving skills with the ability to troubleshoot, analyze, and deliver practical solutions.
  • Proactive approach to identifying risks, inefficiencies, and improvement opportunities.
  • Curiosity and desire to continuously learn systems, tools, and reliability engineering practices.
  • Clear communicator who collaborates effectively across engineering and operations teams.
  • Growing autonomy and consistency aligned with progression toward a Grade 5 SRE role.

About the company

We are looking for engineers who solve operational problems by building software. In this role, you will improve reliability, reduce toil, and enhance production systems by writing code, building automations, and leveraging modern AI-assisted development tools.

This is a hands-on engineering role-not a traditional support position. You’ll use Node.js/TypeScript, Python, and AI tools like GitHub Copilot to design and deliver solutions that make systems more reliable and operations more scalable.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · WWC 2023

3:10 min

Understanding the core concepts of API design

Alen Pokos · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

5:00 min

Exploring the specific workplace responsibilities of staff software engineers

Jan Giacomelli · LIVE

Videos

See all

Related articles

See all