SRE Engineer

Balin Technologies Llc
San Jose, CA, United States
about 2 months ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Clean Code Principles Cloud Computing Software Debugging Software Design Patterns File Systems Distributed Data Store Memory Management Python (Programming Language) Linux Kernel Network Programming Ansible Subsystems
+6 more
TCP/IP Gitlab-ci Kubernetes Bare Metal Terraform Programming Languages

Job description

Primary Skills: Python Coding, (Ansible, Terraform, Kubernetes) and CI/CD practices (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure. Secondary Skills: TCP/IP and network programming., SRE Engineer

  • You will engage in incident response drills, post-mortems, and root cause analysis sessions to learn from past issues and prevent future ones.
  • Each morning starts with a structured review of overnight alerts and system performance metrics - identifying any anomalies, triaging what needs attention.
  • You will collaborate with your team in a morning stand-up meeting to discuss ongoing projects, recent incidents, and priorities for the day..
  • Your tasks will include automating routine processes, analyzing system logs, and developing tools to enhance our monitoring capabilities.
  • You’‘ll spend part of your day working closely with software engineers, advising on best practices for resilient code and reviewing changes before deployment.
  • Throughout the day, your focus is on maintaining high SLIs and SLOs, ensuring that our infrastructure remains robust and reliable for our customers.
  • By days end, you will document your work, share insights with your team, and plan for the next days challenges, always with a customer-centric mindset.

Requirements

  • Strong experience with architecture, design patterns, reliability and scaling of new and current systems.
  • Experience leading and commanding incidents, including driving root cause analysis, coordinating cross-functional teams, and ensuring follow-through on corrective actions.
  • Experience building observability from the ground up — defining SLOs/SLIs, closing monitoring gaps, and implementing alerting strategies that catch failures before customers do.
  • Proficiency in Linux kernel internals, with exposure to scheduler, memory allocation, and driver subsystem.
  • Experience writing high quality code with at least one programming language (Python, Go, or similar).
  • Experience with system-level debugging, including kdump, and kernel panic analysis..
  • Proficiency in Infrastructure as Code tooling (Ansible, Terraform, Kubernetes) and CI/CD practices (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure.. Experience with TCP/IP and network programming.
  • Experience with distributed storage systems and understanding of one or more of object, block, and file storage paradigms..
  • Hardware and GPU troubleshooting experience (nice to have).
  • Exposure to OVN/OVS-based networking stack (nice to have).
  • Strong communication skills.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

2:32 min

Overview of Terraform and Terraform Cloud features

Devlin Duldulao · LIVE

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · World Congress 2025

Videos

See all

Related articles

See all