SRE Engineer
Balin Technologies Llc
San Jose, CA, United States
about 2 months ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source
Tech stack
Clean Code Principles
Cloud Computing
Software Debugging
Software Design Patterns
File Systems
Distributed Data Store
Memory Management
Python (Programming Language)
Linux Kernel
Network Programming
Ansible
Subsystems
+6 more
TCP/IP
Gitlab-ci
Kubernetes
Bare Metal
Terraform
Programming Languages
Job description
Primary Skills: Python Coding, (Ansible, Terraform, Kubernetes) and CI/CD practices (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure. Secondary Skills: TCP/IP and network programming., SRE Engineer
- You will engage in incident response drills, post-mortems, and root cause analysis sessions to learn from past issues and prevent future ones.
- Each morning starts with a structured review of overnight alerts and system performance metrics - identifying any anomalies, triaging what needs attention.
- You will collaborate with your team in a morning stand-up meeting to discuss ongoing projects, recent incidents, and priorities for the day..
- Your tasks will include automating routine processes, analyzing system logs, and developing tools to enhance our monitoring capabilities.
- You’‘ll spend part of your day working closely with software engineers, advising on best practices for resilient code and reviewing changes before deployment.
- Throughout the day, your focus is on maintaining high SLIs and SLOs, ensuring that our infrastructure remains robust and reliable for our customers.
- By days end, you will document your work, share insights with your team, and plan for the next days challenges, always with a customer-centric mindset.
Requirements
- Strong experience with architecture, design patterns, reliability and scaling of new and current systems.
- Experience leading and commanding incidents, including driving root cause analysis, coordinating cross-functional teams, and ensuring follow-through on corrective actions.
- Experience building observability from the ground up — defining SLOs/SLIs, closing monitoring gaps, and implementing alerting strategies that catch failures before customers do.
- Proficiency in Linux kernel internals, with exposure to scheduler, memory allocation, and driver subsystem.
- Experience writing high quality code with at least one programming language (Python, Go, or similar).
- Experience with system-level debugging, including kdump, and kernel panic analysis..
- Proficiency in Infrastructure as Code tooling (Ansible, Terraform, Kubernetes) and CI/CD practices (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure.. Experience with TCP/IP and network programming.
- Experience with distributed storage systems and understanding of one or more of object, block, and file storage paradigms..
- Hardware and GPU troubleshooting experience (nice to have).
- Exposure to OVN/OVS-based networking stack (nice to have).
- Strong communication skills.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
IK
Igor Khokhriakov
24 days ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
KM
Kaleb McKelvey
The Best Software Developer Blogs to Read
over 3 years ago
JF
Jonas Fritzsch
Résumé-Driven Development: How IT trends affect the job market for software developers
almost 5 years ago