Site Reliability Engineer
Hptech Inc.
Los Angeles, CA, United States
28 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source
Tech stack
Java (Programming Language)
Apache ActiveMQ
Systems Engineering
Microsoft Azure
Bash Shell
Cloud Computing
Databases
Linux
DevOps
Distributed Systems
Middleware
Fault Tolerance
+21 more
Monitoring of Systems
Web Servers
Python (Programming Language)
PostgreSQL
Linux Kernel
Load Testing
Enterprise Messaging Systems
Operational Databases
Performance Tuning
Windows PowerShell
Reliability Engineering
Software Engineering
Scripting
Load Balancing
Delivery Pipeline
Large Language Models
Database Performance
Indexer
Kubernetes
Low Latency
Data Analytics
Job description
- We are seeking an experienced engineer who can analyze, diagnose, and optimize performance and reliability of large-scale distributed systems. This role requires deep technical understanding across the entire application stack, the ability to read and reason about code, and the capability to provide data-backed answers to both engineering teams and business stakeholders.
- This role goes beyond traditional operations or DevOps. The successful candidate will think like a software engineer, act like a systems engineer, and operate with a production-first mindset., * Performance & Reliability Engineering
- Analyze and resolve performance issues such as high latency, slow login, throughput degradation, and system instability.
- Perform deep, end-to-end investigations across the full stack including:
- Load balancers and traffic routing
- Web server and application runtime configurations
- Middleware and messaging systems
- Database performance (queries, indexing, pooling)
- Kubernetes clusters (pods, resources, scaling behavior)
- Linux OS tuning (CPU, memory, I/O, ulimits, networking)
- Identify root causes and propose clear, actionable engineering solutions.
- Distributed Systems Design
- Design, review, and influence high-performance, highly-available distributed architectures.
- Evaluate trade-offs related to scalability, latency, fault tolerance, and cost.
- Partner with development teams early to prevent reliability and performance issues before production.
- Capacity Planning & Scalability
- Assess system readiness for growth scenarios such as:
- “We plan to onboard 10,000 users in 6 months - can the system support it?”
- Perform capacity and scale analysis for:
- Application tiers
- Databases
- Messaging systems
- Kubernetes compute and storage
- Provide evidence-based recommendations supported by metrics, benchmarks, and production data.
- Engineering Collaboration
- Work closely with software engineering teams to:
- Review performance-critical code paths
- Propose improvements at code, configuration, or infrastructure level
- Improve system observability (metrics, logs, traces)
- Communicate complex technical findings clearly to both engineers and business stakeholders.
Requirements
- Strong understanding of distributed systems and performance engineering
- Ability to read, analyze, and troubleshoot Java code
- Hands-on experience with:
- Kubernetes (resource management, scaling, container behavior)
- Linux internals and tuning
- PostgreSQL (queries, indexing, performance optimization)
- Proven experience building or operating high-availability, high-throughput systems
- Strong analytical and problem-solving skills with a data-driven approach
- Nice to Have
- Experience with Azure cloud services
- Messaging systems such as ActiveMQ
- Load testing and benchmarking experience
- Background in roles such as SRE, Performance Engineering, Platform Engineering, Technical Skills
- Hands-on experience with cloud platforms (Azure.
- Strong scripting skills (e.g., Python, Bash, PowerShell, or similar).
- Experience with deployment pipelines, automation, and monitoring tools.
- Solid understanding of cloud infrastructure, networking, and application operations.
LLM & AI Experience
- Practical experience working with Large Language Models (LLMs).
- Familiarity with applying LLMs to engineering or operational workflows is required.
Professional Attributes
- Strong desire to learn and deeply understand complex systems.
- Self-starter with the ability to take ownership and drive initiatives independently.
- Demonstrates leadership, accountability, and problem-solving mindset.
Strong collaboration and communication skills
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
IK
Igor Khokhriakov
24 days ago
LM
Luis Minvielle
What Are Large Language Models?
almost 3 years ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
CH
Chris Heilmann
Dev Digest 121 - AI goes offline
over 2 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
LM
Luis Minvielle
Why Upskilling And Reskilling is Important For Developers
over 2 years ago