Cloud Infrastructure - Site Reliability Engineer (SRE)-Sunnyvale in Sunnyvale

Energy Jobline
Sunnyvale, CA, United States
19 days ago
Apply on www.energyjobline.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$171,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Cloud Computing Disaster Recovery Distributed Systems Middleware Monitoring of Systems Python (Programming Language) Log Analysis Performance Tuning Reliability Engineering Management of Software Versions Diagnostic Tools
+7 more
Scripting Autoscaling Kubernetes Infrastructure Automation Frameworks Apache Kafka Terraform Dynatrace

Job description

Alibaba Cloud Message Middleware Team is responsible for message products, including RocketMQ and other messaging products. We are committed to creating a more stable, user-friendly, streaming, and large-scale messaging platform for the future., Oversee stability maintenance, performance tuning, and high-availability architecture design for cloud middleware, including messaging middleware (Kafka/RocketMQ).

\n

Manage the containerized middleware lifecycle on Kubernetes clusters: implement deployments, auto-scaling, version upgrades, and resource optimization in K8s environments.

\n

Incident Response & Root Cause Analysis

\n

Lead the troubleshooting of middleware-related incidents (e.g., message backlog, service registration failures) through log analysis, distributed tracing, and monitoring systems.

\n

Develop diagnostic tools using Java/Go to resolve production issues, performance bottlenecks, and compatibility challenges.

\n

Automation & Operational Excellence

\n

Build Python/Go/Shell automation tools to standardize middleware deployment, monitoring, and disaster recovery workflows.

\n

Implement chaos engineering experiments, capacity planning strategies, and failover mechanisms to enhance system resilience.

Requirements

Strong scripting skills in Shell/Python and experience with Infrastructure as Code (IaC) tools (Terraform )., Experience: Over 2 years of experience in distributed systems reliability engineering, familiar with high-availability architecture design, and proficient in at least one of Python, Go, or Java.

\n

Messaging: Cluster management, message reliability assurance, and performance optimization for Kafka/RocketMQ.

\n

Hands-on Experience Deploying Middleware On Kubernetes (Helm/Operator ).

\n

Automation: Ability to convert operations experience into automated solutions and familiarity with various message middleware, e.g., Kafka and RocketMQ.

Benefits & conditions

SRE Practices: Familiar with core SRE practices (incident review, error budgeting, chaos engineering) and experienced in building automated risk control systems.

\n

The pay range for this position at commencement of employment is expected to be between $104,400 and $171,000/year. However, base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience.

\n

If hired, employee will be in an “at-will position” and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.energyjobline.com
Prepare application

Good distractions

Loading talks and stories from around this role…