SRE/DevOps Engineer- Palo Alto, the US

KODY, INC.
Palo Alto, CA, United States
2 months ago
Apply on indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Languages
Chinese, English
Job source

Tech stack

Amazon Web Services Continuous Integration DevOps Github Monitoring of Systems Reliability Engineering Software Configuration Management Software Deployment Data Logging Scripting Delivery Pipeline Git Flow

Job description

We are seeking a high-caliber Senior Site Reliability Engineer (SRE) based in California to ensure the scalability, reliability, and runtime efficiency of our next-generation platform. In this role, you will bridge the gap between development and operations, working closely with our global engineering teams.

We are looking for a unique engineering mindset: someone who brings a positive, collaborative energy to the daily grind, but can instantly pivot into a hyper-focused, high-ownership responder when an incident strikes., * Production Reliability & Guardrails: Partner with the Platform Engineering team to implement reliability guardrails, ensuring applications running on AWS meet strict uptime and SLA requirements.

  • CI/CD & Repository Management: Own the deployment pipelines and code management practices extensively via GitHub.
  • Incident Management: Lead rapid-response troubleshooting during production incidents; conduct thorough blameless post-mortems to continuously harden our systems.
  • Observability & Performance: Implement advanced monitoring, logging, and alerting systems to proactively detect and mitigate system anomalies.
  • Cross-Border Collaboration: Act as a key technical bridge between our US operations and international engineering hubs, leveraging bilingual communication to streamline complex technical alignment.

Requirements

Do you have experience in Triage?, + Ecosystem Expertise (Must-Haves): Deep, practical experience managing application deployment and runtime environments on AWS, alongside master-level knowledge of advanced Git workflows and actions on GitHub.

  • Core Toolkit: Strong proficiency in monitoring tools, log management, and scripting for quick triaging and troubleshooting., + Ownership & Transparency: You are radically open, highly responsive, and communicative. You don’t just clear tickets; you own the production environment’s health end-to-end.
  • Pressure-Resistance: High psychological resilience. You maintain a happy, positive attitude during smooth operations, yet feel a healthy, driving sense of urgency and laser-focus during high-stakes incidents.
  • Bilingual Capability: Absolute fluency in Mandarin and English (verbal and written) is mandatory for effective technical alignment across our global teams.

Benefits & conditions

  • Competitive packages aligned with California market standards
  • Lead a dynamic and innovative team in a very rapidly growing company
  • Collaborative, inclusive environment where your contributions are recognized and valued

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

5:02 min

Mapping Git flow branches to application tester segments

Majid Hajian · LIVE

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:53 min

Introduction to git flow and clean feature branches

Johannes Haux · World Congress 2022

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

Videos

See all

Related articles

See all