> Markdown version of [/jobs/ext/2156258-cloud-sre-engineer](https://www.wearedevelopers.com/jobs/ext/2156258-cloud-sre-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud SRE Engineer - **Company:** After School Matters, Inc. - **Location:** Palo Alto, CA, United States - **Salary:** $145,600.0 - $208,000.0 - **Contract:** Temporary contract - **Skills:** Amazon Web Services, Microsoft Azure, Software Bug Management, Cloud Computing, Cloud Computing Security, Software Debugging, Linux, DevOps, Disaster Recovery, Domain Name System (DNS), Hypertext Transfer Protocols (HTTP), Internet Control Message Protocol, Python (Programming Language), Nagios, Prometheus, Runbook, TCP/IP, Scripting, Google Cloud, Load Balancing, System Availability, Grafana, Software Troubleshooting, Firewalls (Computer Science), Amazon Virtual Private Cloud (VPC), Kubernetes, Information Technology, Cloudwatch - **Published:** August 20, 2026 - **Apply:** https://intellipro.applytojob.com/apply/R4PfTsWHf3/Cloud-SRE-Engineer-Mandarin-Bilingual?source=GS ## About the Role * Some SRE, DevOps, or cloud operations experience - ability to maintain application stability independently is essential given timezone constraints * Mandarin/English bilingual preferred - ability to communicate with teams in China and Singapore is a plus * Strong networking fundamentals (TCP/IP, DNS, HTTP, ICMP, load balancing, firewalls, VPC) OR deep Linux/CVM knowledge - ability to own either the networking or compute side of operations * Hands-on experience with cloud platforms (AWS, GCP, Azure, or equivalent) - deployment, usage, and high availability * Familiarity with Kubernetes and container-based deployments * Proficiency in at least one scripting language (Python, Shell, or Go) with automation experience * Strong troubleshooting and debugging skills across infrastructure layers * Experience with monitoring and alerting tools (Grafana, Prometheus, CloudWatch, or equivalent) * Bachelor's degree or above in Computer Science or a related field * Strong self-directed work ethic - able to operate independently with minimal supervision across time zones ## Description North America cloud operations team is looking for a skilled Cloud SRE Engineer to own the reliability, stability, and continuous improvement of core cloud services - spanning compute infrastructure (CVM/VMs), networking, and cloud security products. You'll work in a production-critical environment where operational excellence, deep technical expertise, and a self-directed mindset are essential. Since the North America team operates independently from teams in China and Singapore with no overlapping hours, we're looking for someone who can hit the ground running with minimal ramp-up time., * Monitor and maintain cloud compute (CVM), networking, and security products in the North America region to ensure high availability and system stability * Respond to and resolve production incidents, customer-reported issues, and system-level outages with urgency and ownership * Perform deep troubleshooting across network, compute, security, and platform layers * Participate in on-call rotation and handle live production issues independently * Deploy new features, bug fixes, and enhancements into production environments using CI/CD pipelines and internal tooling * Develop scripts and automation tools to improve operational efficiency and reduce toil * Build and improve monitoring, alerting, and disaster recovery systems for 24/7 operations * Document operational workflows, runbooks, and best practices * Work closely with R&D, security, and platform teams across time zones to drive service reliability * Communicate technical issues clearly to internal teams and B2B customers ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [We adopted DevOps and are Cloud-native, Now What?](https://www.wearedevelopers.com/videos/485-we-adopted-devops-and-are-cloud-native-now-what) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)