DevOps Engineer (Bilingual Chinese)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+26 more
Job description
The company is seeking a highly skilled and passionate DevOps Engineer to join our cutting-edge research and development team. In this role, you will bridge the gap between traditional cloud infrastructure and the rapidly evolving world of Generative AI. You will be responsible for scaling global containerized environments, automating complex workflows, and building the foundational infrastructure that powers our next-generation AI tools., * Cloud-Native Architecture: Responsible for the planning and management of AWS cloud and private data centers, implementing Infrastructure as Code (IaC).
- Container Orchestration: Maintain large-scale Kubernetes (EKS) clusters, handling cluster upgrades, scaling, network configuration (CNI), and storage management.
- DevOps Development: Develop automated operations tools, CLIs, or platforms using Python or Go to eliminate repetitive tasks and improve R&D efficiency.
- Observability: Build full-link monitoring and logging systems based on Prometheus, Grafana, and ELK to ensure system stability and rapid troubleshooting.
- AI Infrastructure: * Manage GPU server resources (Nvidia A100/T4, etc.), and optimize driver versions and CUDA environments.
- Responsible for the containerized deployment and concurrency optimization of generative AI tools such as ComfyUI.
Requirements
- Background & Experience: Bachelor’s degree or higher in Computer Science or a related field, with 3+ years of experience in DevOps or SRE.
- Cloud & Containers: Proficient in core AWS services (EC2, EKS, S3, VPC); deep understanding of Kubernetes architecture and scheduling principles.
- Development Skills: Must possess strong programming capabilities, proficiency in Python or Go, and hands-on experience in tool development or backend development.
- Systems & Networking: Deep understanding of the Linux operating system, network protocols (TCP/IP, HTTP, DNS), and Shell scripting.
- CI/CD: Highly proficient in designing pipelines using Jenkins, GitLab CI, or GitHub Actions.
Bonus Points (Nice-to-Haves)
- AI/GPU Operations: Experience in large-scale GPU cluster operations, familiar with GPU memory monitoring, resource partitioning, or Spot instance cost optimization.
- AIGC Hands-on Experience: Practical experience with generative AI, familiarity with the deployment architecture of ComfyUI and Stable Diffusion WebUI, and experience in resolving dependency management, multi-user concurrency, or model loading acceleration.
- MLOps/AIOps: Familiarity with Kubeflow, MLflow, or Triton Inference Server is a strong plus.
- High-Performance Computing (HPC): Experience with RDMA networking or large-scale data parallel processing.
Benefits & conditions
$80,000 - $150,000 a year - Full-time, Pulled from the full job description
- 401(k)
- Health insurance
- Retirement plan
- Paid time off
- Vision insurance
- Dental insurance
- Flexible schedule, * 401(k)
- Dental insurance
- Flexible schedule
- Health insurance
- Paid time off
- Retirement plan
- Vision insurance
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 121 - AI goes offline
Highest Paying Tech Companies for Developers
Dev Digest 120 - Apple and peers
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again