> Markdown version of [/jobs/ext/2720775-automated-testing-engineer](https://www.wearedevelopers.com/jobs/ext/2720775-automated-testing-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Automated Testing Engineer - **Company:** Crusoe's Inc - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $172,500.0 - $210,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Automation of Tests, Bash Shell, Cloud Computing, Computer Clusters, Nvidia CUDA, Continuous Integration, Software Debugging, Linux, Memory Management, InfiniBand, Python (Programming Language), PostgreSQL, Node.Js, PCI Express, Remote Direct Memory Access, Virtualization Technology, Network Switches, Graphics Processing Unit (GPU), Computer Network Technologies, Delivery Pipeline, Gitlab, Integration Tests, Kubernetes, Information Technology, Drilldown, Terraform, Docker - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/automated-testing-engineer-compute-crusoe-7919474 ## About the Role * Education & Experience: 5+ YOE demonstrated ability to competently and independently perform responsibilities plus Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related technical field. * Experience building and deploying automated integration testing for an AI Cloud Environment, ranging from low-level Linux Systems up to Distributed Control Planes. * Working knowledge of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres. * CI/CD & Gitlab: Intimate knowledge of CI/CD pipelines and Gitlab Tooling to enable stable infrastructure releases across multiple datacenters. * Automation & Scripting: Advanced proficiency in Python and/or Bash for automating complex cluster-wide test scenarios. * Distributed GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a multi-node context. * Networking Knowledge: Strong understanding of RDMA, RoCE, and InfiniBand protocols and their implementation in virtualized systems. * System Internals: Knowledge of Linux kernel internals, specifically PCIe topology, VFIO, and memory management (HugePages, IOMMU). Bonus Points: * Experience with MNNVL (Multi-Node NVLink) or specialized AI fabric architectures. * Familiarity with hardware-level debugging tools and performance profilers (e.g., NVIDIA Nsight, AMD Omniperf). * Knowledge of containerized orchestration for GPUs (e.g., Kubernetes with specialized device plugins). ## Description As an Automated Testing Engineer, you will be responsible for the end-to-end validation of large-scale, multi-node GPU clusters. You will help own the automated integration testing framework to validate high-performance GPU training, ensuring that distributed workloads scale efficiently across multiple virtualized nodes. Your role is critical in ensuring the stability of the low-level infrastructure and validating the interconnect fabric that powers the world's most demanding AI and HPC applications. San Francisco, Sunnyvale (Onsite) What You'll Be Working On: * CI/CD Automation and Tooling: Build CI/CD platforms that enable developers to quickly test, iterate, and deploy critical, low-level systems and applications. * Multi-Node Scaling Validation: Design and execute large-scale validation tests across multi-node virtualized clusters to ensure linear scaling and stability of GPU workloads. * Cluster Orchestration: Develop and maintain automation frameworks in Python or Go to dynamically provision, configure, and stress-test multi-node virtualized environments. * Interconnect & Fabric Testing: Validate high-speed interconnects-including NVLink, Infinity Fabric, InfiniBand, and RoCE-within virtualized environments to ensure low-latency, high-bandwidth communication. * Collective Communication Benchmarking: Architect and run comprehensive test suites using nccl-tests and rccl-tests (e.g., AllReduce, AllGather) to verify performance across node boundaries. * Performance Bottleneck Analysis: Perform deep-dive analysis of regressions in CPU performance and multi-node communication, identifying root causes across the guest OS, hypervisor, and physical fabric. * Create automated test suites leveraging tools like fio, stress-ng, and iperf to ensure performance and multi-tenant isolation of CPU and GPU hosts. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [13 AI Tools for Developers](https://www.wearedevelopers.com/magazine/302-13-ai-tools-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)