Stop Treating Your GPUs Like Cattle: Why Kubernetes Finally Got Hardware Right
Last year, I spent three weeks debugging a distributed training job that kept failing with cryptic out-of-memory errors on GPUs we should have had enough of. The problem? Kubernetes was treating our eight V100s like eight identical, interchangeable boxes. It didn't know about the...