Detect Representation Collapse in Self-Supervised Learning
Self-supervised learning can train an encoder without manually assigning a class label to every example. But removing labels also removes an obvious force that tells different examples to occupy meaningfully different parts of representation space. A badly designed objective can therefore admit a trivial solution: the encoder maps many or all inputs to essentially the same representation. This failure is called representation collapse. The training loss may even look good, because a model that emits the same vector for two augmented views of every input has achieved perfect agreement without learning useful distinctions.