Paper page - Invisible Shortcuts: Why Vision Encoders Know Your Camera
… We hypothesize that large-scale semantic supervision, whether through categorical labels ImageNet or billion-scale captions LAION , naturally induces metadata-semantics correlations during pretraining, leading models to convert low-level signals into predictive features. …