Adapting I-JEPA for Low-Resolution Image Understanding: A Grid-Constrained Self-Supervised Learning Approach on CIFAR-10
Oluşturanlar
Katkıda Bulunan Kişiler
Araştırmacıs:
Açıklama
Self-supervised learning (SSL) methods have demonstrated remarkable success in learning visual representations without labeled data. The Image-based JointEmbedding Predictive Architecture (I-JEPA) predicts abstract latent representations rather than pixel-level reconstructions, but was designed exclusively for high-resolution datasets (e.g., ImageNet, 224×224). In this work, we present a systematic adaptation of I-JEPA to the low-resolution CIFAR-10 dataset (32×32), introducing a Grid-Constrained Masking strategy that preserves the core predictive learning mechanism on an 8×8 token grid. Our ViT-Tiny-based implementation (∼5.5M trainable parameters) achieves 62.14% linear probe accuracy after 50 epochs of pre-training—a 6.21× improvement over the random baseline—validating that latent-space prediction remains effective under severe spatial constraints.
Dosyalar
uluirmak_kurban_durmus_ICADA2026bildiri_4sayfa.pdf
Dosyalar
(1.3 MB)
| Ad | Boyut | Hepisini indir |
|---|---|---|
|
md5:554735336b46bd4223f22984f48f7a61
|
1.3 MB | Ön İzleme İndir |