Yayınlanmış 20 Nisan 2026 | Sürüm v1
Konferans bildirisi Açık

Adapting I-JEPA for Low-Resolution Image Understanding: A Grid-Constrained Self-Supervised Learning Approach on CIFAR-10

Açıklama

Self-supervised learning (SSL) methods have demonstrated remarkable success in learning visual representations without labeled data. The Image-based JointEmbedding Predictive Architecture (I-JEPA) predicts abstract latent representations rather than pixel-level reconstructions, but was designed exclusively for high-resolution datasets (e.g., ImageNet, 224×224). In this work, we present a systematic adaptation of I-JEPA to the low-resolution CIFAR-10 dataset (32×32), introducing a Grid-Constrained Masking strategy that preserves the core predictive learning mechanism on an 8×8 token grid. Our ViT-Tiny-based implementation (5.5M trainable parameters) achieves 62.14% linear probe accuracy after 50 epochs of pre-traininga 6.21× improvement over the random baselinevalidating that latent-space prediction remains effective under severe spatial constraints.

Dosyalar

uluirmak_kurban_durmus_ICADA2026bildiri_4sayfa.pdf

Dosyalar (1.3 MB)