Yayınlanmış 24 Eylül 2026 | Sürüm v1
Rapor Açık

Klasik Kaynaklar Çalıştay Raporu: Envanterden Korpusa, OCR/HTR'dan Yapılandırılmış Metinlere Dijitalleştirme ve Yapay Zeka

Açıklama

Türkiye Yazma Eserler Kurumu Başkanlığı (TÜYEK) ile Marmara Üniversitesi Dijital Beşeri Bilimler Uygulama ve Araştırma Merkezi (DBB-M) iş birliğinde düzenlenen Klasik Kaynaklar Çalıştayları, 11 Haziran 2026 (Envanterden Korpus Oluşturmaya Dijitalleşme) ve 25 Haziran 2026 (Dijitalleştirme Süreçlerinden İşlenebilir Metinlere) tarihlerinde Rami Kütüphanesi'nde, yaklaşık seksen uzmanın katılımıyla gerçekleştirilmiştir. Çalıştaylar kapsamında kamu kurumları (TÜYEK, T.C. Cumhurbaşkanlığı Devlet Arşivleri Başkanlığı, Türk Tarih Kurumu, Türk Dil Kurumu, Cumhurbaşkanlığı Millet Kütüphanesi), üniversiteler (İstanbul Üniversitesi, Marmara Üniversitesi, İbn Haldun Üniversitesi, İstanbul Medeniyet Üniversitesi, FSM Vakıf Üniversitesi vs.), araştırma merkezleri (IRCICA, İSAM, Koç Üniversitesi Kütüphanesi) ve özel sektör girişimleri (Osmanlica.com, Wikilala, Muteferriqa, LexiQamus) ilk kez bu genişlikte aynı masada buluşmuştur. Birinci çalıştay kaynakların tespiti, kataloglanması ve korpuslara (derlemlere) dönüştürülmesini; ikinci çalıştay yüksek kaliteli dijitalleştirme, çözünürlük iyileştirme ve OCR/HTR teknolojileriyle metinlerin işlenebilir hâle getirilmesi konularını ele almıştır.

Bu rapor, iki çalıştayın müstakil raporlarını ve ortak raporu birleştirmekte ve özellikle somut önerileri tek bir eylem çerçevesinde sıralamaktadır.

Amaç, Kapsam ve Yöntem

İslâm ve Osmanlı medeniyetine ait yazılı miras, bir yandan nitelikli erişilebilirlik sorunlarıyla, diğer yandan doğal dil işleme ve yapay zekâ teknolojilerinin sunduğu tarihî bir fırsatla karşı karşıyadır. İki çalıştaydan oluşan “Klasik Kaynaklar Çalıştayları” dizisi, bu dönüşümün önündeki teknik, metodolojik ve kurumsal engelleri farklı kurum ve kuruluşlardan uzmanların yanı sıra araştırmacılar ile birlikte ele almak üzere tasarlanmıştır.

Çalıştaylar dizisinin kısa vadeli hedefi; alandaki mevcut durumun tespit edilmesi, güncel ihtiyaçların belirlenmesi ve nihai aşamada atılacak somut adımların tartışılmasıdır.

Uzun vadede ise dört hedefi bulunmaktadır: Farklı kurumların kataloglama, envanter ve dijitalleştirme tecrübelerini sistematik biçimde belgelemek; metadata, kataloglama ve OCR/HTR süreçleri için alanda benimsenebilecek standart önerileri geliştirmek; kütüphaneciler, araştırmacılar ve yazılımcılar arasında kalıcı bir iş birliği ağı kurmak ve iki çalıştayın ardından somut proje önerileri ile ortak altyapı girişimleri için zemin hazırlamak.

Birinci çalıştay, korpus oluştururken kaynakların sistematik tespiti, kataloglanması ve derlenmiş dijital korpuslara dönüştürülmesi süreçlerine odaklanmıştır. Program iki ana oturum hâlinde kurgulanmış; her oturum, sunumların ardından katılımcıların serbestçe söz aldığı forum bölümleriyle tamamlanmıştır. Bu format bilinçli bir tercihtir: Amaç, tek yönlü bilgi aktarımı değil, kurumların kendi tecrübelerini ve ihtiyaçlarını ortaya koyarak ortak sorun ve çözüm alanlarının tespit edilmesidir.

İkinci Çalıştay ise dijitalleştirme süreçlerinden işlenebilir metinlere (çözünürlük, süper çözünürlük, OCR/HTR) odaklanmıştır.

Abstract (English)

Organized in collaboration with the Turkish Manuscripts Institution (TÜYEK) and the Marmara University Center for Digital Humanities (DBB-M), the Classical Sources Workshops were held at Rami Library (Istanbul) on June 11, 2026 (From Inventory to Corpus Creation in Digitization) and June 25, 2026 (From Digitization Processes to Structured Texts), with the participation of approximately eighty experts. Within the scope of the workshops, public institutions (TÜYEK, the Presidency of State Archives of the Republic of Turkey, the Turkish Historical Society, the Turkish Language Association, and the Presidential National Library), universities (Istanbul University, Marmara University, Ibn Haldun University, Istanbul Medeniyet University, Fatih Sultan Mehmet Vakıf University, etc.), research centers (IRCICA, İSAM, and Koç University Library), and private sector initiatives (Osmanlica.com, Wikilala, Muteferriqa, and LexiQamus) came together around the same table at this scale for the first time. The first workshop focused on the identification, cataloging, and transformation of sources into corpora (collections); the second workshop addressed high-quality digitization, resolution enhancement, and the transition of texts into structured formats via OCR/HTR technologies.

This report combines the independent reports of both workshops and the general report outlining specific recommendations within a single action framework.

Aim, Scope, and Methodology

The written heritage belonging to Islamic and Ottoman civilization faces both challenges regarding qualified accessibility and a historic opportunity presented by natural language processing and artificial intelligence technologies. The "Classical Sources Workshops" series, consisting of two workshops, was designed to address the technical, methodological, and institutional obstacles to this transformation together with researchers as well as experts from various institutions and organizations.

The short-term goal of the workshop series is to assess the current situation in the field, determine current needs, and discuss concrete steps to be taken at the final stage.

In the long term, it pursues four main objectives: systematically documenting the cataloging, inventory, and digitization experiences of different institutions; developing standard recommendations that can be adopted in the field for metadata, cataloging, and OCR/HTR processes; establishing a permanent cooperation network among librarians, researchers, and software developers; and laying the groundwork for concrete project proposals and joint infrastructure initiatives following the two workshops.

The first workshop focused on the processes of systematic identification, cataloging, and compilation of sources into digital corpora while building a corpus. The program was structured into two main sessions; each session was concluded with open forum segments where participants freely took the floor following presentations. This format was a deliberate choice: the objective was not one-way information transfer, but rather the identification of common problem and solution areas by allowing institutions to articulate their own experiences and needs.

The Second Workshop focused on the transition from digitization processes to structured texts (resolution, super-resolution, OCR/HTR).

Dosyalar

Klasik_Kaynaklar_Calistay_Raporu_Toplu_Dagitim.pdf

Dosyalar (1.2 MB)