Comparative Security Evaluation of Open-Source Large Language Models Against Turkish-Language Adversarial Prompts
Oluşturanlar
- 1. College of Engineering and Architecture, University College Dublin, Dublin, Ireland
Açıklama
Abstract: The rapid proliferation of open-source large language models (LLMs) has raised pressing concerns regarding their robustness against adversarial and malicious prompting, particularly outside the English-language contexts in which most safety evaluations are conducted. This study addresses this gap by systematically evaluating the information-security behavior of three open-source LLMs — Llama-3.1-8B-Instruct, Mistral-7B-Instruct-v0.3, and Qwen2.5-7B-Instruct — using a custom-built Turkish adversarial prompt suite. The suite comprises 45 prompts evenly distributed across three attack categories: jailbreak and prompt-injection techniques, malicious code generation requests, and sensitive or personal information disclosure requests. Each model's responses were automatically classified using a rule-based Turkish refusal-detection method into REFUSED, COMPLIED, or PARTIALLY COMPLIED categories, from which an Attack Success Rate (ASR) was computed per model and category. Results reveal substantial variation in security robustness across models. Llama-3.1-8B-Instruct exhibited a near-total lack of resistance, with a 100% overall ASR across all three attack categories. Mistral-7B-Instruct-v0.3 followed closely with a 97.8% overall ASR, complying fully with sensitive-information and malicious-code requests and resisting only a small fraction of jailbreak attempts. In contrast, Qwen2.5-7B-Instruct demonstrated markedly stronger resistance, with an overall ASR of 51.1%, and was the only model to refuse the majority of malicious-code and sensitive-information requests. These findings suggest that safety alignment achieved in English does not reliably transfer to Turkish, exposing a critical vulnerability in the deployment of open-source LLMs in non-English-speaking markets, and highlight substantial cross-model variation in multilingual safety robustness. The study contributes a reusable Turkish red-teaming prompt dataset and an evaluation pipeline for future work in this area.
Keywords: Large Language Models, Information Security, LLM Red-Teaming, Jailbreak Attacks, Turkish Natural Language Processing, AI Safety Evaluation
Publication details: Conference abstract presented at the 8th International Conference on Engineering and Applied Natural Sciences (ICEANS 2026), 25-26 July 2026, Konya, Türkiye. Published in: Abstract Book of the 8th International Conference on Engineering and Applied Natural Sciences (ICEANS 2026), pp. 65. Publisher: All Sciences Academy. ISBN: 978-625-8839-44-9.
How to cite / Kaynak gösterimi
APA 7: Gülmez, B. (2026, July). Comparative Security Evaluation of Open-Source Large Language Models Against Turkish-Language Adversarial Prompts [Conference abstract]. In Abstract Book of the 8th International Conference on Engineering and Applied Natural Sciences (ICEANS 2026) (p. 65). All Sciences Academy. 8th International Conference on Engineering and Applied Natural Sciences (ICEANS 2026), Konya, Türkiye. https://doi.org/10.5281/zenodo.23036859
IEEE: B. Gülmez, "Comparative Security Evaluation of Open-Source Large Language Models Against Turkish-Language Adversarial Prompts," in Abstract Book of the 8th International Conference on Engineering and Applied Natural Sciences (ICEANS 2026), Konya, Türkiye, Jul. 2026, p. 65, ISBN: 978-625-8839-44-9, doi: 10.5281/zenodo.23036859.
Chicago: Gülmez, Burak. 2026. "Comparative Security Evaluation of Open-Source Large Language Models Against Turkish-Language Adversarial Prompts." In Abstract Book of the 8th International Conference on Engineering and Applied Natural Sciences (ICEANS 2026), 65. Konya, Türkiye: All Sciences Academy. https://doi.org/10.5281/zenodo.23036859.
MLA 9: Gülmez, Burak. "Comparative Security Evaluation of Open-Source Large Language Models Against Turkish-Language Adversarial Prompts." Abstract Book of the 8th International Conference on Engineering and Applied Natural Sciences (ICEANS 2026), All Sciences Academy, 2026, p. 65. https://doi.org/10.5281/zenodo.23036859.
Harvard: Gülmez, B. (2026) 'Comparative Security Evaluation of Open-Source Large Language Models Against Turkish-Language Adversarial Prompts', in Abstract Book of the 8th International Conference on Engineering and Applied Natural Sciences (ICEANS 2026). Konya, Türkiye: All Sciences Academy, pp. 65. doi: 10.5281/zenodo.23036859.
BibTeX:
@inproceedings{Gulmez2026comparative,
author = {Gülmez, Burak},
title = {Comparative Security Evaluation of Open-Source Large Language Models Against Turkish-Language Adversarial Prompts},
booktitle = {Abstract Book of the 8th International Conference on Engineering and Applied Natural Sciences (ICEANS 2026)},
year = {2026},
month = {jul},
pages = {65},
address = {Konya, Türkiye},
publisher = {All Sciences Academy},
isbn = {978-625-8839-44-9},
doi = {10.5281/zenodo.23036859},
url = {https://doi.org/10.5281/zenodo.23036859}
}
Notlar
Dosyalar
Gulmez_2026_Comparative_Security_Evaluation_of_Open_Source_Large_Language_Models_Against_Turkish_Langu.pdf
Dosyalar
(2.4 MB)
| Ad | Boyut | Hepisini indir |
|---|---|---|
|
md5:57fc6d8a289270e772500ad0bfb9af5f
|
2.4 MB | Ön İzleme İndir |
Ek detaylar
İlgili çalışmalar
- Verilen çalışmanın parçasıdır
- Kitap: 978-625--883944-9 (ISBN)