Ağız, Diş ve Çene Cerrahisinde Büyük Dil Modellerinin Yanıt Doğruluğu ve Bibliyografik Güvenilirliği: Chatgpt, Copilot ve Claude’un Karşılaştırmalı Analizi
NEÜ 4. Uluslararası Diş Hekimliği Kongresi, Konya, Türkiye, 1 - 04 Ekim 2026, ss.174-175, (Özet Bildiri)
- Yayın Türü: Bildiri / Özet Bildiri
- Basıldığı Şehir: Konya
- Basıldığı Ülke: Türkiye
- Sayfa Sayıları: ss.174-175
- İstanbul Medipol Üniversitesi Adresli: Evet
Özet
Amaç: Bu çalışmanın amacı, ChatGPT 5.5, Microsoft 365 Copilot ve Claude Sonnet 5 tarafından ağız, diş
ve çene cerrahisinin on temel alt alanındaki standart klinik sorulara verilen yanıtların doğruluğunu ve
önerilen referansların varlığını, bibliyografik hata profilini, konu uygunluğunu ve güncelliğini
karşılaştırmalı olarak değerlendirmektir.
Yöntem: On alt alanı kapsayan 40 standart klinik soru, aynı koşullar altında üç yapay zekâ sistemine
yöneltildi ve her soru için yalnızca ilk yanıt değerlendirildi. Klinik yanıtların doğruluğu, sistem
kimliğine kör deneyimli bir ağız, diş ve çene cerrahisi uzmanı tarafından beş puanlı Likert ölçeği
kullanılarak puanlandı. Her yanıt için üç referans talep edildi ve toplam 360 referans; varlık, altı
bibliyografik hata kategorisi, konu uygunluğu ve 2021–2026 yılları arasındaki güncellik açısından
incelendi. İstatistiksel analizlerde Kruskal–Wallis, Bonferroni düzeltmeli Mann–Whitney U, ki-kare,
Fisher’ın kesin testi ve Spearman korelasyon analizi kullanıldı.
Bulgular: Yanıt doğruluğu, referans varlığı, konu uygunluğu ve güncellik açısından sistemler arasında
anlamlı fark saptandı (tümü p<0,001). Yanıt doğruluğu medyanları ChatGPT, Copilot ve Claude için
sırasıyla 4, 2 ve 5 idi. Referans varlığı oranları sırasıyla %90,0, %79,2 ve %97,5; güncellik oranları ise
%77,5, %66,7 ve %99,2 olarak bulundu. Konu uygunluğu medyanı tüm sistemlerde 5 olmakla birlikte
puan dağılımları anlamlı biçimde farklıydı. İkili karşılaştırmalarda farklar en tutarlı biçimde Copilot ile
Claude arasında gözlendi. DOI/PMID hatası tüm sistemlerde en sık görülen bibliyografik hata
kategorisiydi. Yazar bilgisi hatası açısından sistemler arasında anlamlı fark bulunmadı (p=0,148).
Sonuç: Değerlendirilen yapay zekâ sistemleri klinik yanıt doğruluğu ve referans özellikleri açısından
farklı performans göstermiştir. Claude incelenen ölçütlerde daha yüksek sonuçlar göstermesine karşın,
sistemlerin hiçbiri bibliyografik hatalardan tamamen arınmış değildir. Yapay zekâ tarafından önerilen
referanslar klinik veya akademik kullanımdan önce bağımsız olarak doğrulanmalıdır.
Anahtar Kelimeler: Yapay Zeka, Ağız, Diş Ve Çene Cerrahisi, Büyük Dil Modelleri, Bibliyografik Hata, Kanıta
Dayalı Diş Hekimliği
Objective: This study aimed to comparatively evaluate the accuracy of responses generated by ChatGPT
5.5, Microsoft 365 Copilot, and Claude Sonnet 5 to standardized clinical questions across ten major
domains of oral and maxillofacial surgery, as well as the existence, bibliographic error profile, topic
relevance, and recency of the references provided.
Methods: Forty standardized clinical questions covering ten domains were submitted to the three
artificial intelligence systems under identical conditions, and only the first response to each question was
evaluated. The accuracy of the clinical responses was rated using a five-point Likert scale by an
experienced oral and maxillofacial surgeon blinded to system identity. Three references were requested
for each response, resulting in a total of 360 references, which were assessed for existence, six categories
of bibliographic errors, topic relevance, and recency within the 2021–2026 period. Statistical analyses
included the Kruskal–Wallis test, Bonferroni-adjusted Mann–Whitney U test, chi-square test, Fisher’s
exact test, and Spearman correlation analysis.Results: Significant differences were identified among the systems in response accuracy, reference
existence, topic relevance, and recency (all p < 0.001). Median response accuracy scores for ChatGPT,
Copilot, and Claude were 4, 2, and 5, respectively. Reference existence rates were 90.0%, 79.2%, and
97.5%, respectively, while recency rates were 77.5%, 66.7%, and 99.2%. Although the median topic
relevance score was 5 for all systems, the score distributions differed significantly. Pairwise comparisons
showed the most consistent differences between Copilot and Claude. DOI/PMID errors were the most
frequent category of bibliographic error across all systems. No significant difference was found among
the systems regarding errors in author information (p = 0.148).
Conclusion: The evaluated artificial intelligence systems demonstrated different levels of performance in
clinical response accuracy and reference characteristics. Although Claude showed higher results across
the evaluated measures, none of the systems was completely free of bibliographic errors. References
provided by artificial intelligence systems should therefore be independently verified before clinical or
academic use.
Keywords: Artificial Intelligence, Oral and Maxillofacial Surgery, Large Language Models, Bibliographic Errors,
Evidence-Based Dentistry