Ağız, Diş ve Çene Cerrahisinde Büyük Dil Modellerinin Yanıt Doğruluğu ve Bibliyografik Güvenilirliği: Chatgpt, Copilot ve Claude’un Karşılaştırmalı Analizi


Malkoç Okumuş Y., Güzel C.

NEÜ 4. Uluslararası Diş Hekimliği Kongresi, Konya, Türkiye, 1 - 04 Ekim 2026, ss.174-175, (Özet Bildiri)

  • Yayın Türü: Bildiri / Özet Bildiri
  • Basıldığı Şehir: Konya
  • Basıldığı Ülke: Türkiye
  • Sayfa Sayıları: ss.174-175
  • İstanbul Medipol Üniversitesi Adresli: Evet

Özet

Amaç: Bu çalışmanın amacı, ChatGPT 5.5, Microsoft 365 Copilot ve Claude Sonnet 5 tarafından ağız, diş

ve çene cerrahisinin on temel alt alanındaki standart klinik sorulara verilen yanıtların doğruluğunu ve

önerilen referansların varlığını, bibliyografik hata profilini, konu uygunluğunu ve güncelliğini

karşılaştırmalı olarak değerlendirmektir.

Yöntem: On alt alanı kapsayan 40 standart klinik soru, aynı koşullar altında üç yapay zekâ sistemine

yöneltildi ve her soru için yalnızca ilk yanıt değerlendirildi. Klinik yanıtların doğruluğu, sistem

kimliğine kör deneyimli bir ağız, diş ve çene cerrahisi uzmanı tarafından beş puanlı Likert ölçeği

kullanılarak puanlandı. Her yanıt için üç referans talep edildi ve toplam 360 referans; varlık, altı

bibliyografik hata kategorisi, konu uygunluğu ve 2021–2026 yılları arasındaki güncellik açısından

incelendi. İstatistiksel analizlerde Kruskal–Wallis, Bonferroni düzeltmeli Mann–Whitney U, ki-kare,

Fisher’ın kesin testi ve Spearman korelasyon analizi kullanıldı.

Bulgular: Yanıt doğruluğu, referans varlığı, konu uygunluğu ve güncellik açısından sistemler arasında

anlamlı fark saptandı (tümü p<0,001). Yanıt doğruluğu medyanları ChatGPT, Copilot ve Claude için

sırasıyla 4, 2 ve 5 idi. Referans varlığı oranları sırasıyla %90,0, %79,2 ve %97,5; güncellik oranları ise

%77,5, %66,7 ve %99,2 olarak bulundu. Konu uygunluğu medyanı tüm sistemlerde 5 olmakla birlikte

puan dağılımları anlamlı biçimde farklıydı. İkili karşılaştırmalarda farklar en tutarlı biçimde Copilot ile

Claude arasında gözlendi. DOI/PMID hatası tüm sistemlerde en sık görülen bibliyografik hata

kategorisiydi. Yazar bilgisi hatası açısından sistemler arasında anlamlı fark bulunmadı (p=0,148).

Sonuç: Değerlendirilen yapay zekâ sistemleri klinik yanıt doğruluğu ve referans özellikleri açısından

farklı performans göstermiştir. Claude incelenen ölçütlerde daha yüksek sonuçlar göstermesine karşın,

sistemlerin hiçbiri bibliyografik hatalardan tamamen arınmış değildir. Yapay zekâ tarafından önerilen

referanslar klinik veya akademik kullanımdan önce bağımsız olarak doğrulanmalıdır.

Anahtar Kelimeler: Yapay Zeka, Ağız, Diş Ve Çene Cerrahisi, Büyük Dil Modelleri, Bibliyografik Hata, Kanıta

Dayalı Diş Hekimliği

Objective: This study aimed to comparatively evaluate the accuracy of responses generated by ChatGPT

5.5, Microsoft 365 Copilot, and Claude Sonnet 5 to standardized clinical questions across ten major

domains of oral and maxillofacial surgery, as well as the existence, bibliographic error profile, topic

relevance, and recency of the references provided.

Methods: Forty standardized clinical questions covering ten domains were submitted to the three

artificial intelligence systems under identical conditions, and only the first response to each question was

evaluated. The accuracy of the clinical responses was rated using a five-point Likert scale by an

experienced oral and maxillofacial surgeon blinded to system identity. Three references were requested

for each response, resulting in a total of 360 references, which were assessed for existence, six categories

of bibliographic errors, topic relevance, and recency within the 2021–2026 period. Statistical analyses

included the Kruskal–Wallis test, Bonferroni-adjusted Mann–Whitney U test, chi-square test, Fisher’s

exact test, and Spearman correlation analysis.Results: Significant differences were identified among the systems in response accuracy, reference

existence, topic relevance, and recency (all p < 0.001). Median response accuracy scores for ChatGPT,

Copilot, and Claude were 4, 2, and 5, respectively. Reference existence rates were 90.0%, 79.2%, and

97.5%, respectively, while recency rates were 77.5%, 66.7%, and 99.2%. Although the median topic

relevance score was 5 for all systems, the score distributions differed significantly. Pairwise comparisons

showed the most consistent differences between Copilot and Claude. DOI/PMID errors were the most

frequent category of bibliographic error across all systems. No significant difference was found among

the systems regarding errors in author information (p = 0.148).

Conclusion: The evaluated artificial intelligence systems demonstrated different levels of performance in

clinical response accuracy and reference characteristics. Although Claude showed higher results across

the evaluated measures, none of the systems was completely free of bibliographic errors. References

provided by artificial intelligence systems should therefore be independently verified before clinical or

academic use.

Keywords: Artificial Intelligence, Oral and Maxillofacial Surgery, Large Language Models, Bibliographic Errors,

Evidence-Based Dentistry