Artificial intelligence facing the diagnostic challenge of orofacial pain
The diagnosis of orofacial pain (OFP) constitutes a real clinical challenge due to the close intertwining of dental, musculoskeletal, and neurological symptoms. In the practice, confusion between a classic odontogenic pathology and a temporomandibular disorder (TMD) or neuralgia can lead to therapeutic orientation errors, often cited among the most frequent adverse clinical events.
This systematic review, conducted according to PRISMA-DTA guidelines, specifically aims to evaluate the use, diagnostic accuracy, and clinical utility of artificial intelligence (AI) models — including machine learning and deep learning — for the classification and prediction of TMD conditions. The authors compiled data from 20 studies published between 2000 and 2026 to determine whether these automated tools can secure decision-making when faced with complex clinical presentations.
The central hypothesis rests on the ability of algorithms to outperform or assist conventional subjective assessment. The study tests the efficacy of AI in extracting reliable diagnostic patterns from multimodal sources (imaging, thermography, neuroimaging, questionnaires), specifically aiming to improve the localization of dental pain and the prediction of odontogenic postoperative pain.
Methodology of the systematic review
This systematic review was conducted according to the PRISMA-DTA guidelines and registered on PROSPERO (CRD420261322606). The authors synthesised data from 20 studies selected after an exhaustive electronic search in six databases (PubMed, Scopus, Embase, Cochrane Library, Web of Science, and Google Scholar) covering the period from 1 January 2000 to 1 February 2026.
The analysis was structured around the following PICO framework:
- Population: Patients presenting with orofacial pain (OFP), including temporomandibular disorders (TMD), trigeminal neuralgia, neuropathic pain, and postoperative odontogenic pain.
- Interventions: Artificial intelligence-based systems, including Machine Learning (ML), Deep Learning (DL), artificial neural networks (ANN) and convolutional neural networks (CNN), as well as language models (LLM).
- Comparison: Conventional clinical diagnostics performed by experts and traditional statistical prediction models.
- Evaluation criteria: Diagnostic accuracy, sensitivity, specificity, and area under the curve (AUC).
The methodological quality assessment was performed using the QUADAS-2 tool, and the certainty of evidence was graded via the GRADE approach. Due to the substantial heterogeneity of datasets, AI architectures, and outcome measures, the authors produced a qualitative synthesis.
Diagnostic and predictive performance of AI models
This systematic review, including 20 studies published between 2000 and 2026, reports significant performance variability depending on the data modalities used. The authors indicate that Machine Learning (ML) models and artificial neural networks (ANN) show overall diagnostic accuracies ranging between 75% and 99%.
| Modalité de données | Performance metrics | Clinical application |
|---|---|---|
| Thermography | 99% Precision | Diagnosis of orofacial pain (OFP) |
| MRI | AUC up to 0.899 | Temporomandibular disorders (TMD) |
| Clinical data & Questionnaires | Precision 75 % - 99 % | Classification and localization of pain |
Pour vous équiper
Produits Delynov en lien avec cette thématique :
- x25 sacs hydrosolubles 50 litres - Delynov (Hydro) - Delynov (dispositif de chirurgie dentaire)
- Kit chirurgical Protect par 1 Carton de 5 pièces (kits stériles) - Hygitech - Delynov (dispositif de chirurgie dentaire)
Delynov Chirurgie, votre fournisseur en fils de suture chirurgicale résorbables et non résorbables, consommables et instruments de chirurgie dentaire et implantaire.
The results highlight increased model efficiency when processing imaging datasets or structured clinical records. Convolutional Neural Networks (CNNs) and ANNs have proven particularly effective for interpreting radiographic imaging and neuroimaging.
Methodological quality and limitations of evidence
The quality assessment using the QUADAS-2 tool and the GRADE approach reveals the following points:
- Risk of bias: Judged low for patient selection and index test in the majority of included studies.
- Applicability: Major concerns have been raised regarding the actual clinical relevance and the lack of external validation of the models on independent populations.
- Certainty of evidence: Classified as moderate by the review authors.
The substantial heterogeneity of AI architectures, validation protocols, and outcome measures made any quantitative meta-analysis impossible. Qualitative observations suggest that, while promising, current AI systems still rely heavily on retrospective study designs, limiting their immediate generalizability to the dental practice.
Analysis: AI facing the puzzle of orofacial pain
This systematic review, compiling data from 20 studies published between 2000 and 2026, confirms that artificial intelligence (AI) is no longer a mere theoretical promise. With diagnostic accuracies ranging from 75% to 99%, Machine Learning (ML) models and artificial neural networks are transforming heterogeneous data into clinical decision-making tools. The synthesis highlights a remarkable performance of ML-assisted thermography (99% accuracy) and a high discriminatory capacity of MRI imaging for temporomandibular disorders (AUC up to 0.899).
But what do these figures mean for the practitioner? Clinically, AI excels where humans struggle: identifying subtle patterns in clinical presentations where symptoms of neuralgia, odontalgia, and TMJ disorders overlap. However, the synthesis highlights a major limitation: the heterogeneity of AI architectures and the lack of external validation. Most models have been tested on retrospective databases. The risk of bias remains low for patient selection, but the authors warn about real-world applicability in the dental surgery until prospective multicenter studies validate these algorithms on diverse populations.
Summary of results
This systematic review, synthesizing 20 studies, reveals that AI models achieve diagnostic accuracy ranging from 75% to 99% for orofacial pain. Thermography-based algorithms show the highest performance (99%), while AI-driven MRI analysis offers excellent discrimination capability for temporomandibular disorders (AUC up to 0.899).
In concrete terms, for the practitioner:
- Secure your complex diagnoses: Integrate AI as a second opinion for differential diagnosis between odontogenic pain, neuralgia and TMJ disorders, where symptoms frequently overlap.
- Anticipate postoperative pain: Use predictive models to identify, from the consultation stage, patients at risk of painful complications in order to adjust your preventive analgesic protocol.
- Maintain your critical thinking: These tools are effective on structured datasets (imaging, questionnaires), but the lack of robust external validation requires clinical expertise to be maintained as the final arbiter.
Source
- Original title: Application, performance and limitations of artificial intelligence for the diagnosis and prediction of oro-facial pain: a systematic review
- Authors: Sanjeev B. Khanagar, Areej Alfaifi, Oinam Gokulchandra Singh, Afnan Alayyash, Fayaz Ul Haq, Barrak Alsomaie, Sudarshan Bhat, Vineet Khinda, Satish Vishwanathaiah, Kiran Iyer
- Publication: Frontiers in Oral Health - 2026-08-05
- DOI: https://doi.org/10.3389/froh.2026.1895438
À lire aussi dans le blog Delynov
Intelligence artificielle et bitewings : vers une automatisation du diagnostic parodontal selon la classification AAP/EFP
IA YOLO en dentisterie : quand le deep learning optimise le diagnostic et la détection d'objets
Information intended for healthcare professionals. This content may contain errors or truncated summaries. We recommend always verifying with the original source article. Delynov disclaims all responsibility regarding the use of this information. This document is not intended for patients or the general public.