All published articles of this journal are available on ScienceDirect.
Utilizing Artificial Intelligence to Diagnose Temporomandibular Joint Osteoarthritis via Orthopantomographs
Abstract
Introduction/Background
Temporomandibular Joint Osteoarthritis (TMJOA) is a common temporomandibular disorder that substantially impairs quality of life. Although cone-beam computed tomography is considered the gold standard for detecting osseous changes, its routine use is limited. Orthopantomographs are widely available but have reduced sensitivity for early TMJOA. This study aimed to evaluate the clinical utility of an artificial intelligence model for TMJOA diagnosis using panoramic images and to compare its performance with expert assessment.
Materials and Methods
A total of 651 participants with clinical symptoms suggestive of TMJOA were included. An automated deep learning framework based on YOLOv11 was developed to extract regions of interest encompassing the mandibular condyle, articular fossa, and articular eminence to classify joints as normal or osteoarthritic. Model performance was assessed using five-fold cross-validation, and diagnostic metrics evaluated included accuracy, sensitivity, specificity, area under the curve, Cohen’s kappa, and McNemar’s test.
Results
The AI model achieved a mean accuracy of 0.76, sensitivity of 0.63, specificity of 0.82, and an AUC of 0.72. Compared to experts, the AI demonstrated higher accuracy and sensitivity, while experts retained higher specificity (0.91 compared to 0.82). The agreement between the AI and expert diagnoses was moderate.
Discussion
The AI model achieved moderate diagnostic performance, consistent with prior dental AI literature, and outperformed expert assessment in sensitivity (0.63 vs. 0.51). Conversely, the expert retained higher specificity (0.91), consistent with a conservative interpretive threshold shaped by clinical experience. This sensitivity-specificity trade-off mirrors broader patterns observed in AI-versus-human diagnostic comparisons. These findings support AI as a complementary screening tool, particularly where access to CBCT or specialist expertise is limited.
Conclusion
The proposed AI model improves sensitivity for TMJOA detection on panoramic radiographs and performs comparably to expert clinicians. These findings support the potential role of AI as a cost-effective screening and decision-support tool in routine dental practice, particularly where access to CBCT is limited.
1. INTRODUCTION
Temporomandibular Joint Osteoarthritis (TMJOA) is one of the most common disorders in the Temporomandibular Joint (TMJ) group, leading to many consequences, including joint pain, trismus, malocclusion, and even permanent deformity [1]. TMJOA is a complex condition characterized by the resorption of articular cartilage and the remodeling and ossification of subchondral bone. Over the world, TMJOA significantly affects the quality of life of 10 - 15% of the population [2].
The pathogenesis of TMJOA involves a complex interplay of mechanical, inflammatory, cellular, and tissue-level processes within the joint [3–5]. Although several contributing factors and signaling pathways have been identified, the underlying mechanisms of TMJOA remain incompletely understood. A central feature of TMJOA is abnormal remodeling of the subchondral bone in the mandibular condyle, driven by multiple interacting factors, including excessive mechanical loading, inflammatory responses, and degeneration of the articular disc [6]. In the early stage of the disease, increased osteoclastic bone resorption occurs, initiating degeneration of the overlying articular cartilage. This is followed by a prolonged and inadequate bone repair process, resulting in subchondral bone sclerosis and increased bone mineral density. Consequently, the osteochondral interface becomes thicker and stiffer, further compromising joint biomechanics and accelerating disease progression. Clinically, patients with TMJOA commonly present with restricted mouth opening, pain during mastication, joint sounds such as clicking or crepitus, mandibular deviation during mouth opening, headache, and otology-related symptoms [7–9].
Panoramic X-ray images are widely used in the evaluation of maxillofacial structures; however, this technique has certain limitations. The most notable limitation is its difficulty in detecting small changes on the surface of the temporomandibular joint due to being overlapped by nearby anatomical structures (Fig. 1). Cone-Beam Computed Tomography (CBCT) is considered the reference standard in detecting TMJOA due to its ability to display detailed changes in the cortical and subcortical bone layers [10]. Nevertheless, CBCT is still not the first choice for TMJOA examination in routine clinical settings due to its high radiation exposure and high cost [11]. In this context, the application of new methods to enhance the early and accurate diagnosis of TMJOA is essential.

Orthopantomograph results.
A promising approach is the use of Artificial Intelligence (AI). AI has shown broad potential in dentistry, from analyzing lateral cephalograms to localizing third molars and the inferior alveolar nerve, as well as diagnosing maxillary sinusitis, dental caries, and other conditions [11, 12]. A recent systematic review concluded that Artificial Intelligence (AI) has considerable potential for the imaging-based diagnosis of TMJOA, demonstrating a pooled sensitivity of approximately 80%, a specificity of up to 90%, and an area under the receiver operating characteristic curve as high as 0.92 [13]. Deep learning models employing transfer learning with ResNet architectures have also been applied to panoramic radiographs, achieving diagnostic sensitivities and specificities of approximately 76–79% [11]. More recently, researchers have expanded AI-based diagnostic approaches by integrating multimodal data rather than relying solely on panoramic radiographs. For example, Eunhye Choi and colleagues developed a diagnostic framework that combined orthopantomogram (OPG) image analysis with temporomandibular joint acoustic signals (joint sounds) recorded during mandibular movement, thereby enhancing the comprehensiveness of TMJOA assessment [14]. Most published studies have been conducted in single-center settings with relatively limited datasets, restricting the generalizability and clinical applicability of their findings. Consequently, further investigations are needed to establish their role in routine clinical practice.
Based on the current clinical applications and research background, this study aims to investigate the clinical utility of an AI diagnostic tool developed for TMJOA diagnosis from Orthopantomograph (OPG) images using algorithms that compare the AI diagnosis with that of an expert.
2. METHODOLOGY
2.1. Participants and Sample Size
Radiographic images of the patients with the following selection and exclusion criteria were reviewed.
Selection criteria:
- Visited the Department of Odonto and Stomatology, Hanoi Medical University Hospital with at least 1 of the following clinical symptoms: temporomandibular joint pain, crepitus, and trismus.
- Have panoramic images.
- Accepted to enroll in the study.
Exclusion criteria:
- Had a history of orthopedic surgery.
- Had a history of systemic diseases.
The sample size of the study uses the following formula to evaluate the sensitivity and specificity of a disease diagnosis method:

With:
nse: Sample size for sensitivity,
: standard normal deviate (α = 0.05 then Z = 1.96)
pse: Probability of sensitivity. Take pse = 80% [15].
w: Error of sensitivity and specificity.
pdis: Prevalence of the disease in the population. Take pdis = 0.15 [16].

In reality, the research sample size was 651 participants.
Figure 2 illustrates the step-by-step schema of the study.

Study schema.
2.2. Orthopantomogram Image Criteria and TMJOA Diagnosis Criteria
All panoramic radiographs were acquired with standardized imaging parameters: 12 mA tube current, 85 kV tube voltage, and a mean exposure time of 17.6 seconds. Acceptable image quality includes the following:
- Alveolar processes with all upper and lower teeth must be clearly represented.
- The image should show the entire mandible, including the TMJ.
- The dimensions of anatomical structures in vertical and horizontal planes must be symmetrical.
- Density image appearance should be uniform, free of air over the tongue, with the appearance of transparent tape (black) over the roots of upper teeth.
- The panoramic image should not have artifacts due to dentures or removable orthodontic appliances, glasses, earrings, and other jewelry on the patient.
Patients were diagnosed with TMJOA based solely on radiographic evidence of temporomandibular joint changes.
2.3. AI Model Development
An automated model was developed to extract regions of interest (ROI), including the mandibular condyle and surrounding anatomical structures, and to classify them into two categories: normal and osteoarthritis (OA), from each OPG image using a YOLOv11-based object detection framework. Once an image is input into the system, the YOLOv11 model performs inference to detect regions suspected of containing temporomandibular joint abnormalities. The results are presented visually as bounding boxes overlaid on the image, accompanied by predicted class labels and corresponding confidence scores (Fig. 3).

Results of ROI extraction.
YOLOv11 is a one-stage object detection framework designed for real-time applications by integrating localization and classification into a unified neural network [17]. Unlike two-stage approaches, it eliminates the need for a region proposal mechanism, allowing the model to infer object locations and categories directly from the input image. The detection process begins with preprocessing, where the image is resized and normalized to ensure consistency. During inference, post-processing techniques such as confidence filtering and non-maximum suppression are employed to remove redundant predictions and yield accurate final detections.
Unlike conventional approaches that rely on manual selection of regions of interest (ROIs), the detection-based architecture of YOLOv11 simultaneously localizes the mandibular condyle and performs classification directly on the detected region. This approach is particularly well suited to panoramic radiographs, where the mandibular condyle typically occupies only a small portion of the image and can easily be obscured by surrounding anatomical structures [18, 19]. In this context, whole-image analysis or manual ROI selection may reduce the signal-to-noise ratio and limit the model's ability to focus on diagnostically relevant features. By learning directly from automatically localized regions, YOLOv11 can more effectively capture morphological characteristics associated with TMJOA, including articular surface erosion, subchondral sclerosis, and relative changes in joint space morphology.
The YOLOv11 model was implemented using the PyTorch deep learning framework, enabling efficient model training, optimization, and inference through its dynamic computational graph and GPU acceleration support. All algorithms were programmed and trained on a PC with a GeForce GTX 1080 Ti GPU. Model training was performed using the Adam optimizer with a learning rate of 1.0 × 10−6. All the training data is divided into mini-batches, with a mini-batch size is set to 16 during training. Training was conducted for up to 1,000 epochs until convergence, with early stopping applied based on the validation loss to prevent overfitting.
Data augmentation techniques, including image rotation (±5 degrees), horizontal and vertical shifts (±10%), brightness adjustment (±10%), and contrast adjustment (±10%), were applied to enhance data diversity and improve model robustness.
2.4. Model and Statistical Analysis
Accuracy, precision, recall, and F1 score were calculated for model performance. Accuracy is defined as the ratio of correct predictions. Precision is the ratio of true positives to true positives and false positives. Recall is the ratio of true positives to true positives and false negatives. Finally, the F1 score is a harmonic mean of precision and recall: (2 × precision × recall)/(precision + recall). Accuracy, specificity, and sensitivity were calculated for diagnostic performance; Cohen’s kappa was used to estimate agreement in TMJOA diagnoses between experts and AI reads; and McNemar’s test was used to evaluate the significance of the difference. All p values < 0.05 were considered to be statistically significant.
2.5. Ethics Approval and Consent to Participate
The study was approved by the University Council - Hanoi Medical University (Approval No. 4571/QD-DHYHN – 06/10/2023). All procedures involving human participants were performed in accordance with the ethical standards of the institutional and/or research committees and the 1975 Declaration of Helsinki, as revised in 2013. Agreements for participation were requested, and informed consents were given to all participants.
3. RESULTS
The demographics of participants in the study are shown in Table 1. A total of 651 participants were included in the analysis, comprising 377 osteoarthritic (OA) joints. For k-fold validation, a five-fold cross-validation strategy (k = 5) was adopted to ensure robust model evaluation and to mitigate bias arising from limited data availability. The dataset is divided into 5 (k=5) equal-sized folds, where each fold is used once as a testing set while the remaining folds are used to train the machine learning model (Supplement 1).
| Characteristics | Normal | OA | ||||
|---|---|---|---|---|---|---|
| Female | Male | Total | Female | Male | Total | |
| Number of participants | 226 | 142 | 368 | 199 | 84 | 283 |
| Number of joints | 452 | 284 | 736 | 268 | 109 | 377 |
| Mean age | 38.32 | 33.19 | 36.34 | 33.44 | 31.68 | 32.92 |
| SD | 14.679 | 13.709 | 14.511 | 15.029 | 17.023 | 15.638 |
| 95% CI | 36.40 - 40.25 | 30.92 - 35.46 | 34.85 - 37.83 | 31.34 - 35.54 | 27.98 - 35.37 | 31.09 - 34.74 |
A total of 651 participants were included in the analysis, comprising 368 individuals in the normal group and 283 individuals who had TMJOA. In the normal group, females accounted for 226 participants (61.4%) and males for 142 participants (38.6%). The mean age was 36.34 years (SD = 14.51), with a 95% confidence interval (CI) of 34.85-37.83.
The model's performance is best in fold 2 (0.79), followed by fold 3 (0.78) and fold 1 (0.76) (Table 2). The recall result is found best in fold 2 (0.69), and the highest precision is shown in fold 3 (0.77) (Table 3). The average accuracy, precision, recall, and F1 score were 0.76, 0.73, 0.63, and 0.68, respectively (Table 3).
| Fold | Confusion Matrix | Model Performance | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Actual | Predicted | Precision | Recall | Accuracy | Weighted Average Precision | Weighted Average Recall | F1 score | ||
| Normal | OA | ||||||||
| 1 | Normal | 176 | 22 | 0.89 | 0.78 | 0.76 | 0.80 | 0.68 | 0.67 |
| OA | 37 | 27 | 0.58 | 0.42 | |||||
| 2 | Normal | 172 | 19 | 0.90 | 0.80 | 0.79 | 0.83 | 0.75 | 0.72 |
| OA | 42 | 27 | 0.58 | 0.61 | |||||
| 3 | Normal | 168 | 17 | 0.91 | 0.77 | 0.78 | 0.83 | 0.71 | 0.72 |
| OA | 28 | 47 | 0.63 | 0.57 | |||||
| 4 | Normal | 152 | 19 | 0.89 | 0.68 | 0.71 | 0.77 | 0.63 | 0.63 |
| OA | 42 | 47 | 0.47 | 0.49 | |||||
| 5 | Normal | 160 | 20 | 0.89 | 0.73 | 0.73 | 0.78 | 0.64 | 0.65 |
| OA | 37 | 43 | 0.54 | 0.43 | |||||
| Average | Normal | 165.6 | 19.4 | 0.90 | 0.75 | 0.76 | 0.80 | 0.69 | 0.68 |
| OA | 37.2 | 38.2 | 0.56 | 0.50 | |||||
| Fold | Recall | Precision | F1 Score | Accuracy | AUC |
|---|---|---|---|---|---|
| 1 | 0.60 | 0.73 | 0.67 | 0.76 | 0.71 |
| 2 | 0.69 | 0.76 | 0.72 | 0.79 | 0.75 |
| 3 | 0.67 | 0.77 | 0.72 | 0.78 | 0.76 |
| 4 | 0.59 | 0.68 | 0.63 | 0.71 | 0.69 |
| 5 | 0.58 | 0.71 | 0.65 | 0.73 | 0.70 |
| Average | 0.63 | 0.73 | 0.68 | 0.76 | 0.72 |
The AI in fold 2 has the highest accuracy (0.79) and sensitivity (0.69), but specificity peaks in fold 3, and Cohen’s kappa is also found highest in Fold 3 (Table 3).
The comparison of the diagnostic performance among folds is shown in Table 4, and a visual presentation is shown in Fig. (4). Table 4 illustrates the diagnostic performance of the model across five cross-validation folds. Accuracy ranged from 0.71 to 0.79, with consistently higher specificity compared with sensitivity across all folds. Cohen’s kappa values indicate moderate agreement, ranging from 0.401 to 0.589. The highest agreement is observed in Fold 3 (κ = 0.589), while the lowest is in Fold 4 (κ = 0.401). McNemar’s test demonstrates no statistically significant difference between paired classifications in most folds (p > 0.05), except for Fold 4, which shows a significant difference (p = 0.002). This suggests overall stability of model predictions across folds, with some variability in specific iterations. The visual comparison of sensitivities and specificities across folds is shown in Fig. (4).
| Diagnostic Performance | Cohen’s Kappa |
McNemar’s test (p-Value) |
|||
|---|---|---|---|---|---|
| Accuracy | Sensitivity | Specificity | |||
| Fold 1 | 0.76 | 0.60 | 0.82 | 0.487 | 0.666 |
| Fold 2 | 0.79 | 0.69 | 0.82 | 0.566 | 0.533 |
| Fold 3 | 0.78 | 0.67 | 0.84 | 0.589 | 0.117 |
| Fold 4 | 0.71 | 0.59 | 0.79 | 0.401 | 0.002 |
| Fold 5 | 0.73 | 0.58 | 0.82 | 0.466 | 0.058 |

Comparison of the sensitivities and specificities among folds.
The difference in diagnostic performance is shown in Table 5 and Fig. 5. Cohen’s kappa shows a substantial level of agreement for the expert (0.62) and moderate agreement for the AI (0.50) (Table 5). On average, the expert diagnostic performance is inferior compared to AI in terms of accuracy and sensitivity. However, the diagnosis of true negatives is better read by experts.
| Diagnostic Performance | Cohen’s Kappa | |||
|---|---|---|---|---|
| Accuracy | Sensitivity | Specificity | ||
| AI average | 0.76 | 0.63 | 0.82 | 0.502 |
| Expert | 0.69 | 0.51 | 0.91 | 0.621 |

Comparison of the sensitivities and specificities between AI and Expert.
The AI model demonstrates higher overall accuracy and greater sensitivity than the experts, indicating an improved ability to correctly identify osteoarthritis. In contrast, the experts achieve higher specificity, reflecting a stronger performance in correctly identifying normal joints. Regarding agreement with the reference standard, the experts show a higher Cohen’s kappa score than the AI model, corresponding to substantial agreement, whereas the AI exhibits moderate agreement. These findings suggest that while the AI model offers superior case detection and overall accuracy, expert assessment remains more conservative, resulting in fewer false-positive classifications.
The curves comparing the average performance of the AI system and expert assessment are shown in Fig. (5). Both approaches demonstrated good discriminatory ability, with sensitivity increasing as specificity decreased. The findings suggest that the AI system performs at least comparably to expert evaluation, with a potential advantage in sensitivity.
4. DISCUSSION
The study developed an artificial intelligence model based on the YOLOv11 framework for the automatic detection of Temporomandibular Joint Osteoarthritis (TMJOA) using orthopantomographs (OPG). To our knowledge, this is among the first studies in Vietnam to apply deep learning for this specific diagnostic challenge. The primary finding of this investigation is that the AI model demonstrated superior diagnostic accuracy and sensitivity compared to human expert assessment, although the expert maintained higher specificity.
4.1. Comparison with Previous Deep Learning Models
In this study, our deep learning model demonstrated moderate diagnostic performance in the interpretation of panoramic dental radiographs. While slightly lower than some previous benchmarks, these results fall within the range reported in the literature on AI for dental image analysis. One of the earliest deep learning investigations in dental radiography applied a deep neural transfer network (DeNTNet) for the detection of periodontal bone loss in panoramic radiographs [20]. This model, trained on over 12,000 images, achieved an F1 score of 0.75, indicating robust lesion detection and competitive performance relative to clinicians. The variation in performance across studies may be partly explained by design differences. Models like Mask R-CNN, which perform pixel-level segmentation, often provide more precise boundary delineation, enhancing performance on tasks requiring detailed localization compared to object detection-oriented networks that prioritize speed, such as YOLO variants [21]. A systematic review of deep learning in dental radiographs observed that while standard convolutional and segmentation networks generally yield high pooled diagnostic metrics, their efficacy remains highly sensitive to variations in dataset size, imaging modality, and clinical task [22].
4.2. Performance Comparison between AI and Expert
A prominent finding in our study was the marked difference in diagnostic performance between the AI model and expert clinicians. Expert clinicians exhibited high specificity (0.91) but notably low sensitivity (0.51), consistent with well-recognized limitations of OPGs in detecting subtle osseous changes due to anatomical superimposition and low contrast resolution. This discrepancy likely reflects well-documented limitations of OPGs, where overlapping anatomical structures, low resolution, and geometric distortion reduce the visibility of subtle bone changes, making early disease difficult to detect for human observers. These limitations make early condylar bone changes difficult to visualize, even for experienced interpreters, and echo findings from broader literature showing that standard panoramic imaging has limited reliability compared to volumetric modalities such as CBCT for TMJOA evaluation [23]. Previous research confirms that panoramic imaging often underperforms more advanced modalities such as CBCT for bone pathology, including condylar degeneration and other mandibular changes, largely due to these inherent imaging constraints. CBCT provides multiplanar, three-dimensional visualization without superimposition, resulting in superior diagnostic accuracy for changes that may be indiscernible on 2D OPGs [24, 25].
The study’s AI model increased sensitivity from 0.51 of experts to 0.63, identifying true positive cases that were missed by experts. This enhanced ability to detect subtle radiographic patterns mirrored results from previous systematic studies of AI applications in TMJOA detection. A recent meta-analysis of AI models applied to TMJOA radiographic datasets reported pooled sensitivity and specificity values around 0.80 and 0.79, when using both CBCT and panoramic imaging modalities, underscoring that deep learning systems can achieve meaningful diagnostic performance across variable imaging conditions [11]. Furthermore, this review highlighted that architectures such as ResNet exhibit moderate to high sensitivity when trained on adequately annotated condylar osteoarthritis datasets, supporting our finding that AI may overcome some of the human limitations inherent in two-dimensional radiographs. Numerous studies in dental imaging also highlighted the capacity of CNN-based models to extract non-linear, sub-visual features, enabling more sensitive detection of radiographic anomalies such as periapical lesions, bone loss, or other pathologies. In a recent large-scale evaluation of AI performance on panoramic radiographs, deep learning achieved moderate to high sensitivity and specificity for detecting dental conditions, often approaching or exceeding the performance of human readers in certain tasks [26, 27].
Several individual diagnostic accuracy studies further illustrate the performance landscape of AI in TMJOA detection. One study developed deep learning models for condylar osteoarthritis classification using panoramic TMJ projection images and found that models such as GoogLeNet and VGG-16 achieved AUCs up to 0.89, outperforming less experienced human observers, particularly on standardized TMJ projections [28]. Another study focused specifically on OPGs showed that a ResNet-based algorithm yielded sensitivity to 0.73 and specificity to 0.82 when trained and tested in a clinically relevant setting confirmed by CBCT, thus demonstrating performance comparable to experts while achieving a more balanced trade-off [15]. It is also notable that many approaches utilize advanced transfer learning and pre-trained architectures, such as ResNet-152 and EfficientNet-B7, had demonstrated that model selection and training strategy significantly influence diagnostic outcomes [29]. These findings support the notion that AI systems can extract and integrate subtle texture and morphological cues from low-dimensional imaging that may be below the threshold of visual detection by human readers.
Despite these advances, expert clinicians often retain higher specificity, which highlights a key benefit of human judgment in reducing false positives. In our cohort, experts achieved a specificity of 0.91, indicating a strong ability to correctly identify normal temporomandibular joints and avoid overdiagnosis. This finding aligns with multiple radiology and dental imaging studies reporting that experienced clinicians adopt a conservative interpretative threshold, particularly when evaluating panoramic radiographs for subtle degenerative changes. Several factors may explain this pattern.
First, clinicians rely heavily on well-established radiographic hallmarks of TMJOA, such as clear condylar erosion, flattening, osteophyte formation, or sclerosis, and are often reluctant to label early or ambiguous findings as pathological in the absence of unequivocal structural changes [30]. As a result, borderline radiographic variations, age-related remodeling, or projection artifacts are frequently classified as normal by human observers, thereby increasing specificity at the expense of sensitivity.
Second, expert judgment is often influenced by implicit clinical context, even when formal clinical data are not provided. Years of exposure to normal anatomical variability allow clinicians to recognize patterns that are unlikely to represent clinically meaningful TMJOA [31]. This experiential knowledge, while valuable for ruling out disease, may also contribute to underrecognition of early osteoarthritic changes, particularly when structural alterations are mild or atypical [32]. Consequently, expert readers may prioritize diagnostic certainty over early detection, reinforcing high specificity. Several investigations demonstrate that while AI models frequently achieve higher sensitivity, expert clinicians consistently outperform AI in confirming normal joints. This suggests that specificity may represent a strength of human expertise, rooted in holistic interpretation and risk-averse decision-making, whereas AI excels at identifying subtle deviations from learned patterns, regardless of their clinical salience.
From a clinical perspective, this trade-off has meaningful implications. In definitive diagnostic settings or specialist care, high specificity is essential to prevent misclassification and unnecessary escalation of care. However, in screening or primary care, the cost of missed TMJOA diagnoses, particularly in the early stages, may outweigh the drawbacks of increased false positives. In such scenarios, AI-assisted screening could complement expert interpretation by flagging suspicious cases for further review, while clinicians retain ultimate authority over diagnosis. Overall, given the clinical implications of missing an early TMJOA diagnosis, the improved sensitivity of AI models suggests a valuable role for AI as a decision-support tool, particularly in settings where access to CBCT or highly experienced specialists is limited.
Artificial intelligence has emerged as a promising approach in radiology, a field characterized by large volumes of relatively standardized imaging data. Although the current literature remains limited in both quantity and methodological rigor, existing evidence consistently suggests that AI can achieve diagnostic performance comparable to that of healthcare professionals [33–36]. Notably, several studies have reported improved model performance when the training set consisted of ROI images rather than entire radiographs [37]. However, in most published work, ROIs are manually delineated, which reduces the scalability and clinical practicality of AI-based workflows. To address this limitation, the present study implemented an object-detection algorithm to automatically extract ROIs from orthopantomogram (OPG) images.
Previous research using similar techniques has reported exceptionally high precision in mandibular condyle detection, with average precision values of 99.4% on the right side and 100% on the left side [38]. Importantly, the automatically generated ROIs in our study encompassed not only the mandibular condyle but also the articular fossa and articular eminence. This broader anatomical coverage is clinically relevant, as TMJOA is not confined to condylar changes alone. Degenerative alterations are known to involve additional joint components, including the fossa and the eminence, underscoring the need for AI diagnostic models to analyze the entire joint rather than focusing solely on the condyle [39]. Consistent with this rationale, large-scale reviews have reported pooled diagnostic performance with AUC values exceeding 0.90, underscoring the strong clinical potential of AI-based approaches, although variability related to imaging modality and dataset characteristics remains an important factor [13].
5. LIMITATIONS
This study has several limitations. First, the study’s ground truth was established based on OPG features. OPGs have inherent geometric distortions; therefore, some “false positives” identified by the AI might actually be true positives if validated against a CBCT gold standard, or vice versa. Second, the study was conducted at a single center in Vietnam. While the sample size (N=651) is quite robust, external validation on datasets from different OPG machines and different races is necessary to ensure the model generalizes well to other populations
CONCLUSION
Overall, the combination of the YOLOv11 model, a context-aware ROI selection strategy, and an appropriate data preprocessing pipeline resulted in a diagnostic model well adapted to the characteristics of panoramic radiographs for TMJOA detection. This integrated approach was designed not only to improve diagnostic performance but also to establish a practical and clinically applicable training framework with the potential to facilitate AI-assisted TMJOA diagnosis in routine clinical settings. Although further validation is required to confirm the model’s robustness, its performance on the data suggests it is a viable, cost-effective screening aid in general dental practice.
AUTHORS’ CONTRIBUTIONS
The authors confirm contributions to the paper as follows: T.M.N., T.C.N., H.T.T.L., P.T.T.N., H.M.B.: Study conception and design; N.M.T., Q.T.T.N.: Data collection; T.M.N., T.C.N., H.T.T.L., A.N.T.L., P.H.N.: Analysis and interpretation of results; T.M.N., H.T.P., H.T.D.: Draft manuscript. All authors reviewed the results and approved the final version of the manuscript.
LIST OF ABBREVIATIONS
| AI | = Artificial Intelligence |
| CBCT | = Cone-Beam Computed Tomograph |
| DeNTNet | = Deep Neural Transfer Network |
| OPG | = Orthopantomogram |
| ROI | = Region of Interest (ROI) |
| TMJOA | = Temporomandibular Joint Osteoarthritis |
| TMDs | = Temporomandibular Disorders |
ETHICS APPROVAL AND CONSENT TO PARTICIPATE
The study was approved by the University Council - Hanoi Medical University (Approval No. 4571/QD-DHYHN – 06/10/2023).
HUMAN AND ANIMAL RIGHTS
All procedures involving human participants were performed in accordance with the ethical standards of the institutional and/or research committees and the 1975 Declaration of Helsinki, as revised in 2013.
CONSENT FOR PUBLICATION
Agreements for participation were requested, and informed consents were obtained from all participants.
AVAILABILITY OF DATA AND MATERIALS
The data that support the findings of this study are available from the authors upon reasonable request.
ACKNOWLEDGEMENTS
Declared none.

