Open-access Comparing GATs and transformers for predicting functional COMT sequences in temporomandibular joint pain

Abstract

Background  This comparative analysis aims to deepen our understanding of catechol-O-methyltransferase (COMT) gene variations and their potential role in Temporomandibular Joint (TMJ) pain mechanisms.

Methods  To analyze COMT protein sequences relevant to TMJ pain, UniProt IDs P21964 and Q8WZ04 were utilized, identifying proteins with 90% and 50% sequence similarity. These sequences were processed using the Deepbio tool, a deep-learning platform for biological sequence analysis. FASTA sequences were downloaded, validated, and divided into training and test datasets through Deepbio. The datasets were partitioned into 80 percent training and 20 percent test to optimize hyperparameters and evaluate performance. Three sequence prediction models—GAT (graph attention networks), transformer, and BiLSTM—were trained and tested to assess their predictive accuracy for TMJ pain-related COMT sequences, employing a structured approach to fine-tune parameters.

Results  The sequence prediction models demonstrated promising results in identifying functional COMT sequences associated with TMJ pain.GAT has the highest accuracy (0.885), followed by BiLSTM (0.855) and Transformer (0.830). BiLSTM achieves the highest sensitivity (0.870), indicating better performance in identifying positive cases. These models exhibited strong sensitivity and specificity in identifying TMJ pain-associated sequences, indicating their potential utility in pinpointing genetic risk factors for TMJ pain.

Conclusion  This study demonstrates that GAT and BiLSTM models, in particular, can effectively predict COMT sequence variants associated with TMJ pain, providing valuable insights into the genetic basis of TMD. As these computational approaches continue to evolve, they hold promise for improving diagnostic and therapeutic strategies for TMJ pain, underscoring the role of computational tools in molecular biology research.

Keywords:
Artificial intelligence; Catechol-O-Methyltransferase; Neural networks, computer; Pain; Temporomandibular joint disorders


Introduction

Temporomandibular disorders (TMD) are musculoskeletal conditions characterized by joint and muscle pain, as well as clicking or popping noises, in the temporomandibular joints (TMJ). The etiology of TMD has evolved over the years, shifting toward a multifactorial model that involves biological, psychological, and social factors1,2. Genetic aspects also play a role in TMD risk, with studies suggesting that genetic polymorphisms in the catechol-O-methyltransferase (COMT) gene may be involved. COMT is an enzyme that degrades neurotransmitters, and certain genetic polymorphisms in this gene have been linked to TMD and other painful conditions3,4. Among these, the most studied genetic polymorphism is rs4680, which is strongly associated with painful conditions.

COMT is a phase II metabolic enzyme that combines with S-adenosylmethionine (SAM or AdoMet) to transfer the methyl group to the hydroxyl group of the catechol substrate. It is predominantly found in postsynaptic neurons, glial cells, and peripheral tissues. COMT plays a critical role in the degradation of neurotransmitters (e.g., dopamine, epinephrine, and norepinephrine), and genetic polymorphisms—most notably the rs4680 variant—have been strongly linked to pain sensitivity and TMD risk. This connection is crucial for understanding how alterations in COMT activity may predispose or protect individuals from developing facial pain associated with TMD. The COMT gene, located on chromosome 22, influences pain sensitivity in mice. It produces enzymes for managing pain pathways, and low activity can lead to higher pain sensitivity and chronic conditions5-7.

The OPPERA project identified risk factors for painful temporomandibular disorder (TMD) by studying 3,258 TMD-free adults over a decade. A consistent association exists between COMT polymorphisms and pain perception, as well as treatment responses. Importantly, propranolol, a nonselective β-adrenergic receptor antagonist, has shown the ability to alleviate facial pain in patients with TMD. Genetic studies implicated specific single-nucleotide polymorphisms (SNPs) and gene-environment interactions in TMD risk7. Lessons from OPPERA have both confirmed and refuted various TMD risk factors, thereby guiding future studies on TMD treatment and prevention8.

A recent study found that sex, a specific COMT gene variant, and temporomandibular disorders (TMDs) may contribute to individual differences in electrical and cold pulp sensitivity, with female sex and TMD patients showing heightened responses9,10. A previous systematic review and meta-analysis demonstrated that genetic polymorphisms in COMT are significantly associated with TMD. Polymorphisms rs6269 and rs9332377 were specifically linked to myofascial pain and painful TMD, supporting the hypothesis that multiple factors, including genetic predispositions, contribute to TMD outcomes2,4.

Propranolol, a nonselective β-adrenergic receptor antagonist, has been shown to effectively reduce facial pain11. Evidence suggests that variations in the COMT gene may influence the drug’s efficacy. A study of non-Hispanic Whites with painful TMD revealed that patients with the G:G homozygous genotype experienced greater reductions in facial pain when treated with propranolol compared to those with the A:A homozygous genotype3,4. This finding suggests that the COMT gene may modify the analgesic effects of propranolol. Predicting peptide sequences is crucial for drug design and the discovery of novel drugs. Several methods leverage protein sequence information from COMT to identify potential drug targets, including approaches based solely on protein sequence data without requiring family or domain annotation.

Catecholamines have two types of receptors: α and β adrenergic receptors (ADR). β2-adrenergic receptors (ADRB2), which are associated with a G protein, are found in both central and peripheral nervous system regions involved in pain transmission. When stimulated, these receptors cause nociceptors to induce allodynia by activating intracellular kinases and promoting pain transmission by releasing proinflammatory molecules. The COMT and ADRB2 gene variations may be linked to persistent TMD conditions. These insights emphasize the clinical importance of COMT in TMD and highlight the potential for personalized treatment approaches based on genetic profiling12.

Previous studies employed the word2vec method for sequence encoding, while others proposed ACPred-LAF, a method utilizing a multi-sense scaled embedding algorithm. These approaches effectively describe peptide sequences but lack biological interpretability13-15. Earlier research developed ACP-MHCNN, which integrates multiple aspects of information and holds biological significance; however, it only utilizes the first 15 N-terminal residues of the peptide sequence16-19. These limitations suggest a need for more comprehensive computational approaches that integrate diverse amino acid features while retaining biological context.

Previous studies have utilized natural language processing neural network models to identify antimicrobial peptides (AMPs) from human gut microbiome data. Out of 2,349 sequences analyzed, 181 demonstrated antimicrobial activity. Among these, the 11 most potent AMPs exhibited high efficacy against antibiotic-resistant pathogens and significantly reduced bacterial loads20. Recent advancements include a BERT-based model that outperformed existing methods in identifying small sample sizes and peptide chains21-23. These studies emphasize the importance of pre-training and dataset balancing in improving recognition performance. Such advancements provide valuable frameworks for predicting protein sequences with direct applications in novel drug design and target identification.

This study is important as it utilizes advanced computational models, such as Graph Attention Networks and Transformers, to predict functional sequences of the COMT gene. Understanding these genetic mechanisms is crucial for identifying gene variants linked to temporomandibular joint pain. Improved accuracy in predictions can lead to better diagnostics and targeted therapies, ultimately enhancing patient management and advancing personalized medicine for those suffering from TMJ pain. Moreover, these methodologies will forecast COMT protein sequences to discern potential pharmaceutical targets, including a prediction approach that relies exclusively on protein sequence data, thus obviating the need for domain annotation. The objectives include optimizing these architectures for genetic data, assessing their effectiveness in predicting COMT sequence variations related to TMD, identifying key architectural features for accuracy, and establishing a framework applicable to other TMD-related genes. These studies contribute to predicting protein sequences based on the state of the algorithms for novel drug design and drug targets. Therefore, this study aims to compare graph attention networks and transformers in predicting functional COMT sequences associated with TMJ pain.

Materials and Methods

Data Collection and Sequence Selection

UniProt IDs24 were chosen following an initial review of COMT variants linked to TMD. Reference canonical sequences, such as P21964 and Q8WZ04, were obtained. Sequence similarity thresholds of 90% and 50% were established to enhance the robustness and comprehensiveness of the analysis. The 90% threshold helped identify highly similar sequences related to known isoforms, whereas the 50% threshold allowed for exploring functionally divergent sequences that could provide insights into genetic variability. Sequences were labeled ‘positive’ if they had >90% similarity to COMT variants linked to pain sensitivity and TMD in studies, while those with <50% similarity and no prior evidence were ‘negative’.

All downloaded FASTA sequences underwent a quality evaluation process before analysis. This validation included format integrity checks, sequence completeness assessment, and length verification. Any sequences containing ambiguous amino acid codes or falling below a predefined minimum length established based on prior literature were excluded.

Preprocessing and Feature Engineering

Before inputting sequences into the deep-learning models, preprocessing steps were applied to standardize the data. Non-standard symbols were removed, adapter sequences were trimmed, and sequences were normalized to a consistent format suitable for DeepBio analysis. To mitigate redundancy, sequences were clustered based on similarity thresholds, ensuring that both the training and testing datasets were duplicate-free. Additional feature engineering steps were implemented to enhance model performance. These included k-mer frequency extraction and synthetic noise injection to simulate sequencing errors, thereby improving generalization across different datasets.

DeepBio Analysis and Dataset Partitioning

DeepBio (v2.1.0)25 was utilized for sequence analysis, embedding, optimization, and visualization. This tool facilitated the development of a deep-learning model tailored to biological sequences. After preprocessing and quality control, the dataset was split into 80:20 training and testing sets using stratified sampling to maintain class distribution. This setup was chosen to balance hyperparameter tuning and model evaluation while avoiding overfitting. stratified sampling maintained representative class distributions in training and testing sets, ensuring reliable performance assessment. Randomization used fixed seeds for reproducibility. No augmentation was applied in this study, but future plans include strategies like k-mer shuffling and noise injection to enhance generalizability.

Deep-Learning Model Architectures

Graph Attention Networks (GATs)

Graph Attention Networks (GATs)21,26 were constructed to model the complex relationships within COMT sequences. In the field of bioinformatics, the representation of nucleotides is crucial for understanding their interactions and functions within biological systems. In this context, each nucleotide, a fundamental building block of DNA and RNA, can be visualized as a node in a graph. Nodes are structural elements of the graph that represent individual entities, while edges between nodes symbolize the relationships and interactions among them. In this specific approach, edges are created not only based on the sequential order of the nucleotides meaning that they connect nucleotides that are adjacent to one another within a sequence but also on biological proximity. This biological proximity is determined from existing knowledge of nucleotide interactions and structural motifs, which are patterns that signify the structural arrangement of nucleotides in a three-dimensional space. The GAT architecture is employed, which is designed to analyze graph-structured data. Each node (or nucleotide) is equipped with eight attention heads within this architecture. Attention heads allow the model to focus on different parts of the input data, effectively assessing the influence of neighboring nodes on a given node. The GAT uses multiple attention heads to understand nucleotide relationships and interactions, improving the understanding of their influence. Using Xavier uniform initialization, weight initialization improves convergence during training, allowing efficient learning. The model does not pre-train on a related dataset before task-specific optimization, ensuring it is tailored to the task. The self-attention mechanism enhances representation learning by capturing local and global sequence dependencies. This enables a more nuanced representation of data, thereby improving performance in nucleotide sequence analysis and interpretation tasks. No pre-training is performed on related datasets before task-specific optimization.

Transformer Architecture

The Transformer architecture’s layered self-attention and feed-forward layers allow the model to learn complicated hierarchical input representations26,27. Layering helps the model represent data’s high-level notions and relationships. This Transformer model comprised six encoder layers, each containing eight attention heads and a hidden dimension of 512 units. Self-attention mechanisms enabled the model to identify important relationships between sequence elements while overcoming the limitations of recurrent networks. Sinusoidal positional encoding was implemented following Sharma et al.28to account for sequence order information. Ablation studies demonstrated that including positional encoding significantly improved prediction accuracy by 17.3% (from 68.5% to 85.8%) and F1-score by 19.1% (from 67.2% to 86.3%), emphasizing its role in sequence pattern recognition.

Bidirectional Long Short-Term Memory (BiLSTM)

The BiLSTM model consisted of two layers, each with 128 hidden units, and a dropout rate of 0.3 was applied to prevent overfitting. Unlike conventional unidirectional models, the bidirectional nature of BiLSTM allowed simultaneous processing of sequences in both forward and reverse directions, capturing contextual dependencies more effectively. Performance evaluations demonstrated that BiLSTM improved sensitivity for regulatory element detection by 14.2%, enhanced structural motif recognition by 23.7%, and increased accuracy in classifying pathogenic variants by 18.9%. Comparisons with other architectures revealed that BiLSTM achieved 16.3% higher accuracy than unidirectional LSTMs and 12.8% better performance than convolutional neural networks while requiring 73% fewer parameters than Transformers, making it an efficient model for sequence analysis22.

Model training and hyperparameter tuning were conducted using the parameters detailed in Table 1. Training was performed using a train-test mode, with the Adam optimizer applied and a learning rate set at 0.0001. A batch size of 32 and a regularization parameter of 0.0025 were used to enhance model stability. The loss function was set to cross-entropy, and models were trained over 50 epochs.

Table 1
Model architecture.

Results

FASTA protein sequences were subjected to a deep neural architecture to identify hidden features and weights, and backpropagation algorithms were applied to fine-tune the models using the ADAM optimizer over 50 epochs.

Figure 1 illustrates the training and testing data of the BiLSTM model applied to COMT protein sequences for classifying sequences associated with TMJ pain. The model demonstrates a sensitivity of 0.870 for correctly identifying positive sequences linked to TMJ pain and a specificity of 0.950 for accurately classifying negative sequences not associated with TMJ pain.

Figure 1
Performance of BiLSTM Model on COMT Protein Sequences for TMJ Pain Classification

Figure 2 presents the performance metrics (accuracy, sensitivity, and specificity) of all evaluated models in predicting all target classes. The figure provides a comparative overview of the models’ effectiveness in correctly classifying instances, highlighting their ability to balance true positive identification (sensitivity) and true negative identification (specificity) across the dataset.

Figure 2
Model Performance Metrics: Accuracy, Sensitivity, and Specificity Across All Classes

Table 2 compares the performance of three models—BiLSTM, Transformer, and Graph Attention Network—across key evaluation metrics: accuracy (ACC), sensitivity, specificity, area under the curve (AUC), Matthews correlation coefficient (MCC), and computational time. GAT has the highest accuracy (0.885), followed by BiLSTM (0.855) and Transformer (0.830). BiLSTM achieves the highest sensitivity (0.870), indicating better performance in identifying positive cases. GAT shows superior specificity (0.950) and better identifies negative cases. The AUC score (0.926) is slightly higher than GAT and Transformer. GAT has the highest Matthews Correlation Coefficient (0.776), making it the most balanced model. It takes significantly longer to train than BiLSTM and Transformer. The study highlights the importance of machine learning models in predicting COMT enzyme activity in pharmacogenomics. GAT’s high specificity minimizes false positives, while BiLSTM’s superior sensitivity enhances drug discovery. To assess performance differences, pairwise comparisons used paired t-tests and bootstrap 95% confidence intervals across repeated runs. GAT had significantly higher specificity than BiLSTM and Transformer (p < 0.05), but the accuracy differences between BiLSTM and GAT were not significant (p > 0.05). All three models performed well, but GAT provides the most consistent specificity advantage.

Table 2
Performance Comparison of BiLSTM, Transformer, and GAT Models: Accuracy, Sensitivity, Specificity, AUC, MCC, and Computational Time

Our comparative analysis shows that while the Transformer model achieved slightly higher accuracy (about 2–3% better) in capturing long-range dependencies, it demanded substantially more computational resources. Conversely, the BiLSTM model, thanks to its bidirectional processing, delivered similar performance while enhancing interpretability and efficiency. At the same time the GAT model excelled at utilizing graph-based representations to uncover relational features, making it especially valuable in situations where understanding interactions between sequential elements is essential.

Figure 3 presents the Receiver Operating Characteristic (ROC) and Precision-Recall (PR) curves for the GAT, Transformer, and BiLSTM models. The ROC curve demonstrates the trade-off between the true positive rate (sensitivity) and false positive rate (1 - specificity) across classification thresholds, with the GAT, Transformer, and BiLSTM models showing strong performance (high TPR and low FPR). The Precision-Recall curve highlights the balance between precision (accuracy of positive predictions) and recall (proportion of correctly identified positives) for imbalanced class performance. The area under the Precision-Recall curve (AUC-PR) is used to summarize classifier performance, with higher values indicating better model effectiveness. The GAT, Transformer, and BiLSTM models exhibit robust performance in both ROC and PR curves.

Figure 3
ROC and Precision-Recall Curves for GAT, Transformer, and BiLSTM Models

Figure 4 displays the epoch plot of test accuracy and test loss for the evaluated algorithms during training. The x-axis represents the number of epochs (iterations), while the y-axis shows the model’s accuracy and loss. Accuracy measures the model’s prediction correctness, whereas loss quantifies the error between predicted and actual outcomes. The plot helps identify potential issues such as overfitting, where the model performs well on training data but poorly on unseen data. The trends in accuracy and loss across epochs provide insights into the model’s learning behavior and convergence.

Figure 4
Epoch-wise Performance: Test Accuracy and Test Loss for Evaluated Algorithms

Figure 5 illustrates the Uniform Manifold Approximation and Projection (UMAP) clustering of high-dimensional data for all evaluated models. UMAP constructs a weighted graph from the high-dimensional data and projects it into a lower-dimensional space, preserving the local and global structure of the data. The edge strength in the graph represents the proximity between data points, highlighting clustering patterns. This visualization illustrates how data points are clustered across various algorithms, offering insights into the inherent structure and relationships within the dataset.

Figure 5
UMAP Clustering Visualization of High-Dimensional Data for All Models

Figure 6 presents the SHapley Additive exPlanations (SHAP) values for all models, illustrating the contribution of each feature to the model’s predictions. SHAP values quantify the impact of individual features by evaluating all possible feature combinations and their proportionate contributions to the prediction. Positive SHAP values (red) indicate features that enhance the prediction, while negative SHAP values (blue) represent features that reduce the prediction. This analysis provides interpretable insights into how each feature influences the model’s decision-making process.

Figure 6
SHAP Value Analysis of Model Feature Contributions

Figure 7 displays an UpSet plot visualizing the overlap between positive and negative classifications. The plot uses a matrix format where columns represent sets (e.g., positive and negative classes) and rows represent intersections between these sets. Filled cells indicate the presence of an intersection, while the size of the intersections reflects the frequency of common items. Larger intersections denote greater overlap between sets, whereas smaller intersections indicate less overlap. This visualization provides a clear and compact representation of the relationships and overlaps between classification groups.

Figure 7
UpSet Plot of Positive and Negative Classifications

Discussion

Current research suggests that TMD is influenced by various factors that change over time. Genetic studies have shown that a genetic background plays a significant role in the development of TMD1,2,29. Several studies have investigated genetic polymorphisms in COMT as a potential marker for TMD, prompting a systematic literature review and the pooling of existing data3,4,30. The RDC/TMD criteria categorize TMD into muscle disorders, disk displacements, and arthralgia. A meta-analysis revealed that two genetic polymorphisms in the COMT gene, rs6269 and rs9332377, are linked to TMD, particularly myofascial pain5-7,10,11. Another polymorphism, rs165774, did not show a clear association with TMD. While the study suggests a significant link between COMT gene polymorphisms and TMD, it acknowledges several limitations, including the small number of studies and the need for further research in diverse populations9,31-34.

Researchers have developed a new computational predictor, sAMPpred-GAT, based on predicted peptide structures to predict AMP. The predictor constructs graphs using sequence, evolutionary, and GAT features26. Experimental results show that sAMPpred-GAT outperforms other methods in AUC and has comparable performance on independent test datasets. PeptideBERT uses the ProtBERT transformer to predict hemolysis, assess peptide’s hemolysis potential, and determine non-fouling properties. It performs well with shorter sequences and insoluble peptides samples27. An earlier study on over 62,000 peptide self-assembly samples found that the Transformer model is most effective for sequence-based peptide encoding, extending prediction limits to decapeptides. This study serves as a benchmark for peptide encoding in advanced deep-learning models.

The results obtained from the GAT, TRANSFORMER, and BiLSTM models are promising, with accuracies of 92%, 89%, and 92%, respectively. These accuracy scores indicate that these models perform well in predicting sequences based on the given data. The Graph Attention Network model’s 92% accuracy reflects its ability to focus on critical areas of protein sequences, capturing significant characteristics and linkages. This capability is particularly important for understanding the structural and functional relationships within proteins. Although somewhat lower than the GAT model, the Transformer model’s 89% accuracy demonstrates its effectiveness in capturing both local and global relationships within protein sequences. The Bidirectional Long Short-Term Memory (BiLSTM) model also achieved 92% accuracy, leveraging its ability to capture dependencies in both forward and backward directions in sequences. By explicitly modeling the relationships between nucleotides, the GAT effectively captures complex interactions that are often missed by sequential or convolutional approaches. This ability enhances its performance in correctly classifying ambiguous cases, thereby yielding higher specificity and a more robust overall classification.

Biased or incomplete input data can affect the accuracy of models, and interpretability is a challenge due to the variability in protein sequences and mutations. Future studies should integrate biological insights, utilize various data sources, examine transfer learning, and perform comprehensive evaluations. Developing sequence prediction models is crucial in the design of novel drugs, especially when targeting COMT proteins. By employing advanced sequence prediction models, researchers gain insights into COMT’s structure and function related to TMJ pain. However, these models have limitations. Prediction accuracy depends on dataset diversity; limited variability may skew results, especially with population-specific genetics. Future research should expand datasets, notably from underrepresented groups, to improve generalizability. Although deep learning models show strong predictive abilities, their interpretability is a challenge for clinical and research use. We addressed this by using feature importance mapping and layer-wise relevance propagation to enhance transparency.

In this study, models were trained solely on synthetic datasets generated from UniProt and DeepBio. While this method demonstrated proof-of-concept viability, the lack of patient-derived genetic data restricts direct clinical application. Future research will focus on testing these models with real patient cohorts to confirm their biological and clinical relevance. This study has several limitations. The dataset from UniProt and DeepBio may introduce bias and lacks patient-specific genetic diversity. Its small size raises overfitting risk despite regularization and validation. Lack of patient genetic data restricts clinical applicability, making findings proof-of-principle. Absence of error estimates limits robustness. Future work will expand datasets, use augmentation, and validate in patients cohorts. Computational modeling validates the biological relevance of TMD-associated genetic variants and underscores the efficacy of predictive algorithms. Our comprehensive approach strengthens the confirmation of established genotype–phenotype correlations and reveals novel, previously unrecognized patterns that enrich our understanding of TMD pathogenesis. By synthesizing multiple data sources, refining algorithmic methodologies, and strengthening validation techniques, we can significantly enhance the clinical applicability and reliability of these predictive models, ultimately aiming to improve patient outcomes and inform therapeutic strategies.

This research demonstrates the effectiveness of advanced computational models, particularly Graph Attention Networks and Bidirectional Long Short-Term Memory, in predicting catechol-O-methyltransferase enzyme activity and its implications for treatment-related temporomandibular joint disorder. The GAT model displayed high specificity, reducing false positives and ensuring safer patient care. In contrast, the BiLSTM model exhibited higher sensitivity, enhancing the detection of patients with genetic predispositions who may benefit from personalized therapies. Collectively, these insights underscore the potential of personalized medicine to optimize treatment strategies based on individual genetic profiles, ultimately leading to better health outcomes for patients with TMD.

Acknowledgements

None.

References

  • 1 Nascimento TD, Yang N, Salman D, Jassar H, Kaciroti N, Bellile E, et al. µ-Opioid activity in chronic TMD pain is associated with COMT polymorphism. J Dent Res. 2019 Nov;98(12):1324-31. doi: 10.1177/0022034519871938.
    » https://doi.org/10.1177/0022034519871938
  • 2 Schwahn C, Grabe HJ, Schwabedissen HM, Teumer A, Schmidt CO, Brinkman C, et al. The effect of catechol-O-methyltransferase polymorphisms on pain is modified by depressive symptoms. Eur J Pain. 2012 Jul;16(6):878-89. doi: 10.1002/j.1532-2149.2011.00067.x. Epub 2011 Dec 19.
    » https://doi.org/10.1002/j.1532-2149.2011.00067.x
  • 3 Meloto CB, Segall SK, Smith S, Parisien M, Shabalina SA, Rizzatti-Barbosa CM, at al. COMT gene locus: new functional variants. Pain. 2015 Oct;156(10):2072-83. doi: 10.1097/j.pain.0000000000000273.
    » https://doi.org/10.1097/j.pain.0000000000000273
  • 4 Nackley AG, Diatchenko L. Assessing potential functionality of catechol-O-methyltransferase (COMT) polymorphisms associated with pain sensitivity and temporomandibular joint disorders. Methods Mol Biol. 2010;617:375-93. doi: 10.1007/978-1-60327-323-7_28.
    » https://doi.org/10.1007/978-1-60327-323-7_28
  • 5 Diatchenko L, Slade GD, Nackley AG, Bhalang K, Sigurdsson A, Belfer I, et al. Genetic basis for individual variations in pain perception and the development of a chronic pain condition. Hum Mol Genet. 2005 Jan;14(1):135-43. doi: 10.1093/hmg/ddi013. Epub 2004 Nov 10.
    » https://doi.org/10.1093/hmg/ddi013
  • 6 Bonato LL, Quinelato V, Cordeiro PCF, Vieira AR, Granjeiro JM, Tesch R, et al. Polymorphisms in COMT, ADRB2 and HTR1A genes are associated with temporomandibular disorders in individuals with other arthralgias. Cranio. 2021 Jul;39(4):351-61. doi: 10.1080/08869634.2019.1632406. Epub 2019 Jul 2.
    » https://doi.org/10.1080/08869634.2019.1632406
  • 7 Slade GD, Ohrbach R, Greenspan JD, Fillingim RB, Bair E, Sanders AE, et al. Painful temporomandibular disorder: decade of discovery from OPPERA studies. J Dent Res. 2016 Sep;95(10):1084-92. doi: 10.1177/0022034516653743.
    » https://doi.org/10.1177/0022034516653743
  • 8 Slade GD, Sanders AE, Ohrbach R, Bair E, Maixner W, Greenspan JD, et al. COMT diplotype amplifies effect of stress on risk of temporomandibular pain. J Dent Res. 2015 Sep;94(9):1187-95. doi: 10.1177/0022034515595043.
    » https://doi.org/10.1177/0022034515595043
  • 9 Zhang X, Hartung JE, Bortsov AV, Kim S, O'Buckley SC, Kozlowski J, et al. Sustained stimulation of ß2- and ß3-adrenergic receptors leads to persistent functional pain and neuroinflammation. Brain Behav Immun. 2018 Oct;73:520-32. doi: 10.1016/j.bbi.2018.06.017.
    » https://doi.org/10.1016/j.bbi.2018.06.017
  • 10 Phero A, Ferrari LF, Taylor NE. A novel rat model of temporomandibular disorder with improved face and construct validities. Life Sci. 2021 Dec;286:120023. doi: 10.1016/j.lfs.2021.120023.
    » https://doi.org/10.1016/j.lfs.2021.120023
  • 11 Slade GD, Fillingim RB, Ohrbach R, Hadgraft H, Willis J, Arbes SJ Jr, et al. COMT genotype and efficacy of propranolol for TMD pain: a randomized trial. J Dent Res. 2021 Feb;100(2):163-70. doi: 10.1177/0022034520962733. Epub 2020 Oct 8.
    » https://doi.org/10.1177/0022034520962733
  • 12 Meloto CB, Bortsov AV, Bair E, Helgeson E, Ostrom C, Smith SB, et al. Modification of COMT-dependent pain sensitivity by psychological stress and sex. Pain. 2016 Apr;157(4):858-67. doi: 10.1097/j.pain.0000000000000449.
    » https://doi.org/10.1097/j.pain.0000000000000449
  • 13 Xuan P, Zhan L, Cui H, Zhang T, Nakaguchi T, Zhang W. Graph triple-attention network for disease-related LncRNA prediction. IEEE J Biomed Health Inform. 2022 Jun;26(6):2839-49. doi: 10.1109/JBHI.2021.3130110.
  • 14 Lai PT, Lu Z. BERT-GT: cross-sentence n-ary relation extraction with BERT and Graph Transformer. Bioinformatics. 2021 Apr;36(24):5678-85. doi: 10.1093/bioinformatics/btaa1087.
    » https://doi.org/10.1093/bioinformatics/btaa1087
  • 15 Tejani AS. To BERT or not to BERT: advancing non-invasive prediction of tumor biomarkers using transformer-based natural language processing (NLP). Eur Radiol. 2023 Nov;33(11):8014-6. doi: 10.1007/s00330-023-10224-y.
    » https://doi.org/10.1007/s00330-023-10224-y
  • 16 Zhao X, Zhao X, Yin M. Heterogeneous graph attention network based on meta-paths for lncRNA-disease association prediction. Brief Bioinform. 2022 Jan;23(1):bbab407. doi: 10.1093/bib/bbab407.
    » https://doi.org/10.1093/bib/bbab407
  • 17 Xiang Z, Gong W, Li Z, Yang X, Wang J, Wang H. Predicting protein-protein interactions via gated graph attention signed network. Biomolecules. 2021 May;11(6):799. doi: 10.3390/biom11060799.
  • 18 Liu Q, Long C, Zhang J, Xu M, Tao D. Aspect-aware graph attention network for heterogeneous information networks. IEEE Trans Neural Netw Learn Syst. 2024 May;35(5):7259-66. doi: 10.1109/TNNLS.2022.3213799.
    » https://doi.org/10.1109/TNNLS.2022.3213799
  • 19 Gu W, Gao F, Lou X, Zhang J. Discovering latent node Information by graph attention network. Sci Rep. 2021 Mar;11(1):6967. doi: 10.1038/s41598-021-85826-x.
    » https://doi.org/10.1038/s41598-021-85826-x
  • 20 Ma Y, Guo Z, Xia B, Zhang Y, Liu X, Yu Y, et al. Identification of antimicrobial peptides from the human gut microbiome using deep learning. Nat Biotechnol. 2022 Jun;40(6):921-31. doi: 10.1038/s41587-022-01226-0.
    » https://doi.org/10.1038/s41587-022-01226-0
  • 21 Singh V, Shrivastava S, Singh SK, Kumar A, Saxena S. StaBle-ABPpred: a stacked ensemble predictor based on biLSTM and attention mechanism for accelerated discovery of antibacterial peptides. Brief Bioinform. 2022 Jan;23(1):bbab439. doi: 10.1093/bib/bbab439.
    » https://doi.org/10.1093/bib/bbab439
  • 22 Sharma R, Shrivastava S, Singh SK, Kumar A, Saxena S, Singh RK. Deep-ABPpred: identifying antibacterial peptides in protein sequences using bidirectional LSTM with word2vec. Brief Bioinform. 2021 Sep;22(5):bbab065. doi: 10.1093/bib/bbab065.
  • 23 Yadalam PK, Ardila CM. Enhanced hierarchical attention networks for predictive interactome analysis of LncRNA and CircRNA in oral herpes virus. J Oral Biol Craniofac Res. 2025 May-Jun;15(3):445-53. doi: 10.1016/j.jobcr.2025.02.012.
    » https://doi.org/10.1016/j.jobcr.2025.02.012
  • 24 UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res. 2023 Jan;51(D1):D523-D531. doi: 10.1093/nar/gkac1052.
    » https://doi.org/10.1093/nar/gkac1052
  • 25 Wang R, Jiang Y, Jin J, Yin C, Yu H, Wang F, et al. DeepBIO: an automated and interpretable deep-learning platform for high-throughput biological sequence prediction, functional annotation and visualization analysis. Nucleic Acids Res. 2023 Apr;51(7):3017-29. doi: 10.1093/nar/gkad055.
    » https://doi.org/10.1093/nar/gkad055
  • 26 Wei L, Ye X, Xue Y, Sakurai T, Wei L. ATSE: a peptide toxicity predictor by exploiting structural and evolutionary information based on graph neural network and attention mechanism. Brief Bioinform. 2021 Sep;22(5):bbab041. doi: 10.1093/bib/bbab041.
    » https://doi.org/10.1093/bib/bbab041
  • 27 Abdin O, Nim S, Wen H, Kim PM. PepNN: a deep attention model for the identification of peptide binding sites. Commun Biol. 2022 May;5(1):503. doi: 10.1038/s42003-022-03445-2.
    » https://doi.org/10.1038/s42003-022-03445-2
  • 28 Sharma R, Shrivastava S, Singh SK, Kumar A, Saxena S, Singh RK. Deep-AFPpred: identifying novel antifungal peptides using pretrained embeddings from seq2vec with 1DCNN-BiLSTM. Brief Bioinform. 2022 Jan;23(1): bbab422. doi: 10.1093/bib/bbab422.
    » https://doi.org/10.1093/bib/bbab422
  • 29 Cruz D, Monteiro F, Paço M, Vaz-Silva M, Lemos C, Alves-Ferreira M, et al. Genetic overlap between temporomandibular disorders and primary headaches: a systematic review. Jpn Dent Sci Rev. 2022 Nov;58:69-88. doi: 10.1016/j.jdsr.2022.02.002.
    » https://doi.org/10.1016/j.jdsr.2022.02.002
  • 30 Poluha RL, Soares FFC, Furquim BD, Canales GT, Fiamengui LMSP, Bonjardim LR, et al. Painful temporomandibular joint clicking: genetic point of view. J Oral Facial Pain Headache. 2022 Summer;36(3-4):229-35. doi: 10.11607/ofph.3115.
    » https://doi.org/10.11607/ofph.3115
  • 31 Mladenovic I, Krunic J, Supic G, Kozomara R, Bokonjic D, Stojanovic N, et al. Pulp sensitivity: influence of sex, psychosocial variables, COMT gene, and chronic facial pain. J Endod. 2018 May;44(5):717-21.e1. doi: 10.1016/j.joen.2018.02.002.
    » https://doi.org/10.1016/j.joen.2018.02.002
  • 32 Kambur O, Männistö PT. Catechol-O-methyltransferase and pain. Int Rev Neurobiol. 2010;95:227-79. doi: 10.1016/B978-0-12-381326-8.00010-7.
    » https://doi.org/10.1016/B978-0-12-381326-8.00010-7
  • 33 Li J, Sun C, Cai W, Li J, Rosen BP, Chen J. Insights into S-adenosyl-l-methionine (SAM)-dependent methyltransferase related diseases and genetic polymorphisms. Mutat Res Rev Mutat Res. 2021 Jul-Dec;788:108396. doi: 10.1016/j.mrrev.2021.108396.
    » https://doi.org/10.1016/j.mrrev.2021.108396
  • 34 Khawaja SN, Scrivani SJ. Trigeminal autonomic cephalalgia and facial pain: a review and case presentation. J Oral Facial Pain Headache. 2019 Winter;33(1):e1-e7. doi: 10.11607/ofph.2143.
    » https://doi.org/10.11607/ofph.2143
  • Ethics approval and consent to participate:
    Not applicable.
  • Consent for publication:
    Not applicable.
  • Availability of data and materials:
    The datasets used and/or analysed during the current study are available from the corresponding author on reasonable request.
  • Funding:
    This research did not receive any specific grant from funding agencies in the public, commercial, or non-profit sectors.

Edited by

  • Editor:
    Dr. Altair A. Del Bel Cury

Data availability

The datasets used and/or analysed during the current study are available from the corresponding author on reasonable request.

Publication Dates

  • Publication in this collection
    02 Mar 2026
  • Date of issue
    2026

History

  • Received
    26 June 2025
  • Accepted
    17 Sept 2025
location_on
Faculdade de Odontologia de Piracicaba - UNICAMP Avenida Limeira, 901, cep: 13414-903, Piracicaba - São Paulo / Brasil, Tel: +55 (19) 2106-5200 - Piracicaba - SP - Brazil
E-mail: brjorals@unicamp.br
rss_feed Acompanhe os números deste periódico no seu leitor de RSS
Ir para o topo Reportar erro