Abstract
Deinococcus gobiensis is a major source of protease, which has several applications in biotechnology. The 3D structure of DgokerA significantly impacts its efficacy. In this study, we evaluated DgokerA from Deinococcus gobiensis I-0 (strain DSM 21396) using bioinformatic methods. The results revealed that the DgokerA sequence from Deinococcus gobiensis I-0 contains highly conserved amino acids, including arginine (R), aspartic acid (D), and glutamic acid (E). DgokerA cleaves the bonds between hydrophobic and polar amino acids in keratin-associated proteins. The effects of point mutations on the crystal structure of DgokerA showed that mutations leading to changes of asparagine D171 and N260 to isoleucine, valine, and isoleucine, respectively, were the most stabilizing. Genomic analysis of the DgokerA-coding gene (DgokerA) indicated that this gene and the malL gene are expressed through the same promoter. Analysis of transcription factor binding sites (TFBSs) revealed AbrB, PucR, Spo0A, Hpr, SigA, and ComA as relevant regulators of DgokerA expression. Phylogenetic analysis showed that DgokerA has highly conserved structures among Deinococcus species, sharing a common ancestor. Their coding genes were duplicated and evolved within the Deinococcus genus. In conclusion, our in silico study demonstrated that the overexpression of the DgokerA gene and the stability of the produced DgokerA can be improved through systems biology methods, such as point mutations and identifying the involved transcription factors (TFs) or TFBSs.
Keywords:
keratinase; DgokerA; Deinococcus gobiensis I-0; bioinformatics; structural analysis
Resumo
Deinococcus gobiensis é uma importante fonte de protease, com diversas aplicações em biotecnologia. A estrutura tridimensional da DgokerA impacta significativamente sua eficácia catalítica. Neste estudo, avaliamos a DgokerA de Deinococcus gobiensis I-0 (cepa DSM 21396) utilizando métodos bioinformáticos. Os resultados revelaram que a sequência de DgokerA de Deinococcus gobiensis I-0 contém aminoácidos altamente conservados, incluindo arginina (R), ácido aspártico (D) e ácido glutâmico (E). A DgokerA cliva as ligações entre aminoácidos hidrofóbicos e polares em proteínas associadas à queratina. Os efeitos de mutações pontuais na estrutura cristalina de DgokerA mostraram que as mutações que levam à substituição dos resíduos de asparagina D171 e N260 por isoleucina, valina e isoleucina, respectivamente, foram as que conferiram maior efeito estabilizador. A análise genômica do gene que codifica DgokerA indicou que este gene e o gene malL são expressos a partir do mesmo promotor. A análise dos sítios de ligação de fatores de transcrição (TFBSs) revelou AbrB, PucR, Spo0A, Hpr, SigA e ComA como reguladores relevantes da expressão de DgokerA. A análise filogenética mostrou que DgokerA possui estruturas altamente conservadas entre as espécies de Deinococcus, compartilhando um ancestral comum. Seus genes codificadores foram duplicados e evoluíram dentro do gênero Deinococcus. Em conclusão, nosso estudo in silico demonstrou que a superexpressão do gene DgokerA e a estabilidade do DgokerA produzido podem ser melhoradas através de métodos de biologia de sistemas, como mutações pontuais e identificação dos fatores de transcrição (TFs) envolvidos ou de seus respectivos TFBSs.
Palavras-chave:
queratinase; DgokerA; Deinococcus gobiensis I-0; bioinformática; análise estrutural
1. Introduction
Nature has bestowed countless blessings upon humanity in the form of microbes. Extremophilic microorganisms can survive under harsh conditions due to their specific molecular physiology. The enzymes produced by extremophiles are known as extremozymes, which possess unique properties that enable them to withstand high temperatures, salt concentrations, and pressures (Dumorné et al., 2017). Amylases, proteases, lipases, pullulanases, cellulases, chitinases, xylanases, pectinases, isomerases, esterases, and dehydrogenases offer significant advantages in various industries, including agriculture, chemicals, biomedicine, pharmaceuticals, cosmetics, food, detergents, biomaterials, and biotechnology (Dumorné et al., 2017; Marques, 2018; Wang et al., 2019). According to a recent survey by Mordor Intelligence, the proteases market is projected to grow at a Compound Annual Growth Rate (CAGR) of about 5.8% during the period from 2022 to 2027 (Protease Market - Size, Share & Analysis).
Keratinase (DgokerA) belongs to the protease family and is derived from Deinococcus gobiensis I-0, found in the Gobi Desert. The major sources of keratinases include bacterial genera such as Bacillus, Vibrio, Chryseobacterium, Brevibacillus, Pseudomonas, Serratia, Fervidobacterium, Microbacterium, Aeromonas, Burkholderia, Stenotrophomonas, Rhodococcus, Geobacillus, Amycolatopsis, Meiothermus, and Paenibacillus (Fang, Sha, Peng, Zhang, & Du, 2019; Friedrich & Antranikian, 1996; Yamamura, Morita, Hasan, Yokoyama, & Tamiya, 2002; Williams, Richter, Mackenzie & Shih, 1990). Actinobacteria are also reported as commercial microbial strains including Streptomyces, Arthrobacter, Bevibacterium, and Nocardiopsis. However, some fungal strains have also displayed keratinolytic activity like Aspergillus, Paecilomyces, Doratomyces, Trichoderma, Fusarium, Acremonium, Onygena, Cladosporium, Microsporum, Lichtheimia, Chrysosporium, Aphanoascus, Trichophyton, and Scopulariopsi. It was revealed that the feather-degrading bacteria D. ficus and D. radiodurans R1 (Dalmaso et al., 2015) are capable of breaking down complete feathers at 30 °C while keratinase genes are present in extreme microorganisms such as D. geothermalis (Tang et al., 2021) and Deinococcus sp. D7000.
Previous studies have reported that DgokerA consists of three domains: a signal peptide (pre), an N-terminal pro-peptide (N-pro), and a mature domain. The mature domain contains three highly conserved catalytic triad residues formed by Asp171, His203, and Ser358 (Meng et al., 2022). Previous studies proved that site-directed mutagenesis is beneficial to increase enzyme activity under extreme conditions (Liu et al., 2013). Meng et al. (2022) conducted molecular modifications on DgokerA to create a mutant (A350S) that exhibited 30% higher enzyme activity. Their results provide a theoretical basis for the development and optimization of keratinase applications (Meng et al., 2022).
Moreover, Yang et al. (2021) employed bioinformatics approaches to extract relevant information related to lipases cold adaptation. They predicted the amino acid composition and the secondary and tertiary structures (Yang et al., 2021). Likewise, Sharma et al. (2022) conducted an in silico structural-functional analysis of keratinase, discovering distinct amino acid residues and their positions. Additionally, keratinase possesses a high-affinity calcium-binding site (Asp128, Leu162, Asn164, Ile166, and Val168) and a catalytic triad consisting of Asp119, His151, and Ser308, which are characteristic features of serine proteases (subtilisin family) (Sharma et al., 2022). In this study, we describe the genomic and proteomic features of keratinase (DgokerA) from Deinococcus gobiensis I-0 using bioinformatics approaches. Additionally, we predict the protease specificity, function, and cleavage sites of the enzyme on keratin substrate.
2. Materials and Methods
2.1. Structural analysis of DgokerA
A proprotein convertase subtilisin/kexin type 9 preproprotein (DgokerA) protein sequence (UniProt ID: H8GWN3) from Deinococcus gobiensis I-0 (strain DSM 21396) was used for the initial UniProt BLAST protein homology search http://www.uniprot.org/blast/ (The UniProt Consortium, 2018). In the next step, 20 DgokerA protein sequences (including the first three amino acids prior to histidine (H) as the first highly conserved amino acid) were aligned using the Clustal Omega algorithm to create a multiple sequence alignment (MSA) within the Jalview open-source program (version 2.10.3) http://www.jalview.org/ (Waterhouse et al., 2009). MtaKerA from M. taiwanensis WR-220 (UniProt ID: AWR86689.1) was used as the query sequence in this analysis.
AlphaFold2 was used to simulate the three-dimensional (3D) structures of DgokerA. AlphaFold2 is a neural network-based protein structure prediction tool developed by Google DeepMind. The "single sequence" mode was used for the prediction. During the sequence insertion process, the algorithm first searches for homologous sequences with preexisting structures. The algorithm was also instructed to perform an Amber relaxation procedure to correct any structural errors in the predicted model. Encoding these operations in separate programs has some important advantages. Once an approximation of the structure (AlphaFold, homology model, etc.) has been generated, unrealistic close contacts, poor atom overlaps, or strained geometry in bond lengths and angles may exist. The energy minimization of AMBER reduces the potential energy of the system such that such clashes are relieved and local geometry is enhanced. First, it allows individual pieces to be upgraded or replaced with minimal impact on other parts of the program suite; this has happened several times in Amber's history (Case et al., 2005). AlphaFold2 compiled a set of templates, and the one with the highest overall pLDDT scores was selected as the best template. AlphaFold2 provides a per-residue estimate of its confidence, ranging from 0 to 100. The model's predicted score corresponds to this confidence level indicator, or pLDDT. High-accuracy modeling is expected for regions with a pLDDT score > 90. Regions with pLDDT scores between 70 and 90 are considered well-modeled (mean good backbone prediction). Low-confidence regions, with pLDDT values between 50 and 70, should be handled with caution. To assess the structural similarity of proteins, the final predicted structure of DgokerA was superimposed using Chimera X.
2.2. Substrate cleavage sites of DgokerA
2.2.1. Effect of point mutations on the stability of DgokerA
The representative sequences of keratin and keratin-related proteins were chosen to predict possible sites of cleavage of DgokerA depending on their biological significance in relation to keratin degradation and species-specific structural variation. Predicting the stability of the crystal and 3D-visualized structure of DgokerA revealed that the subtilisin E-propeptide complex includes two chains or asymmetric units: subtilisin E domain (A) and E-propeptide (B), which are autoprocessed by the enzyme cleavage activity. In the present analysis, the sensitivity of point mutations and their stabilization were predicted at pH 7.0 using the MAESTRO (https://biwww.che.sbg.ac.at/maestro/web) (Laimer et al., 2016) web tool. The stability predictions were carried out at pH of 7.0 which is the physiological pH of most bacterial cytoplasmic environments and is usually used as a universal standard of stability evaluation of a protein. We performed this prediction to evaluate the effect of point mutations on the stability of DgokerA. MAESTRO, as a multi-agent machine learning system, provides information including (1) predicted ∆∆G (∆∆Gpred (kcal/mol), values of changes in standard binding free energy upon mutations), where ∆∆Gpred < 0.0 and ∆∆Gpred > 0.0 reveal the stabilizing and destabilizing mutations, respectively, (2) confidence estimation (Cpred), and (3) disulfide bond score (Sss). The scoring of disulfide bonds was also a part of the stability analysis since such covalent bonds are particularly important towards the stability of the tertiary structure and stability of proteins as a whole. Disulfide bonds help to increase resistance to thermal or chemical stress denaturation in extracellular enzymes like DgokerA and improve the structural integrity of the enzyme.
2.2.2. Genomic location of DgokerA gene
The genomic location of Deinococcus gobiensis I-0 (DgokerA) gene was identified using the Pathosystems Resource Integration Center (PATRIC) is the all-bacterial Bioinformatics Resource Center (BRC) (http://www.patricbrc.org) (Wattam et al., 2017). A joint effort by two of the original National Institute of Allergy and Infectious Diseases-funded BRCs, PATRIC provides researchers with an online resource that stores and integrates a variety of data types [genomics, transcriptomics, protein-protein interactions (PPIs), three-dimensional protein structures and sequence typing data] and associated metadata (Wattam et al., 2014).
2.3. Phylogenetic analysis
A total of 20 sequences were retrieved from the UniProt database. The phylogenetic tree was constructed by using the Molecular Evolutionary Genetics Analysis (MEGA) software (version 7.0) (http://www.megasoftware.net/) (Kumar et al., 2016; Tamura et al., 2021) through the maximum likelihood method. This method applies Bayesian inference to predict relative branch-length differences and presents a high-quality tree. In the current analysis, the alignment transformation environment (ALTER) tool (http://sing.ei.uvigo.es/ALTER/) (Glez-Peña et al., 2010). It was used to convert MSA of the coding DNA sequences (CDSs) into an aligned codon. DgokerA from Deinococcus gobiensis I-0 (UniProt ID: H8GWN3) was included in the analysis as an outgroup.
2.4. Subcellular localization and signal peptide
The subcellular localization and signal peptide were predicted by PSORT (http://psort.hgc.jp/) (Nakai and Kanehisa, 1991), and Signal P (http://www.cbs.dtu.dk/services/SignalP/) (Almagro Armenteros et al., 2019) respectively. Prediction of protein sorting signals and localization sites in amino acid sequences (PSORT) is a computer program for the prediction of protein localization sites in cells. It receives the information of an amino acid sequence and its source of origin, e.g., Gram-negative bacteria, as inputs. Then, it analyzes the input sequence by applying the stored rules for various sequence features of known protein sorting signals. Finally, it reports the possibility for the input protein to be localized at each candidate site with additional information. The SignalP 5.0 server predicts the presence of signal peptides and the location of their cleavage sites in proteins from Archaea, Gram-positive Bacteria, Gram-negative Bacteria, and Eukarya.
2.5. Prediction of transcription binding factors (TBFs)
The subtilisin E (aprE) gene from B. subtilis subsp. subtilis str. 168 was evaluated by DBTBS (http://dbtbs.hgc.jp/) (Sierro et al., 2008), which is a reference database used for the study of transcriptional regulation in B. subtilis based on characterized TBFs, targeted genes for regulation, and motifs recognized by transcription factors (TFs). The latest DBTBS release (release 5.0) can provide some updated information, including the number and name of TFs, position-specific scoring matrices, promoters, regulated operons, and terminators.
2.6. Sequence analysis and physiochemical properties
Physicochemical properties such as molecular weight (MW), theoretical isoelectric point (pI), amino acid composition, atomic composition, extinction coefficient, estimated half-life, instability index, aliphatic index, and grand average of hydropathicity index (GRAVY) were calculated using ProtParam (http://web.expasy.org/protparam/) (Gasteiger et al., 2005).
2.7. Docking studies
2.7.1. Receptor preparation
DgokerA crystal structure was generated using AlphaFold2 and then prepared for docking via AutoDock 4.2. This involved the removal of inhibitor and water molecules, the addition of hydrogen atoms to correct the ionization and tautomeric states of amino acid residues and Kollman charges. The protein was saved in pdbqt format.
2.7.2. Ligand preparation and prediction
Ligand was predicted by the structure-based method COACH, Protein-ligand-binding site recognition, TM-SITE, S-SITE, COFACTOR, ConCavity, and FINDSITE.
2.7.3. Receptor-ligand interaction by molecular docking
Molecular docking was used to estimate the scoring function and evaluate protein-ligand interactions to predict the binding affinity and activity of the ligand molecule. The AutoDock Tools was used to generate the bioactive binding poses of the ligands data set in the active site of DgokerA. The protein coordinates from the bound ligand were used to define the binding site. So, the scoring function was calculated using the standard protocol of the Lamarckian genetic algorithm. The grid map for docking calculations was centered on the target protein. Accelrys Discovery Studio 2021 software was used to model non-bonded polar and hydrophobic contacts in the inhibitor site of 6LU7. The docking results were visualized using PyMOL 2.3.4.0 and Discovery Studio Visualizer 4.0.
3. Results
3.1. Prediction and evaluation of 3D-structure of DgokerA
A 3D structural model of the DgokerA was generated using AlphaFold2. The AlphaFold2-based prediction was run with the “mmseqs2” mode by ColabFold. Furthermore, to further improve protein structure quality, we used amber force fields to fix structural violations in the 3D structure of nucleoproteins. Later, AlphaFold2 generated five models for each input sequence, and the top-ranked models were selected based on the pLDDT values. Figure 1 shows the 3D structure and the position of motifs and motif-like sequences which are indicated with arrows.
3.2. Cleavage site in the substrates of DgokerA
Prediction of cleavage sites of DgokerA from Deinococcus gobiensis I-0 in keratin and keratin-associated proteins from different hosts are shown in Figure 2 and Table 1.
The cleavage sites of DgokerA on the keratin or keratin-associated proteins. The cleaved proteins are from A. P07200 (Sus scrofa-Pig), B. P78386 (Homo sapiens), C. P02448 (Ovis aries), D. Q5IRA7 (Brachydanio rerio), and E. P49339 (Gallus gallus).
3.3. Effect of point mutations on the stability of DgokerA
The stability prediction of the crystal structure of DgokerA (UniProt ID: H8GWN3) revealed that point mutations can increase the stability of the subtilisin E protein structure. The stabilizing mutations (∆∆Gpred < 0.0) and confidence estimation (Cpred) for the chain of DgokerA are shown in Supplementary Data Table S1 (Supplementary Material). This evaluation revealed that the point mutations leading to changes of asparagine D171, and N260 to Isoleucine, Valine, and Isoleucine with ∆∆Gpred=-2.033, -2.001, and -1.997 respectively are the most stabilizing point mutations in DgokerA.
3.4. Genomic location of DgokerA (DgokerA) gene
Analysis of the genomic location of the DgokerA-coding gene (DgokerA) revealed that the DgokerA gene is located on the genome of Deinococcus gobiensis I-0 between 1,500,000 and 2,000,000 nucleotides (1,239 bp) (Figure 3). This genomic visualization showed that the DgokerA gene is expressed from a SigA (rA) promoter. Moreover, the transporter genes and unrecognized proteins, including malL, are located upstream of the DgokerA gene, whereas unidentified amino acid transporters are downstream of the aprE gene.
3.5. Phylogenetic analysis
The DgokerA showed significant relatedness to peptidase from the genus Deinococcus; they were localized at the same branch in the phylogenetic tree, indicating a common evolutionary origin. The protein sequence of DgokerA displayed 100% similarity to the S8 family peptidase from Deinococcus WR-420. DgokerA had a certain homology with the keratinases from the Deinococcus sp., such as S8 peptidase (WP_184136467) from Deinococcus humi and peptidase (WP_188973309) from D. aerolatus with % and 57.44% identity, respectively. The amino acid sequence of DgokerA, compared with those of keratinases, was peptidase from Deinococcus sp. Figures 4 and 5.
MSA analysis of 20 DgokerA protein sequences. Three highly conserved amino acids include arginine, aspartic acid, and glutamic acid.
3.6. Subcellular localization and signal peptide
Our subcellular localization results proved that DgokerA showed an extracellular localization score (10.00). SCL-BLAST analysis showed that DgokerA is (matched with 62906871, an extracellular serine proteinase precursor) extracellular proteases having non-cytoplasmic signal peptides as shown in Table 2.
3.7. Prediction of transcription binding factors (TBFs)
Five binding factors, AbrB, PucR, Spo0A, Hpr, SigA, and ComA, were found by prediction of the TBFs of the DgokerA gene as shown in Tables 3 and 4. These factors actively bind to the pertinent locations, serve as negative or positive regulators, or recognize the DgokerA gene promoter.
Transcriptional factors (TFs) of the DgokerA gene were obtained from the experimental result.
3.8. Sequence analysis and physiochemical properties
Physiochemical features of any protein play an integral role in tracing particular proteins' uniqueness. Expasy ProtParam provided the total amino acid composition of DgokerA protein as displayed in Table 5. Our results displayed that DgokerA has an equal number of negatively charged residues (Asp + Glu) and positively charged residues (Arg + Lys). DgokerA was shown 1.000 M-1 cm-1 assuming all pairs of Cys residues form cystines and 0.993 M-1 cm-1 excluding Cys residues at 280nm measured in water. Additionally, the instability index (Ii), which was less than 60 and equal to 19.64, showed that DgokerA is a stable protein. Our findings indicate that DgokerA has an aliphatic index of 82.72, which is favorable for thermostability. Our Grand Average of Hydropathy (GRAVY) result is positive (+0.035), indicating that the protein is hydrophobic. The theoretical pI of DgokerA was 6.07 to represent the acidic nature of the protein. Ala, Thr, Val, Ser, Leu amino acids play an integral role in the production of DgokerA.
3.9. Energetics and Ligand-protein interaction
One of the key methodologies in the field of computational protein engineering is molecular docking. This approach assists in estimating the biomolecular interactions of ligands with a specific binding site of a peculiar target protein. It predicts the binding potential of small molecules by placing them on the target site in a non-covalent fashion and predicts whether it forms a stable complex with high specificity or not. As a result of the obtained information from the docking experiment, it becomes relatively easier to predict the binding affinity, stability, and free energy of a ligand with the active site of the target protein. Currently, the approximation of binding free energy is an important objective of docking protocols, which is revealed in terms of hydrogen bonds, total internal energy, energy of dispersion and repulsion, torsional free energy, electrostatic force, energy of desolvation, and unbound system’s energy. In the current study, the ligands obtained from the COACH database were docked against the proteins and docked ligands were ranked based on the binding energies (kcal/mol). Only those ligands were selected which showed stronger binding energy than native ligands. The ZINC000016051516 and ZINC000196075144 are at the top of the list owing to their highest affinity and binding score (Table 6 and Figure 6). ZINC000016051516 formed one H-bond with Gly A: 236 with a binding affinity of -8.11758 kcal/mol. ZINC000041147074 formed interaction with Pro A: 266, Gly A: 265, Ser A: 235, Gly A: 264 residues. While ZINC000196075144formed H-bonds with Asn A:238 and Pro A: 266 residues of the target protein. Overall, the ligands ZINC000016051516, ZINC000041147074, and ZINC000196075144 demonstrate strong binding affinities, as evidenced by their hydrogen bonding interactions with specific residues within the target protein. These findings highlight their potential as promising candidates for further investigation and development as therapeutic agents.
Molecular docking of the selected best ligands into the target protein binding site. (A) 3D binding poses and (B) 2D ligand–protein interaction diagrams are shown for the top-ranked compounds. (Top) ZINC000016051516: The ligand (shown in red sticks) binds within the protein cavity (tan ribbon), with a magnified inset highlighting the binding orientation. The 2D interaction map reveals that a robust conventional hydrogen bond between Serine 300 (Ser A:300) and the ligand significantly contributes to stabilizing the complex. Additional interactions include van der Waals contacts with Gly A:234, Leu A:230, Val A:241, Gly A:264, Leu A:263, and Ser A:235; a carbon hydrogen bond with Gly A:236; and amide-π stacked and π-alkyl interactions with Val A:271, Pro A:266, Asn A:238, and Ser A:268. (Bottom) ZINC000041147074: The ligand (shown in blue sticks) occupies the binding pocket with favorable π-conjugated geometry. The presence of aromatic rings facilitates π–π stacking and π–alkyl interactions, as evidenced by amide-π stacking with Val A:271 and Gly A:236, and π-alkyl contacts with Leu A:263 and Val A:271. Conventional hydrogen bonds are formed with Ser A:268 and Ser A:300, while van der Waals interactions occur with Pro A:266, Gly A:265, Gly A:264, and Ser A:235. A carbon hydrogen bond is also observed with Ser A:235. (Not shown in image but per your text) ZINC000196075144: Key stabilizing forces include hydrogen bonds with Ser237 and Ser268, van der Waals contacts, and hydrophobic π-alkyl/amide-π stacking interactions with Val271, Pro266, and Gly236. The interaction legend indicates: light green = van der Waals; dark green = conventional hydrogen bond; pale green = carbon hydrogen bond; magenta = amide-π stacked; pink = π-alkyl.
As mentioned in Table 6, all three ligands ZINC000016051516, ZINC000041147074, and ZINC000196075144 conform to Lipinski’s rule of five, which is a rule of thumb to evaluate the drug-likeness or pharmaceutical potential of a compound. None of these ligands show any violation of the rule, implying that they possess considerable potential for being good protein-engineered candidates. Ligand ZINC000016051516, with a molecular weight of 148.16 Da, a partition coefficient (Log P) of 1.79, two hydrogen bond donors (HBD), one hydrogen bond acceptor (HBA), and MlogP of 1.91, complies well with the rule of five. Ligand ZINC000041147074 shows similar compliance with a slightly higher molecular weight (242.32 Da), a slightly elevated partition coefficient (Log P of 2.35), two hydrogen bond donors (HBD), no hydrogen bond acceptors (HBA), and MlogP of 2.21. In contrast, ligand ZINC000196075144, although in conformity with Lipinski's rule, has a significantly lower partition coefficient (Log P of -2.66), more hydrogen bond donors and acceptors (6 HBD and 5 HBA), and a much lower MlogP (-4.01). Its lower Log P and higher number of hydrogen bond donors and acceptors might suggest better solubility and possibly higher bioavailability.
4. Discussion
Industrial enzymes have become the core "chip" for bio-manufacturing technology. With the development of bioinformatics and computational technology, the catalytic mechanism of the enzyme can be solved by various calculation methods (Chen et al., 2019).
The MSA analysis of DgokerA protein sequences revealed that three amino acids including Glutamic acid, aspartic acid, and arginine are highly conserved in all Deinococcus sp. as well as Deinococcus gobiensis I-0. Moreover, the results from the prediction of cleavage sites in the keratin and keratin-associated proteins indicated that this enzyme cleaves the binding sites between hydrophobic [valine (V), isoleucine (I), glycine (G), and proline (P)] and polar [glutamine (G), asparagine (N), serine (S), threonine (T), cysteine (C), and tryptophan (W)] amino acids. The overall observations made by Qiu et al. (2020) and Li (2021) affirm that the conserved charged residues tend to be very important in protease active sites or binding pockets of the keratinases (El-Seedi et al., 2021; Qiu et al., 2020). In addition, these analyses revealed that DgokerA cleaves the hair keratins of humans (Homo sapiens) and Pig hair (Sus scrofa) as well as the wool keratin of sheep (Ovis aries) to further segments. Therefore, the hair and wool from human, sheep, and chicken feather sources could be more suitable substrates for DgokerA. The Keratinases (that degrade feather keratin) and the structural composition of keratins in these tissues (enriched with hydrophobic and polar amino acids, crosslinking with disulfide) support the idea that the substrates of DgokerA should be hair, wool, and feathers (Banasaz and Ferraro, 2024).
In addition, point mutation analysis revealed that the most stabilizing mutations occurred by changing D171 and N260 to Isoleucine, Valine, and Isoleucine, respectively. This result indicated that the bond between two hydrophobic residues translocates the enzyme cleavage site from the surface to the core, generating a more stable protein structure. Thus, it has been hypothesized that mutation in a hydrophobic amino acid could be the relevant target of further studies involving cleavage and site-directed mutagenesis, leading to improved protein stability and higher yield. Protein thermodynamic stability has long been driven by hydrophobic burial and enhanced core packing The incorporation of smaller or more closely packed hydrophobic side chains into the core is often found to enhance folded-state stability with an enhanced van der Waals contact and reduction of cavity formation (Pace et al., 2011).
Based on the genomic location of the transporter and carbon metabolism genes, such as maIL, downstream and upstream of the DgokerA gene. Therefore, an association between the expression of the DgokerA gene and transporter genes has been speculated, emphasizing the regulatory role of secretory genes in the expression of the DgokerA gene. Transcriptional binding factors analysis proved that ComA, Mta, PucR, and SigF overexpressed the DgokerA gene to increase the enzyme yield. The same pattern of gene clustering has been also described in Deinococcus and Bacillus species where metabolism and transporter genes lying near protease genes control the uptake of substrates and the release of enzymes (Mankge, 2024).
Regarding to molecular docking analysis, based on Lipinski's rule of five, while all three compounds have the potential to be successfully engineered proteins, the unique properties of ZINC000196075144 may offer better protein potential owing to its potential higher solubility and bioavailability (Daoud et al., 2021). However, a comprehensive analysis involving in vitro and in vivo studies, along with absorption, distribution, metabolism, excretion, and toxicity (ADMET) profiling, would be required to determine the definitive drug potential of these compounds (Veber et al., 2002).
5. Conclusion
Site-directed mutagenesis and TBFs, among other genomic and structural variables, can have a significant influence on the expression and stability of the recombinant DgokerA. Consequently, the use of genetic and protein engineering as well as system biology techniques could become able to increase the stability of Deinococcus gobiensis I-0 in both bench and large-scale production.
Supplementary Material
Supplementary material accompanies this paper.
Table S1
This material is available as part of the online article from https://doi.org/10.1590/1519-6984.303604
Acknowledgements
The authors appreciate the Faculty of Veterinary Medicine, University of Airlangga for supporting the APC of the manuscript. This study was funded by the Faculty of Veterinary Medicine, University of Airlangga under Contract No [297/UN3/HK.07.00/2025].
Data Availability Statement
This manuscript contains all the data.
References
-
ALMAGRO ARMENTEROS, J.J., TSIRIGOS, K.D., SØNDERBY, C.K., PETERSEN, T.N., WINTHER, O., BRUNAK, S., VON HEIJNE, G. and NIELSEN, H., 2019. SignalP 5.0 improves signal peptide predictions using deep neural networks. Nature Biotechnology, vol. 37, no. 4, pp. 420-423. https://doi.org/10.1038/s41587-019-0036-z PMid:30778233.
» https://doi.org/10.1038/s41587-019-0036-z -
BANASAZ, S. and FERRARO, V., 2024. Keratin from animal by-products: structure, characterization, extraction and application—a review. Polymers, vol. 16, no. 14, pp. 1999. https://doi.org/10.3390/polym16141999 PMid:39065316.
» https://doi.org/10.3390/polym16141999 -
CASE, D.A., CHEATHAM III, T.E., DARDEN, T., GOHLKE, H., LUO, R., MERZ JUNIOR, K.M. and WOODS, R.J., 2005. The Amber biomolecular simulation programs. Journal of Computational Chemistry, vol. 26, no. 16, pp. 1668-1688. https://doi.org/10.1002/jcc.20290 PMid:16200636.
» https://doi.org/10.1002/jcc.20290 -
CHEN, Q., LI, C., ZHENG, G., YU, H. and XU, J., 2019. Computational analysis of structure-activity relationship of industrial enzymes. Sheng Wu Gong Cheng Xue Bao, vol. 35, no. 10, pp. 1829-1842. https://doi.org/10.13345/j.cjb.190329 PMid:31668032.
» https://doi.org/10.13345/j.cjb.190329 -
DALMASO, G.Z., LAGE, C.A., MAZOTTO, A.M., DIAS, E.P., CALDAS, L.A., FERREIRA, D. and VERMELHO, A.B., 2015. Extracellular peptidases from Deinococcus radiodurans. Extremophiles : Life Under Extreme Conditions, vol. 19, no. 5, pp. 989-999. https://doi.org/10.1007/s00792-015-0773-y PMid:26216108.
» https://doi.org/10.1007/s00792-015-0773-y -
DAOUD, N.E.-H., BORAH, P., DEB, P.K., VENUGOPALA, K.N., HOURANI, W., ALZWEIRI, M. and TIWARI, V., 2021. ADMET profiling in drug discovery and development: perspectives of in silico, in vitro and integrated approaches. Current Drug Metabolism, vol. 22, no. 7, pp. 503-522. https://doi.org/10.2174/1389200222666210705122913 PMid:34225615.
» https://doi.org/10.2174/1389200222666210705122913 -
DUMORNÉ, K., CÓRDOVA, D.C., ASTORGA-ELÓ, M. and RENGANATHAN, P., 2017. Extremozymes: a potential source for industrial applications. Journal of Microbiology and Biotechnology, vol. 27, no. 4, pp. 649-659. https://doi.org/10.4014/jmb.1611.11006 PMid:28104900.
» https://doi.org/10.4014/jmb.1611.11006 -
EL-SEEDI, H.R., YOSRI, N., KHALIFA, S.A.M., GUO, Z., MUSHARRAF, S.G., XIAO, J., SAEED, A., DU, M., KHATIB, A., ABDEL-DAIM, M.M., EFFERTH, T., GÖRANSSON, U. and VERPOORTE, R., 2021. Exploring natural products-based cancer therapeutics derived from egyptian flora. Journal of Ethnopharmacology, vol. 269, pp. 113626. https://doi.org/10.1016/j.jep.2020.113626 PMid:33248183.
» https://doi.org/10.1016/j.jep.2020.113626 -
FANG, Z., SHA, C., PENG, Z., ZHANG, J. and DU, G., 2019. Protein engineering to enhance keratinolytic protease activity and excretion in Escherichia coli and its scale-up fermentation for high extracellular yield. Enzyme and Microbial Technology, vol. 121, pp. 37-44. https://doi.org/10.1016/j.enzmictec.2018.11.003 PMid:30554643.
» https://doi.org/10.1016/j.enzmictec.2018.11.003 -
FRIEDRICH, A.B. and ANTRANIKIAN, G., 1996. Keratin Degradation by Fervidobacterium pennavorans, a Novel Thermophilic Anaerobic Species of the Order Thermotogales. Applied and Environmental Microbiology, vol. 62, no. 8, pp. 2875-2882. https://doi.org/10.1128/aem.62.8.2875-2882.1996 PMid:16535379.
» https://doi.org/10.1128/aem.62.8.2875-2882.1996 -
GASTEIGER, E., HOOGLAND, C., GATTIKER, A., DUVAUD, S., WILKINS, M.R., APPEL, R.D., and BAIROCH, A. Protein identification and analysis tools on the ExPASy server. In: J.M. WALKER, editor. The proteomics protocols handbook New York (NY): Springer; 2005. p. 571-607. https://doi.org/10.1385/1-59259-890-0:571
» https://doi.org/10.1385/1-59259-890-0:571 -
GLEZ-PEÑA, D., GÓMEZ-BLANCO, D., REBOIRO-JATO, M., FDEZ-RIVEROLA, F. and POSADA, D., 2010. ALTER: program-oriented conversion of DNA and protein alignments. Nucleic Acids Research, vol. 38, no. Web Server issue, suppl. 2, pp. W14-W18. https://doi.org/10.1093/nar/gkq321 PMid:20439312.
» https://doi.org/10.1093/nar/gkq321 -
KUMAR, S., STECHER, G. and TAMURA, K., 2016. MEGA7: Molecular Evolutionary Genetics Analysis Version 7.0 for Bigger Datasets. Molecular Biology and Evolution, vol. 33, no. 7, pp. 1870-1874. https://doi.org/10.1093/molbev/msw054 PMid:27004904.
» https://doi.org/10.1093/molbev/msw054 -
LAIMER, J., HIEBL-FLACH, J., LENGAUER, D. and LACKNER, P., 2016. MAESTROweb: a web server for structure-based protein stability prediction. Bioinformatics (Oxford, England), vol. 32, no. 9, pp. 1414-1416. https://doi.org/10.1093/bioinformatics/btv769 PMid:26743508.
» https://doi.org/10.1093/bioinformatics/btv769 -
LI, Q., 2021. Structure, application, and biochemistry of microbial keratinases. Frontiers in Microbiology, vol. 12, pp. 674345. https://doi.org/10.3389/fmicb.2021.674345 PMid:34248885.
» https://doi.org/10.3389/fmicb.2021.674345 -
LIU, B., ZHANG, J., FANG, Z., GU, L., LIAO, X., DU, G. and CHEN, J., 2013. Enhanced thermostability of keratinase by computational design and empirical mutation. Journal of Industrial Microbiology & Biotechnology, vol. 40, no. 7, pp. 697-704. https://doi.org/10.1007/s10295-013-1268-4 PMid:23619970.
» https://doi.org/10.1007/s10295-013-1268-4 - MANKGE, M.E., 2024. Screening of alkaline proteases from Bacillus bacterial endophytes and delineation of their genetic features South Africa: University of Johannesburg.
-
MARQUES, C.R., 2018. Extremophilic microfactories: applications in metal and radionuclide bioremediation. Frontiers in Microbiology, vol. 9, pp. 1191. https://doi.org/10.3389/fmicb.2018.01191 PMid:29910794.
» https://doi.org/10.3389/fmicb.2018.01191 -
MENG, Y., TANG, Y., ZHANG, X., WANG, J. and ZHOU, Z., 2022. Molecular identification of keratinase DgokerA from Deinococcus gobiensis for feather degradation. Applied Sciences (Basel, Switzerland), vol. 12, no. 1, pp. 464. https://doi.org/10.3390/app12010464
» https://doi.org/10.3390/app12010464 -
NAKAI, K. and KANEHISA, M., 1991. Expert system for predicting protein localization sites in gram‐negative bacteria. Proteins, vol. 11, no. 2, pp. 95-110. https://doi.org/10.1002/prot.340110203 PMid:1946347.
» https://doi.org/10.1002/prot.340110203 -
PACE, C.N., FU, H., FRYAR, K.L., LANDUA, J., TREVINO, S.R., SHIRLEY, B.A. and SCHOLTZ, J.M., 2011. Contribution of hydrophobic interactions to protein stability. Journal of Molecular Biology, vol. 408, no. 3, pp. 514-528. https://doi.org/10.1016/j.jmb.2011.02.053 PMid:21377472.
» https://doi.org/10.1016/j.jmb.2011.02.053 -
QIU, J., WILKENS, C., BARRETT, K. and MEYER, A.S., 2020. Microbial enzymes catalyzing keratin degradation: Classification, structure, function. Biotechnology Advances, vol. 44, pp. 107607. https://doi.org/10.1016/j.biotechadv.2020.107607 PMid:32768519.
» https://doi.org/10.1016/j.biotechadv.2020.107607 -
SHARMA, C., TIMORSHINA, S., OSMOLOVSKIY, A., MISRI, J. and SINGH, R., 2022. Chicken Feather Waste Valorization Into Nutritive Protein Hydrolysate: Role of Novel Thermostable Keratinase From Bacillus pacificus RSA27. Frontiers in Microbiology, vol. 13, pp. 882902. https://doi.org/10.3389/fmicb.2022.882902 PMid:35547122.
» https://doi.org/10.3389/fmicb.2022.882902 -
SIERRO, N., MAKITA, Y., DE HOON, M. and NAKAI, K., 2008. DBTBS: a database of transcriptional regulation in Bacillus subtilis containing upstream intergenic conservation information. Nucleic Acids Research, vol. 36, no. Database issue, suppl. 1, pp. D93-D96. https://doi.org/10.1093/nar/gkm910 PMid:17962296.
» https://doi.org/10.1093/nar/gkm910 -
TAMURA, K., STECHER, G. and KUMAR, S., 2021. MEGA11: molecular evolutionary genetics analysis version 11. Molecular Biology and Evolution, vol. 38, no. 7, pp. 3022-3027. https://doi.org/10.1093/molbev/msab120 PMid:33892491.
» https://doi.org/10.1093/molbev/msab120 -
TANG, Y., GUO, L., ZHAO, M., GUI, Y., HAN, J., LU, W. and WANG, J., 2021. A novel thermostable keratinase from Deinococcus geothermalis with potential application in feather degradation. Applied Sciences (Basel, Switzerland), vol. 11, no. 7, pp. 3136. https://doi.org/10.3390/app11073136
» https://doi.org/10.3390/app11073136 -
THE UNIPROT CONSORTIUM, 2018. UniProt: the universal protein knowledgebase. Nucleic Acids Research, vol. 46, no. 5, pp. 2699. https://doi.org/10.1093/nar/gky092 PMid:29425356.
» https://doi.org/10.1093/nar/gky092 -
VEBER, D.F., JOHNSON, S.R., CHENG, H.-Y., SMITH, B.R., WARD, K.W. and KOPPLE, K.D., 2002. Molecular properties that influence the oral bioavailability of drug candidates. Journal of Medicinal Chemistry, vol. 45, no. 12, pp. 2615-2623. https://doi.org/10.1021/jm020017n PMid:12036371.
» https://doi.org/10.1021/jm020017n -
WANG, J., SALEM, D.R. and SANI, R.K., 2019. Extremophilic exopolysaccharides: a review and new perspectives on engineering strategies and applications. Carbohydrate Polymers, vol. 205, pp. 8-26. https://doi.org/10.1016/j.carbpol.2018.10.011 PMid:30446151.
» https://doi.org/10.1016/j.carbpol.2018.10.011 -
WATERHOUSE, A.M., PROCTER, J.B., MARTIN, D.M., CLAMP, M. and BARTON, G.J., 2009. Jalview Version 2: a multiple sequence alignment editor and analysis workbench. Bioinformatics (Oxford, England), vol. 25, no. 9, pp. 1189-1191. https://doi.org/10.1093/bioinformatics/btp033 PMid:19151095.
» https://doi.org/10.1093/bioinformatics/btp033 -
WATTAM, A.R., ABRAHAM, D., DALAY, O., DISZ, T.L., DRISCOLL, T., GABBARD, J.L., GILLESPIE, J.J., GOUGH, R., HIX, D., KENYON, R., MACHI, D., MAO, C., NORDBERG, E.K., OLSON, R., OVERBEEK, R., PUSCH, G.D., SHUKLA, M., SCHULMAN, J., STEVENS, R.L., SULLIVAN, D.E., VONSTEIN, V., WARREN, A., WILL, R., WILSON, M.J., YOO, H.S., ZHANG, C., ZHANG, Y. and SOBRAL, B.W. 2014. PATRIC, the bacterial bioinformatics database and analysis resource. Nucleic Acids Research, vol. 42, no. Database issue, pp. D581-D591. https://doi.org/10.1093/nar/gkt1099 PMid:24225323.
» https://doi.org/10.1093/nar/gkt1099 -
WATTAM, A.R., DAVIS, J.J., ASSAF, R., BOISVERT, S., BRETTIN, T., BUN, C., CONRAD, N., DIETRICH, E.M., DISZ, T., GABBARD, J.L., GERDES, S., HENRY, C.S., KENYON, R.W., MACHI, D., MAO, C., NORDBERG, E.K., OLSEN, G.J., MURPHY-OLSON, D.E., OLSON, R., OVERBEEK, R., PARRELLO, B., PUSCH, G.D., SHUKLA, M., VONSTEIN, V., WARREN, A., XIA, F., YOO, H. and STEVENS, R.L., 2017. Improvements to PATRIC, the all-bacterial bioinformatics database and analysis resource center. Nucleic Acids Research, vol. 45, no. D1, pp. D535-D542. https://doi.org/10.1093/nar/gkw1017 PMid:27899627.
» https://doi.org/10.1093/nar/gkw1017 -
WILLIAMS, C.M., RICHTER, C.S., MACKENZIE, J.M. and SHIH, J.C., 1990. Isolation, identification, and characterization of a feather-degrading bacterium. Applied and Environmental Microbiology, vol. 56, no. 6, pp. 1509-1515. https://doi.org/10.1128/aem.56.6.1509-1515.1990 PMid:16348199.
» https://doi.org/10.1128/aem.56.6.1509-1515.1990 -
YAMAMURA, S., MORITA, Y., HASAN, Q., YOKOYAMA, K. and TAMIYA, E., 2002. Keratin degradation: a cooperative action of two enzymes from Stenotrophomonas sp. Biochemical and Biophysical Research Communications, vol. 294, no. 5, pp. 1138-1143. https://doi.org/10.1016/S0006-291X(02)00580-6 PMid:12074595.
» https://doi.org/10.1016/S0006-291X(02)00580-6 -
YANG, G., MOZZICAFREDDO, M., BALLARINI, P., PUCCIARELLI, S. and MICELI, C., 2021. An in-silico comparative study of lipases from the antarctic psychrophilic ciliate Euplotes focardii and the mesophilic congeneric species Euplotes crassus. Marine Drugs, vol. 19, no. 2, pp. 67. https://doi.org/10.3390/md19020067 PMid:33513970.
» https://doi.org/10.3390/md19020067
Edited by
-
Editor:
Takako Matsumura Tundisi












