Open-access Protein selection gain in soybean grains based on segregating generations

Abstract

Soybean (Glycine max L.) is a major global crop due to its diverse uses, high nutritional value, and strong production potential. This study quantified selection gains for grain protein content across segregating generations and proposed breeding strategies adjusted to heterozygosity levels. The experiment was conducted in Ijuí, Rio Grande do Sul, using an augmented block design with common controls, including 170 progenies and 44 cultivars. Genetic effects, heritability, and selection gains were estimated through Bayesian inference with Markov Chain Monte Carlo. Broad-sense heritability (H2=0.451) indicated moderate genetic control of protein content. Greater selection intensity was suitable for the F10 generation due to low heterozygosity (0.19%), whereas earlier generations (F5, F7, and F8) required milder selection to preserve genetic variability. Adjusting selection intensity across generations ensures reliable parameter estimation and consistent genetic gain, highlighting the usefulness of Bayesian approaches in protein-oriented soybean breeding.

Keywords:
Yield; biochemical composition; heterozygosity

INTRODUCTION

Soybean (Glycine max L.) is one of the most important commodities in the global agricultural landscape. Its prominence is mainly due to its wide range of uses, nutritional quality, and production potential. Considered the main source of vegetable protein, soybeans play a fundamental role in both human and animal nutrition, in addition to being used in various industries, such as pharmaceuticals and cosmetics (Kim et al. 2021). Furthermore, the biochemical composition of soy protein includes the amino acids leucine, isoleusine, lysine, methionine, phenylalanine, threonine, tryptophan, and valine (Kudelka et al. 2021). These attributes make soy an important ingredient in the formulation of products such as protein beverages, flour, and nutritional supplements.

Historically, breeding programs have prioritized increasing productivity, which has reduced the focus on quality traits. Thus, selection strategies for protein should be conducted carefully, especially over segregating generations, where heterozygosity is greater and genetic variability can be exploited more efficiently.

In early segregating generations, elevated heterozygosity not only increases genetic variability but also affects the expression of genotype × environment (G×E) interaction. Heterozygous loci may exhibit dominance or partial dominance effects that are differentially expressed across environments, thereby amplifying phenotypic plasticity and inflating the apparent G×E interaction. As a result, phenotypic performance in these generations may reflect environment-specific dominance effects rather than stable additive genetic values, reducing the accuracy of selection for protein content. As homozygosity increases in later generations, dominance effects are progressively reduced, and the additive genetic signal becomes more consistent across environments, allowing a clearer discrimination of genotypic performance.

Previous studies on protein selection in soybeans demonstrate modest gains, especially when selection is applied late or without adjusting the intensity to the level of available genetic variability (Assefa et al. 2019). However, few studies have evaluated how selection intensity should vary between generations (F₂, F₃, F₄...), nor how different strategies can maximize additive gain in the face of progressively decreasing heterozygosity.

Furthermore, selection strategies applied to each generation still lack comparative studies, especially considering that the heritability of protein content is influenced by genotype × environment interaction and the polygenic nature of the trait (Pathan et al. 2013). This gap hinders the definition of optimized protocols for segregating populations.

Bayesian methods have gained prominence in plant breeding for offering greater stability in scenarios of small samples, hierarchical structure, and high uncertainty typical conditions of segregating populations. Unlike frequentist approaches, these methods allow the incorporation of prior information, estimation of the complete distribution of parameters, and better handling of variability between generations, making them a promising alternative for predicting selection gains.

Given these limitations and opportunities, this study aimed to quantify selection gains for protein content across different segregating generations and propose selection strategies adjusted to the heterozygosity level of each generation, using Bayesian modeling to improve the accuracy and robustness of the estimates.

MATERIAL AND METHODS

The study was conducted in Ijuí (lat 28° 53’ 10” S and long 52° 59’ 55” W, alt 300 m asl), in the state of Rio Grande do Sul, by the genetic improvement program of the Regional University of Northwest of the State of Rio Grande do Sul. The soil is classified as a Typical Dystrophic Red Latosol, and the climate is characterized as humid subtropical Cfa. The progenies were sown in the first half of November 2023 and harvested in the first half of April 2024.

The experimental design was an augmented block design with common controls. The regular treatments consisted of 170 soybean progenies (Figure 1), and the common treatments consisted of 44 commercial soybean cultivars (Tables 1 and 2), arranged in four blocks. The experimental units consisted of a ten-meter-long seeding row spaced 0.5 m apart. For all genotypes, a seeding density of 18 seeds per linear meter was used, with a base fertilizer of 250 kg ha-1 of NPK (05-20-20). Phytosanitary management was carried out preventively to mitigate biotic effects (insect pests, diseases, and invasive plants) in the experiment.

Table 1
Soybean control cultivars and the relative maturity group
Table 2
Genealogy of soybean segregating lines and populations

Figure 1
Chronological diagram of the genetic improvement process using hybridizations, F1 to F10.

The progenies were stratified into 38 F5 lines (87.5% homozygosity and 6.25% ​​heterozygosity), 40 F7 lines (96.88% homozygosity and 1.56% heterozygosity), 63 F8 lines (98.44% homozygosity and 0.78% heterozygosity) and 29 F10 lines (99.61% homozygosity and 0.19% heterozygosity) (Tables 1 and 2). The main variable evaluated after harvest was grain protein, obtained by assessing 50 representative plants from each experimental unit, where the grains were manually threshed and grouped, and a 100 g sample of grains was separated and subjected to a drying process until it reached 13% moisture.

Subsequently, the percentage of total crude protein (%) was determined using the acid-catalytic digestion method, followed by distillation and titration using the Kjeldahl method, employing a conversion factor of 6.25, indicated for soybean cultivation (Kjeldahl 1883). The data obtained were subsequently subjected to analysis of the assumptions of the statistical model, normality and homogeneity of residual variances, and model additivity. Afterwards, the selection differentials were calculated using the overall mean of the commercial controls (C) and the segregating generations (F5, F7, F8, and F10). Thus, the overall standard deviation of the experiment was calculated for the next steps (Table 3).

Table 3
Selection differentials (SD) used for line selection

Subsequently, the Bayesian inference model based on the Markov Chain Monte Carlo (MCMC) algorithm was used, through the MCMCglmm function, to estimate the genetic effects. Thus, informative prior matrix, random effects for the genotypes, 100,000 iterations and a burn-in of 10,000 were considered. The model used was based on Bandeira et al. (2025a), as follows:

Y = X β + Z 1 δ 1 + Z 2 δ 2 + e

Where: y is the vector of phenotypic values, 𝛽 is the vector of the incidence matrix and corresponding to the vector of systematic effects (general mean), Z1 and Z2 are the incidence matrices of random effects, δ₁ is the vector of block effects and δ₂ is the vector of genetic values, e is the residual vector.

In this model, the mean parameters of the protein percentage referring to the posterior distribution (post mean), 95% credible interval (UP-95% CrI), significance of the probabilistic model by the Markov chain Monte Carlo method (pMCMC) at 5% probability and broad-sense heritability (H2) were estimated. To understand the transgressive profile of the lines and promote selection, the method of identifying transgressives based on selection differentials graphically, segregated by generation, was used, using the transg function through the EstimateBreed package (Bandeira et al. 2025b). All analyses were performed using R software (R Core Team 2023).

RESULTS AND DISCUSSION

The use of selection differentials enabled the evaluation of the productive performance of progenies in relation to the objectives established by the breeding program, allowing the estimation of potential genetic gains for grain protein content. This methodology complements the selection strategies applied across generations (Table 4), including F5, F7, F8, and F10, and enables comparisons with the overall mean of the controls (C = 35.66%), the best-performing control (BC), and reference thresholds based on the standard deviation of the controls (C + 1S, C + 2S, C + 3S).

Table 4
Estimated genetic parameters for the categories and segregating generations of the studied genotypes

The analysis of genetic parameters revealed that the F5 generation presented the highest posterior mean for protein content (up to 36.47%), whereas F10 showed the lowest mean (35.82%). The overall category presented the narrowest credible interval (0.81%), while segregating generations particularly F7 (2.07%) and F10 (2.87%) displayed broader intervals. This pattern is consistent with the increased homozygosity and the expression of additive genetic variance as lines become more genetically fixed (Sun et al. 2025).

Although the F5 generation exhibited the highest mean protein content and F10 the lowest, these differences should not be interpreted as being driven exclusively by heterozygosity. The segregating generations evaluated in this study originated from different parental combinations, and variation in parental genetic background may have contributed to the observed contrasts among generations. Therefore, the higher protein concentration observed in F5 may reflect, at least in part, the superior genetic potential of specific crosses rather than an intrinsic advantage of higher heterozygosity. In this context, heterozygosity should be interpreted as a contributing factor rather than a causal determinant, and the results represent generation-level patterns within a heterogeneous breeding population rather than direct within-cross comparisons.

Broad-sense heritability (H2=0.451) demonstrated moderate genetic control of protein content when considering the 170 progenies and 44 controls. Although lower than the values ​​reported by Zhang et al. (2025) (H2=0.7-0.9), the magnitude still indicates the presence of exploitable genetic variability. In evaluations conducted in a single year and environment, it is expected that the reduction in environmental variance may inflate the H² estimates due to confounding between genetic effects and genotype × environment interaction. However, the moderate value observed in this study suggests that the genetic variance for protein content in this specific population is relatively limited and/or that the residual variance remains significant, even under controlled environmental conditions. This behavior may be associated with the high level of homozygosity of the evaluated lines, the relatively narrow genetic base of the material, and the highly polygenic nature of protein content. Previous studies also highlight that heterogeneity between environments is substantial for this trait, due to the effects of temperature and nitrogen availability during grain filling (Carneiro et al. 2019). Therefore, validation in multiple environments and growing seasons is recommended to confirm the stability of the lines and ensure more robust and generalizable heritability estimates.

Selection gains differed across generations and intensities (Table 5). The highest gains (up to 4.23%) were obtained under strong selection pressure (1%), especially in F10 and F7, which presented greater fixation and narrower phenotypic variance. In contrast, early selection in F5 requires lower selection pressure (80-90%) to maintain genetic diversity, followed by progressive intensification in F7 and rigorous selection in F8 and F10. This strategy aligns with the dynamics of heterozygosity (6.25% in F5 to 0.19% in F10) and ensures the retention of favorable recombinants. Similar patterns were described by Carvalho et al. (2024), who observed meaningful genetic gains for protein content under indirect and stringent selection strategies.

Table 5
Selection gain, stratified by segregating generations

The protein data also highlighted contrasts between population-level and individual selection (Figure 2). The use of C + 1S (39.24%), C + 2S (42.29%), and C + 3S (45.34%) as thresholds enabled identification of high-performing populations (e.g., IRC 007 and IRC 047) and detection of transgressive recombinant inbred lines. Notably, line L25 exceeded the C + 3S threshold and achieved values above 45%, originating from DM 7.0 BMX Magna RR × Monasca RR, indicating strong complementarity between parental alleles affecting nitrogen assimilation and storage protein pathways (Port et al. 2024). Although values above 45% are uncommon in commercial soybean, typically ranging from 38-42%, high-protein breeding lines have occasionally surpassed these levels in controlled environments, particularly when derived from crosses specifically targeted for seed quality traits. Thus, the value for L25 appears plausible rather than resulting from measurement error, though replication in multi-environment trials is advisable.

Figure 2
a. Transgressive selection of populations, subjected to selection differentials (SD): SD 5: Overall mean of 44 commercial controls - C; SD 6: Mean of the best commercial control - BC; SD 7: Reference mean to obtain 40% protein - Breeding; SD 8: Overall mean of controls plus one standard deviation (3.62%) - (C + 1S); SD 9: Overall mean of controls plus two standard deviations (2 x 3.62%) - (C + 2S); SD 10: Overall mean of controls plus three standard deviations (3 x 3.62%) - (C + 3S); b. Transgressive selection of 170 segregating lines, F5 (93.75% homozygosity and 6.25% heterozygosity), 40 lines F7 (98.44% homozygosity and 1,56% heterozygosity), 63 lines F8 (99.22% homozygosity and 0.78% heterozygosity) and 29 lines F10 (99.80% homozygosity and 0.20% heterozygosity), subjected to the selection differentials (SD) SD 5: General mean of the 44 commercial controls - C; SD 6: Mean of the best commercial control - BC; SD 7: Reference mean to obtain 40% protein - Breeding); SD 8: General mean of the controls plus one standard deviation (3.62%) - (C + 1S); SD 9: General mean of the controls plus two standard deviations (2 x 3.62%) - (C + 2S): SD 10: General mean of the controls plus three standard deviations (3 x 3.62%) - (C + 3S).

Comparison with commercial controls (Figure 3) showed that only two cultivars (DM5958 RSF IPRO and M6410 IPRO) approached the C + 2S threshold, confirming their potential for use in quality-oriented germplasm banks. The relatively low mean of the control group (35.66%) below typical commercial standards likely reflects the specific environmental conditions during grain filling, especially temperature and water availability, which are known to reduce protein levels (Carneiro et al. 2019). This reinforces the role of G×E in the expression of this trait and highlights the need for expanded trials to validate the superiority of high-protein lines.

Overall, the results confirm the presence of exploitable genetic variability for seed protein content, support moderate heritability, and highlight the identification of superior transgressive lines such as L25. These findings demonstrate potential for genetic progress, though future evaluations must incorporate multi-environment trials to assess stability, quantify G×E interactions, and reconcile the documented negative genetic correlation between protein and oil content an important factor for breeding decisions.

Figure 3
Descriptive analysis of controls according to selection differences (SD) SD 1: General mean for the segregating generation F5; SD 2: General mean for the segregating generation F7; SD 3: General mean for the segregating generation F8; SD 4: General mean for the segregating generation F10; SD 5: General mean of the 44 commercial controls - C; SD 6: Mean of the best commercial control - BC; SD 7: Reference mean to obtain 40% protein - Breeding); SD 8: General mean of the controls plus one standard deviation (3.62%) - (C + 1S); SD 9: General mean of the controls plus two standard deviations (2 x 3.62%) - (C + 2S): SD 10: General mean of the controls plus three standard deviations (3 x 3.62%) - (C + 3S).

CONCLUSION

The selection gain based on the F10 generation with 0.19% heterozygosity appears promising and supports the use of higher selection intensity. However, the segregating generations F5, F7, and F8 should be cautious regarding soybean protein selection and employ milder pressures to ensure genetic gain.

Data Availability Statement

The datasets generated and/or analyzed during the current research are available from the corresponding author upon reasonable request.

REFERENCES

  • Assefa ST, Yang EY, Chae SY, Song M, Lee J, Cho MC, Jang S2019 Alpha glucosidase inhibitory activities of plants with focus on common vegetablesPlants 9:2
  • Bandeira WJA, Carvalho IR, Pradebon LC, Sausen NH, Sangiovo JP, Bruinsma GMW, Loro MV2025a Estimation of genetic parameters and selection pressure in canary seed (Phalaris canariensis L.)Pesquisa Agropecuária Brasileira 60:e03781
  • Bandeira WJA, Carvalho IR, Loro MV, Pradebon LC, Silva JAG2025b EstimateBreed: Estimation of environmental variables and genetic parameters. Available at: Available at: https://cran.r-project.org/web/packages/EstimateBreed/index.html Accessed on January 30, 2025.
    » https://cran.r-project.org/web/packages/EstimateBreed/index.html
  • Carneiro AK, Bruzi AT, Pereira JLAR, Zambiazzi EV2019 Stability analysis of pure lines and a multiline of soybean in different locationsCrop Breeding and Applied Biotechnology 19:395-401
  • Carvalho IR, Silva JAG, Loro MV, Sarturi MVDR, Hutra DJ, Lautenchleger F2021 Soybean canonical nutraceutical interrelations and their reflections on breedingAgropecuária Catarinense 34:67-75
  • Kjeldahl J1883 Neue methode zur bestimmung des stickstoffs in organischen Körpern. Fresenius, Zeitschrift f. analChemie 22:366-382
  • Kim SH, Tripathi P, Yu S, Park J-M, Lee JD, Chung YS, Chung G, Kim Y2021 Selection of tolerant and susceptible wild soybean (Glycine soja Siebold & Zucc.) accessions under waterlogging condition using vegetation indicesPolish Journal of Environmental Studies 30:3659-3675
  • Kudełka W, Kowalska M, Popis M2021 Quality of soybean products in terms of essential amino acids compositionMolecules 26:5071
  • Pathan SM, Vuong T, Clark K, Lee J, Shannon JG, Roberts CA, Ellersieck MR, Burton JW, Cregan PB, Hyten DL, Nguyen HT, Sleper DA2013 Genetic mapping and confirmation of quantitative trait loci for seed protein and oil contents and seed weight in soybeanCrop Science 53:765-774
  • Port ED, Carvalho IR, Pradebon LC, Loro MV, Colet CF, Silva JAG, Sausen NH2024 Early selection of resilient progenies to seed yield in soybean (Glycine max L.) populationsCiência Rural 54:e20230287
  • R Core Team2023 R: a language and environment for statistical computing. R Foundation for Statistical Computing, Vienna.
  • Sun X, Hu B, Li WX, Ning HL2025 Genetic basis and simulated breeding strategies for enhancing soybean seed protein content across multiple environmentsPlants 14:2117
  • Zhang D, Sun X, Hu B, Li WX, Ning H2025 QTN Mapping, gene prediction and molecular design breeding of seed protein content in soybeanThe Crop Journal 13:1116-1126

Edited by

  • SCIENTIFIC EDITOR:
    Luiz Antônio dos Santos Dias

Publication Dates

  • Publication in this collection
    20 July 2026
  • Date of issue
    2026

History

  • Received
    10 Oct 2025
  • Accepted
    18 Feb 2026
  • Published
    24 Feb 2026
location_on
Crop Breeding and Applied Biotechnology Universidade Federal de Viçosa, Departamento de Fitotecnia, 36570-000 Viçosa - Minas Gerais/Brasil, Tel.: (55 31)3899-2611, Fax: (55 31)3899-2611 - Viçosa - MG - Brazil
E-mail: cbab@ufv.br
rss_feed Acompanhe os números deste periódico no seu leitor de RSS
Ir para o topo Reportar erro