Food & Agri
Finding Candidate Genes for Breeding and Ingredient Development Through Meta-Analysis of Public Data
By reanalyzing public RNA-seq data with a common procedure, we narrowed down candidate stress-tolerance markers in turfgrass and candidate markers for evaluating functional ingredients, without running new experiments.
Narrowing Down Candidates with Public Data Before New Experiments
In breeding and in ingredient development alike, deciding which genes to use as markers usually takes many experiments. Meanwhile, public databases hold a large and growing body of RNA-seq (RNA sequencing) data deposited by researchers around the world. PtBio collected and reanalyzed this data through evidence marker discovery (meta-analysis) to narrow down candidate genes before collecting any new samples. We used the same method in two cases with very different subjects.
The Shared Method
- Search public databases and select relevant studies
- Collect the RNA-seq data from those studies and reanalyze it with a common procedure
- Calculate an integrated score for each gene
- Extract genes whose expression changes consistently
- Narrow down candidates for breeding or ingredient evaluation
The integrated score is the TN score (T for treated, N for non-treated), calculated for each gene as the number of treatment–control comparisons showing upregulation minus the number showing downregulation (Shintani et al., 2024). A result from a single experiment can be skewed by its conditions; the score summarizes the direction of expression changes across many treatment–control comparisons. The TN2 score counts comparisons with a TN ratio (treated to non-treated expression) above 2 as upregulated and those below 0.5 as downregulated. The reanalyzed data and the scores for each gene are stored in MATOI (Meta-Analytic Transcriptome Ortholog Index), a meta-analysis database we developed in house.
Case 1: Candidate Stress-Tolerance Markers in Turfgrass
Turfgrass production and management face challenges from heat, drought, salt damage, and disease. Breeding stress-tolerant turfgrass needs molecular markers that can support selection, and selection still relies largely on the breeder’s experience and intuition. Developing a stress-tolerant variety takes many years.
We applied the same method to two turfgrass species.
| Species A | Species B | |
|---|---|---|
| Question | A marker for selecting stress-tolerant varieties | The molecular basis of salt tolerance across species |
| Data used | 65 RNA-seq datasets from 35 publications | 90 data pairs (treatment vs. control) built from 117 samples in three projects |
| Comparison | Across three stresses: salt, cold, and cadmium | Four groups defined by salt tolerance (tolerant/sensitive) and tissue (shoot/root) |
- In species A, one gene that responds to all three stresses (salt, cold, and cadmium) remained. It is a candidate marker for further evaluation in breeding.
- In species B, the analysis identified 179 candidate genes associated with salt-stress responses, including genes for ion transport, osmotic adjustment, abscisic acid (ABA) signaling, and antioxidant defense. These genes may promote or suppress tolerance and are candidates to test as breeding targets.
Case 2: Candidate Markers for Evaluating Functional Ingredients
A common problem in developing functional ingredients is the lack of markers that show how an ingredient acts in the body. Collecting human clinical data in house is expensive. Meanwhile, results from a single public study may depend on the age group, sex, and region of the participants, and need further validation before use as evaluation markers.
We therefore reanalyzed public RNA-seq data on human obesity to find genes with consistent expression changes across studies. We searched public databases for keywords such as obesity and adipose tissue, then reviewed the results by hand and selected four studies that differ in age group, sex, region, and metabolic status. The same four studies were also identified through metadata screening using a large language model (LLM). In each study, we ranked genes with increased expression in obese versus non-obese participants by TN2 score and took about the top 500. Because many genes share the same TN2 score around rank 500, we chose the TN2 score cutoff that gave a gene count closest to 500, keeping genes with the same score together. Each set therefore contains between 409 and 575 genes. We then identified the genes shared by the four gene sets.
| BioProject accession | Comparison | Source |
|---|---|---|
| PRJNA434431 | Non-obese vs. obese | Gao et al., EBioMedicine 2018 |
| PRJNA641129 | Non-obese vs. obese (including groups with different metabolic status) | Cifarelli et al., J Clin Invest 2020 |
| PRJNA682573 | Non-obese vs. obese | Fisk et al., EBioMedicine 2022 |
| PRJNA847049 | Non-obese vs. obese | Yang et al., Nat Metab 2022 |

Figure: About the top 500 genes by TN2 score in each study (409 to 575 genes, depending on the study), overlapped across the four studies. Numbers are gene counts for each overlap; circle labels are BioProject accessions. This figure was generated by reanalyzing public RNA-seq data from the four studies listed above.
Six genes showed increased expression in all four studies. These genes are candidates associated with obesity-related expression changes that were consistent across studies with different participant backgrounds. Whether they can serve as markers for evaluating functional ingredients requires validation in experiments that test responses to those ingredients.
One Method, Different Subjects
In two very different subjects, stress tolerance in plants and obesity in humans, the same method narrowed down candidates to test in breeding and ingredient evaluation. When sufficient, comparable public RNA-seq data are available for the research question, reanalyzing them may reduce the cost and time needed for candidate discovery compared with generating new experimental data. This approach narrows down candidate genes and potential markers for evaluating functional ingredients, helping guide the next stage of experimental validation.
Related information
- Partner: PtBio (reanalysis of public data)
- Fields: digital breeding, functional foods, biomarkers
- Service used: Candidate Gene & Marker Discovery (evidence marker discovery)
- TN score method: Shintani M, Tamura K, Bono H. Meta-analysis of public RNA sequencing data of abscisic acid-related abiotic stresses in Arabidopsis thaliana. Front Plant Sci. 2024;15:1343787. doi:10.3389/fpls.2024.1343787
- Method for extracting public data with LLMs: Shintani M, Andrade D, Bono H. A Pilot Evaluation of Open-Weight Large Language Models for Screening RNA-seq Metadata in Public Databases. bioRxiv 2026. doi:10.64898/2026.02.16.706241