Food & Agri

Finding Candidate Genes for Breeding and Ingredient Development Through Meta-Analysis of Public Data

By reanalyzing public RNA-seq data with a common procedure, we narrowed down candidate stress-tolerance markers in turfgrass and candidate markers for evaluating functional ingredients, without running new experiments.

Narrowing Down Candidates with Public Data Before New Experiments

In breeding and in ingredient development alike, deciding which genes to use as markers usually takes many experiments. Meanwhile, public databases hold a large and growing body of RNA-seq (RNA sequencing) data deposited by researchers around the world. PtBio collected and reanalyzed this data through evidence marker discovery (meta-analysis) to narrow down candidate genes before collecting any new samples. We used the same method in two cases with very different subjects.

The Shared Method

  1. Search public databases and select relevant studies
  2. Collect the RNA-seq data from those studies and reanalyze it with a common procedure
  3. Calculate an integrated score for each gene
  4. Extract genes whose expression changes consistently
  5. Narrow down candidates for breeding or ingredient evaluation

The integrated score is the TN score (T for treated, N for non-treated), calculated for each gene as the number of treatment–control comparisons showing upregulation minus the number showing downregulation (Shintani et al., 2024). A result from a single experiment can be skewed by its conditions; the score summarizes the direction of expression changes across many treatment–control comparisons. The TN2 score counts comparisons with a TN ratio (treated to non-treated expression) above 2 as upregulated and those below 0.5 as downregulated. The reanalyzed data and the scores for each gene are stored in MATOI (Meta-Analytic Transcriptome Ortholog Index), a meta-analysis database we developed in house.

Case 1: Candidate Stress-Tolerance Markers in Turfgrass

Turfgrass production and management face challenges from heat, drought, salt damage, and disease. Breeding stress-tolerant turfgrass needs molecular markers that can support selection, and selection still relies largely on the breeder’s experience and intuition. Developing a stress-tolerant variety takes many years.

We applied the same method to two turfgrass species.

Species ASpecies B
QuestionA marker for selecting stress-tolerant varietiesThe molecular basis of salt tolerance across species
Data used65 RNA-seq datasets from 35 publications90 data pairs (treatment vs. control) built from 117 samples in three projects
ComparisonAcross three stresses: salt, cold, and cadmiumFour groups defined by salt tolerance (tolerant/sensitive) and tissue (shoot/root)
  • In species A, one gene that responds to all three stresses (salt, cold, and cadmium) remained. It is a candidate marker for further evaluation in breeding.
  • In species B, the analysis identified 179 candidate genes associated with salt-stress responses, including genes for ion transport, osmotic adjustment, abscisic acid (ABA) signaling, and antioxidant defense. These genes may promote or suppress tolerance and are candidates to test as breeding targets.

Case 2: Candidate Markers for Evaluating Functional Ingredients

A common problem in developing functional ingredients is the lack of markers that show how an ingredient acts in the body. Collecting human clinical data in house is expensive. Meanwhile, results from a single public study may depend on the age group, sex, and region of the participants, and need further validation before use as evaluation markers.

We therefore reanalyzed public RNA-seq data on human obesity to find genes with consistent expression changes across studies. We searched public databases for keywords such as obesity and adipose tissue, then reviewed the results by hand and selected four studies that differ in age group, sex, region, and metabolic status. The same four studies were also identified through metadata screening using a large language model (LLM). In each study, we ranked genes with increased expression in obese versus non-obese participants by TN2 score and took about the top 500. Because many genes share the same TN2 score around rank 500, we chose the TN2 score cutoff that gave a gene count closest to 500, keeping genes with the same score together. Each set therefore contains between 409 and 575 genes. We then identified the genes shared by the four gene sets.

BioProject accessionComparisonSource
PRJNA434431Non-obese vs. obeseGao et al., EBioMedicine 2018
PRJNA641129Non-obese vs. obese (including groups with different metabolic status)Cifarelli et al., J Clin Invest 2020
PRJNA682573Non-obese vs. obeseFisk et al., EBioMedicine 2022
PRJNA847049Non-obese vs. obeseYang et al., Nat Metab 2022

Venn diagram of genes with increased expression in four public RNA-seq studies of human obesity. Six genes are shared by all four studies.

Figure: About the top 500 genes by TN2 score in each study (409 to 575 genes, depending on the study), overlapped across the four studies. Numbers are gene counts for each overlap; circle labels are BioProject accessions. This figure was generated by reanalyzing public RNA-seq data from the four studies listed above.

Six genes showed increased expression in all four studies. These genes are candidates associated with obesity-related expression changes that were consistent across studies with different participant backgrounds. Whether they can serve as markers for evaluating functional ingredients requires validation in experiments that test responses to those ingredients.

One Method, Different Subjects

In two very different subjects, stress tolerance in plants and obesity in humans, the same method narrowed down candidates to test in breeding and ingredient evaluation. When sufficient, comparable public RNA-seq data are available for the research question, reanalyzing them may reduce the cost and time needed for candidate discovery compared with generating new experimental data. This approach narrows down candidate genes and potential markers for evaluating functional ingredients, helping guide the next stage of experimental validation.

  • Partner: PtBio (reanalysis of public data)
  • Fields: digital breeding, functional foods, biomarkers
  • Service used: Candidate Gene & Marker Discovery (evidence marker discovery)
  • TN score method: Shintani M, Tamura K, Bono H. Meta-analysis of public RNA sequencing data of abscisic acid-related abiotic stresses in Arabidopsis thaliana. Front Plant Sci. 2024;15:1343787. doi:10.3389/fpls.2024.1343787
  • Method for extracting public data with LLMs: Shintani M, Andrade D, Bono H. A Pilot Evaluation of Open-Weight Large Language Models for Screening RNA-seq Metadata in Public Databases. bioRxiv 2026. doi:10.64898/2026.02.16.706241
← Back to Projects