Bio-DX

Candidate Gene & Marker Discovery

We combine RNA-seq with meta-analysis of public data to identify candidate genes for genomic breeding and evidence markers, which are marker genes whose expression changes are reproduced across multiple studies.

Service Overview

We combine gene expression data with public databases to identify evidence markers: marker genes whose expression changes consistently across multiple independent studies. Because they rest on reproduced results rather than a single experiment, they make more reliable breeding targets and indicators. Short-read RNA-seq is our standard method. We add Iso-Seq with long reads (PacBio HiFi) when the goal calls for it.

1. RNA-seq analysis

We profile gene expression across the transcriptome.

Ideal for:

  • Finding genes whose expression differs between treated and control groups
  • Identifying the gene sets involved in a physiological state or a stimulus
  • Examining non-coding RNAs

Key analyses: expression quantification, differential expression analysis (identifying differentially expressed genes), visualization (principal component analysis [PCA], heatmaps, volcano plots, clustering), functional analysis (Gene Ontology [GO] enrichment and pathway analysis), and gene network analysis

We run every step from expression quantification through differential expression, visualization, and functional analysis on Basepair, a cloud platform for RNA-seq analysis from Basepair Inc. (US). Our RNA-seq expertise, combined with Basepair’s browser-based interactive visualization, lets you explore and reanalyze your results in the browser yourself, even without bioinformatics staff in house.

Options: non-coding RNA (ncRNA) analysis and fusion gene analysis. Neither is part of the standard Basepair workflow; we run both as custom analyses.

2. Iso-Seq analysis (full-length transcripts)

We read transcripts end to end with long reads to resolve isoform structure.

Ideal for:

  • Telling apart isoforms produced by alternative splicing
  • Finding new transcripts
  • Comparing expression between isoforms

Key analyses: identification of full-length transcripts (isoforms), matching to known genes and quality-based classification, isoform-level expression quantification, and function prediction

3. Evidence marker discovery (meta-analysis)

We reanalyze and integrate public RNA-seq data to find markers with stronger evidence.

Ideal for:

  • Adding statistical power when your own data are not enough
  • Finding markers shared across several independent studies
  • Backing a hypothesis with stronger evidence

Key analyses: integration of multiple studies, higher detection power, lower bias, and identification of evidence markers

For this search we use MATOI (Meta-Analytic Transcriptome Ortholog Index), a meta-analysis database we developed in house. We reanalyze public RNA-seq data with a common procedure and score each candidate gene with an integrated score, the TN score (T for treated, N for non-treated; Shintani et al., 2024), calculated for each gene as the number of experiments showing upregulation minus the number showing downregulation. This narrows the candidates to genes that move consistently across conditions, which a single experiment cannot show.

The PtBio Approach

Incomplete reference transcriptomes, sparse gene annotation, and scarce public data can complicate expression analysis of organisms used in industry, and a generic pipeline can miss findings as a result.

PtBio designs each analysis around your research goal and the biology of your organism, and works with you as a partner to draw out findings a standard analysis would miss.

What we resolve in the initial consultation

Your challengeWhat PtBio proposes
Not sure which RNA-seq method fitsA method matched to your goal, budget, and organism
Too many differentially expressed genes to narrow downPrioritization and stronger evidence through meta-analysis
Want to use public data but don’t know where to startSelection, retrieval, and integration of datasets that fit your goal
Unsure what the results mean biologicallyFunctional interpretation and proposals for the next experiment or analysis

Why PtBio

StrengthWhat it means for you
Collaborative custom analysisOne team supports you from experimental design to interpretation, tailored to your organism and research goal
Know-how with public databasesWe select and integrate data that fit your goal from GEO (Gene Expression Omnibus) and other repositories
Joint research lab at Hiroshima UniversityWe apply methods developed at the Laboratory of Bio-DX (Professor Hidemasa Bono, Graduate School of Integrated Sciences for Life)
Support for R&D decisionsEvery delivery pairs the data analysis with biological interpretation and proposed next steps

Technologies

MethodSequencing platformRead typeMain use
RNA-seqIlluminaShort readsExpression quantification, differentially expressed genes, pathway analysis
Iso-SeqPacBio (HiFi)Long readsFull-length transcripts, isoform analysis

Process

  1. Contact us — We propose an analysis method that fits your goal.
  2. Experimental design & sample preparation (approx. 1–2 weeks) — Consultation on your objectives, experimental design, high-quality RNA extraction and quality assessment
  3. Transcriptome data acquisition (approx. 1–2 weeks) — Selection of RNA-seq or Iso-Seq, library preparation, sequencing, quality control (QC) report preparation
  4. Data analysis & gene expression profiling (approx. 1–2 weeks) — Expression quantification, differential expression analysis, isoform analysis, functional enrichment analysis
  5. Result interpretation & report preparation (approx. 1–2 weeks) — Functional analysis and annotation of candidate genes, pathway analysis, custom report preparation, proposals for future research strategy

A meta-analysis-only project has no sequencing step. We hold the initial consultation, select public data (1–2 weeks), run the integrated analysis (2–4 weeks), and deliver.

Plans & Pricing

RNA-seq data acquisition — 3 Gb per sample (about 20 million reads)

From ¥35K (excluding tax, per sample) / Delivery: 1 month

We acquire RNA-seq data with short reads.

Candidate gene and marker discovery for genomic breeding

From ¥80K (excluding tax) / Delivery: 1 month

We identify differentially expressed genes, analyze their function, and visualize the data.

Cost Examples

Costs vary by project. Use these examples as a guide and contact us for a quote.

AnalysisExampleSample sizeCost (excl. tax)
RNA-seq expression analysisEnvironmental response in a marine organism, including custom analysis and a report tailored to the research goal9 samplesFrom approx. ¥1.0M
Evidence markers (meta-analysis)Public data integration for an industrial microbe30 data pairsFrom approx. ¥900K

Note: Prices include sequencing where the project involves it, and vary with sample count, analysis scope, and options.

The RNA-seq example covers more than standard expression analysis. It includes custom analysis designed around the research goal and a report that interprets the results and proposes the next experiments. We don’t stop at a list of differentially expressed genes: we work with you on what the data mean and what to do next.

Deliverables

You receive an analysis report that summarizes the results, and the full set of data behind it. We organize the data by analysis step in folders and deliver it on a hard drive or through cloud storage, in a form your team can reanalyze or hand to another provider.

Analysis report (PDF and HTML)

One document covers how we sequenced, what data we obtained, how we analyzed it, and what we found. Results come with discussion and suggestions for what to examine next. Technical terms are explained in an appendix.

Included in every project

  • For projects that include sequencing, raw sequencer output with checksums so you can confirm the files are intact (how to verify them, in Japanese)
  • Data volume and quality checks
  • A list of the software and databases we used, so the analysis can be reproduced later

When we compare gene activity between conditions (RNA-seq)

  • A table of how strongly each gene is expressed in each sample
  • A list of genes whose expression changed between conditions, with the size of the change and its statistical confidence, so you can shortlist candidates
  • Figures that show the results at a glance: overall sample trends, differences between conditions, and genes that changed in more than one comparison
  • An analysis of what the changed genes do (GO, Kyoto Encyclopedia of Genes and Genomes [KEGG])
  • Reads aligned to the reference genome, with alignment rates

When we examine transcript structure (Iso-Seq)

  • Sequences and structures of the detected transcripts (isoforms)
  • Their correspondence to known genes and a quality classification, for discovering new transcripts and shortlisting candidates
  • Expression levels per isoform (Excel)
  • Predicted function of each isoform

When we search public data for candidates (meta-analysis)

  • A list of the public datasets and condition pairs we used
  • A table of how expression changed between conditions in each study
  • A list of genes that consistently go up or down across studies (Excel), ranked so you can start from the strongest evidence
  • Functional trends among the candidate genes (GO, KEGG)
  • The reference sequences and annotation we used

On request

  • A cloud environment where your team can analyze its own data in a browser
  • A genome browser
  • A searchable database of the meta-analysis results

FAQ

Q: How many samples do we need?

A: We recommend at least 3 samples per condition for statistical power. A pilot analysis can start with 3 samples per group.

Q: Can you analyze an organism with no reference genome?

A: Yes. We use de novo transcriptome assembly.

Q: Can you combine our data with public data?

A: Yes. We can integrate your RNA-seq data with public datasets.

Q: We find the results hard to interpret.

A: We work through them with you, from biological interpretation to proposals for the next experiment or analysis.

Q: How does Iso-Seq differ from short-read RNA-seq?

A: Iso-Seq captures full-length transcripts and suits isoform analysis. Short-read RNA-seq is well suited to quantifying expression. We propose whichever fits your goal.

Q: How much public data is available for meta-analysis?

A: GEO and similar databases hold millions of samples. We select datasets that match your research theme and propose them to you.

Q: What is the typical timeline?

A: As a guide, a whole project takes 1–2 months from sample receipt to report delivery for RNA-seq, and 1–1.5 months from the start of the project for meta-analysis. The delivery times under Plans & Pricing cover single steps only. The timeline depends on sample count and analysis scope, so we give a project-specific estimate in the initial consultation.

Q: Can we get a cost estimate first?

A: Yes. We provide a rough estimate in the initial consultation.

Contact Us

For RNA-seq and meta-analysis inquiries and quotes, please use our inquiry form.

← Back to Services