Does anyone have go-to public datasets with explicit batch annotations (center, lane, date) for evaluating normalization and batch correction in bulk RNA‑seq pipelines? I’m comparing ComBat, RUVSeq, and limma::removeBatchEffect in a Snakemake+Conda setup on 100–500 samples and need phenotype labels to benchmark simple classifiers afterwards; GEO/SRA accessions or curated links appreciated.
emma21
November 24, 2025, 4:03am
2
GEUVADIS E-GEUV-1 has lab/lane, pop/sex labels (about 465); great for ComBat: BioStudies < The European Bioinformatics Institute < EMBL-EBI .
GTEx v8 via recount3 (https://rna.recount.bio/ ) — SMCENTER/SMGEBTCH available; ComBat-friendly if tissue labels work; easy to subset 100–500.
karen90
November 28, 2025, 2:16am
4
Try the SEQC/MAQC‑III UHRR/HBRR study — it was designed for cross‑site RNA‑seq benchmarking and SRA metadata includes center, platform, and lane; curated overview: A comprehensive assessment of RNA-seq accuracy, reproducibility and information content by the Sequencing Quality Control Consortium | Nature Biotechnology . It’s great for ComBat/RUVSeq comparisons, but phenotypes are simple (UHR vs brain mixes), so if you need a tougher benchmark I can share a Snakemake‑ready manifest to subsample 100–500.