Welcome to scRiskDB
A platform for displaying Human Disease-Related Tissue-Specific Resources.
scRiskDB provide the annotated,prioritized tissue-,development-specific data
Gene to Function
Exploring disease risk geneset related pathways
Go >>>Risk to Relevant Cells
Exploring disease associated high-risk celltypes based on SC-VAR:sce-DRS approach
Go >>>Latest database update
Version 2026.01
Reference
Gefei Zhao, Binbin Lai SC-VAR: a computational tool for interpreting polygenic disease risks using single-cell epigenomic data, Briefings in Bioinformatics, Volume 26, Issue 2, March 2025, bbaf123, https://doi.org/10.1093/bib/bbaf123
Repository content goes here.
SNV to Disease
Here you can explore tissue specific disease single nucleotide variation (SNV)
RISK GENES
Here you can explore tissue specific disease risk genes
Risk Gene List
Adult-Specific Risk Genes
Fetal-Specific Risk Genes
RISK CREs
Here you can explore tissue specific disease risk CREs
Adult-Specific Risk CREs
Fetal-Specific Risk CREs
Risk Pathways
Here you can explore tissue specific disease high-risk Pathways
Enrichment Analysis Results
RISK CELLS
Here you can explore tissue specific disease risk Cell Types
Cell Type-specific Information
Select a cell type and disease to view cell type specific Trait-associated CREs, genes, and SNPs.
Download Selected Rows
Additional Risk Cell Analysis
Explore disease-critical celltypes of your own data >>>Analyze Your Single-Cell Data
Here you can explore high-risk celltypes of your own single cell data
Upload & Submit
(Note: This column name must exist in adata.obs)
Peaks will be automatically harmonized via overlap mapping.
Check Status & Download
Enter your Task ID below to retrieve results (even after closing the page).
User Guide & Documentation
👋 Welcome to scRiskDB
scRiskDB is a comprehensive databsase platform designed to interpret polygenic disease risks using single-cell data.
Key Features:
-
Prioritized Annotation: Connects GWAS SNPs to candidate Risk Genes and cis-Regulatory Elements (CREs) in specific tissue/cell-type contexts.
-
Functional Insight: Uncovers disease-enriched Pathways (GO terms) and highlights Critical Cell Types that drive disease susceptibility.
-
Multi-dimensional Validation: Integrated with Open Targets/LDexpress (for genes) and 3D Genome Browser (for CREs) to provide causal evidence.
-
Custom Analysis: Upload your own single-cell data (RNA or ATAC) to calculate disease enrichment scores.
📊 Data Sources
GWAS: FINNGEN R11
We selected 318 traits with clear ICD-11 codes, summarized as follows:
| ICD11-Code | Num. | Description |
|---|---|---|
| M13 | 42 | Diseases of the musculoskeletal system and connective tissue |
| K11 | 41 | Diseases of the digestive system |
| I9 | 36 | Diseases of the circulatory system |
| N14 | 28 | Diseases of the genitourinary system |
| O15 | 25 | Pregnancy, childbirth and the puerperium |
| J10 | 23 | Diseases of the respiratory system |
| H7 | 23 | Diseases of the eye and adnexa |
| G6 | 13 | Diseases of the nervous system |
| F5 | 15 | Mental and behavioural disorders |
| E4 | 14 | Endocrine, nutritional and metabolic diseases |
| L12 | 11 | Diseases of the skin and subcutaneous tissue |
| H8 | 10 | Diseases of the ear and mastoid process |
| AB1 | 10 | Certain infectious and parasitic diseases |
| C3 | 8 | Neoplasms, from cancer register |
| ST19 | 6 | Injury, poisoning and certain other consequences of external causes |
| CD2 | 6 | Neoplasms from hospital discharges |
| D3 | 4 | Diseases of the blood and blood-forming organs and immune mechanism disorders |
| Z21 | 1 | Factors influencing health status and contact with health services |
| R18 | 1 | Symptoms, signs and abnormal clinical and laboratory findings, NEC |
Single-Cell Atlas Data:
We included single-cell data with atlas level, summarized as follows:
-
single cell ATAC+RNA data of human Brain, GEO: GSE204682
-
sci-ATAC-seq data of human adult tissues, GEO: GSE184462, GEO: GSE160472, dbGaP: phs001961, dbGaP: phs002204
-
sci-ATAC-seq data of human fetal tissues, dbGaP: phs002003
🧬 Risk Genes Module
This module enables the exploration of tissue-specific disease candidate risk genes identified by SC-VAR framework.
📋 How to Use
- Select Context: Choose a Tissue (e.g., Brain) and a Disease (e.g., Schizophrenia) from the dropdown menus.
- Browse Results: The table displays candidate risk genes sorted by
ZSTAT(Association Weight).And you can search interested genes by their gene name in the Search Box. - Analyze & Validate: Use the validation links described below to gather more causal evidence.
Interface Guide:
What’s More?
- Developmental Comparison: Click the “Compare Adult vs Fetal” button.
- This generates two side-by-side tables showing genes specific to the Adult stage versus the Fetal stage, helping you identify developmental-specific risk drivers.
- Click “Download Results” to export the full gene list for your own analysis.
🌟 Functional & Causal Validation
We provide two powerful external integration to help users move from “Predicted Candidate Gene” to “Validated Target”:
1. Open Targets Platform (Therapeutic Evidence)
-
How to access: Click on the Blue Gene Name (e.g.,
GRIN2A) in the results table. -
Usage: Directly navigates to the Associations page for that gene.
-
Why use it? (Example: GRIN2A in Schizophrenia)
-
Assess Druggability (ChEMBL): Check the ChEMBL column to see if there are known drugs or clinical compounds targeting this gene (e.g., GRIN2A is a known target for NMDA receptor modulators).
-
Verify Genetic Evidence: The GWAS and Gene Burden columns provide orthogonal genetic support, confirming the gene’s association across independent cohorts.
-
Literature Mining: The Europe PMC column aggregates text-mining evidence, helping you quickly find relevant publications linking the gene to the disease.
-
2. NCI LDexpress (eQTL & LD Analysis)
- How to access: Click the
Check LDbutton in the Validation column. - How to use:
- First, click the Copy Icon(1) to automatically copy the top 10 risk SNVs associated with this gene.
- Then, click the Link(2) to open LDexpress. Paste the SNVs into the input box.
- eQTL Validation: It searches if your risk variants (or variants in Linkage Disequilibrium with them) are associated with gene expression changes in GTEx v8 tissues.
- Mechanism: LDexpress helps determine if a non-coding variant functions by regulating the expression levels of nearby genes.
- LD Calculation: LDexpress calculates \(R^2\) and \(D'\) statistics based on 1000 Genomes Project populations to find proxy variants.
🔗 Risk cis-Regulatory Elements (CREs)
This module links non-coding GWAS variants to functional regulatory elements (Enhancers/Promoters).
📋 General Usage
-
Select Context: Choose a Tissue and Disease from the dropdown menus.
-
Filter & Search: Use the Search box to find specific Genes or genomic regions.
-
Download: Click “Download Results” to save the annotated CRE list for downstream analysis.
-
Developmental Comparison: Click the “Compare Adult vs Fetal” button.
- This generates two side-by-side tables showing CREs specific to the Adult stage versus the Fetal stage(The same as gene page).
- Click “Download Results” to export the full gene list for your own analysis.
🌟 Key Feature 1: Cell-Type Specificity
-
How to use: Check the
Linked Cell Typescolumn in the results table. -
What it tells you: We explicitly label which specific cell types (e.g., Astrocyte, Excitatory Neuron) show chromatin accessibility at this CRE region.
-
Why it matters: This helps you determine if the disease risk is driven by a specific cellular context, rather than a tissue effect.
🌟 Key Feature 2: Physical Validation (Hi-C Tutorial)
To prove that a CRE actually regulates a distant gene, we need physical evidence that they touch each other in 3D space. We integrate the 3D Genome Browser for this purpose.
Follow this Step-by-Step Guide to visualize the interaction:
Step 1: Get the Coordinates
-
Find the Validation column in the result table.
-
Click the Copy Icon.
-
Note: The copied coordinates are automatically expanded by ±100kb (e.g.,
chr1:182,100,000-182,400,000) to ensure you can see the surrounding chromatin loops and TAD structures.
Step 2: Select the Dataset
-
Click the
3D Viewbutton to open the 3D Genome Browser. -
Configure the browser based on your query:
-
Species: Select
Human. -
Organ: Select the tissue you are interested in (e.g.,
Brain). -
Cell Type: Select the cell type that matches the
Linked Cell Typescolumn (e.g.,Astrocyte). -
Click
+ Addto load the dataset.
-
Visual Guide: Dataset Selection
(Make sure to match the Cell Type with our “Linked Cell Types” recommendation)
Step 3: Visualize & Interpret
- Paste the coordinates you copied in Step 1 into the search bar at the top.
Visual Guide: Interpretation
(Example: A clear chromatin interaction loop in Cerebellum Astrocytes)
1. The “Red Triangle” (Heatmap) :
The map shows interaction frequency. The x-axis is the genomic coordinate. The intense red pixels indicate regions that physically touch each other. What to look for: Look for a dark red dot or cluster off the diagonal line that aligns with your CRE and the target Gene. This “Spark” confirms they are interacting in the nucleus.
2. TADs (Topologically Associating Domains) : You will see large, triangular blocks of red signal along the diagonal. These are TADs—structural “neighborhoods” of the genome.
Interpretation: DNA sequences within the same TAD interact frequently, while interactions across TAD boundaries are insulated (blocked).
Functionality Rule: An Enhancer (CRE) and its Target Gene usually reside within the same TAD. If your CRE and Gene are in different TADs (separated by a white boundary), the regulatory link is less likely.
3. Chromatin Loops (The Arcs) : Look at the bottom track showing purple/blue arcs.
Interpretation: These arcs represent high-confidence loops (e.g., Enhancer-Promoter loops).
Functionality Rule: If you see an arc where one end anchors at your CRE and the other end anchors at the Gene Promoter, this is the strongest evidence of functional regulation.
4. Summary: Is this CRE functional?
| Evidence Level | Observation in Browser |
|---|---|
| High Confidence | A clear Loop (Arc) connects the CRE to the Gene Promoter; both are inside the same TAD. |
| Medium Confidence | No distinct Loop, but they are in the same TAD with high interaction frequency (red signal) between them. |
| Low Confidence | The CRE and Gene are in different TADs (separated by a boundary) or show no interaction signal (white). |
SNV to Function (📊 Pathways & Critical Cells)
Pathways
-
Select a disease and tissue from the drop-down menus.
-
You can choose either the
Display Tablebottom to view the GO results. -
When results is ready
-
Choose the terms you are interested in.Then Click
Plot Selected Termsbottom to Visualize the GO results.
Aggregate SNV information to trait-associated Cells
-
Select a tissue with any diseases you interested in the tab.
-
Click Generate Plot.
-
A heatmap will be generated showing the disease-related risk scores across cell types.
- Asterisks (*) mark cell types significantly associated with the selected disease.
-
Use the dropdown to select a specific cell type and explore associated risk genes and CREs.
Example:
Selected: Left_Ventricle (Adult) → Cell type: Fibroblast → Disease: I9_AF
Output: - CRE: chr10-103448764-103449065\
- GENE:
SH3PXD2A\ - SNPs:
rs555347613,rs114681462,rs148371916
Gene SH3PXD2A has been reported to be associated with atrial fibrillation (AF) in GWAS.
SNV Annotation
Focus on SNV
This module allows you to query specific single nucleotide variations (SNVs) to explore their associations with risk genes and cis-regulatory elements (CREs) in a specific tissue-disease context.
How to use:
-
Select Context: Choose a target Tissue and Disease from the dropdown menus (e.g
Adrenal_Glands(Fetal),I9_ANGINA) -
Search by ID: Enter a valid
rsID(e.g.,rs1000000851) in the input box and click the Search button. -
View Results: The table displays the query results, including:
- CRE: The associated Cis-Regulatory Element ID.
- GENE: The linked Risk Gene symbol.
- Weight: The Z-score indicating the significance of the SNP in this context.
Interactive Features (Jump & Filter):
The result table is interactive to facilitate deep diving into specific targets:
- Jump to Gene Analysis: Click on any Gene Name (in blue blod) to automatically switch to the SNV to Risk Genes tab. The system will auto-select the current tissue/disease and filter for that specific gene.
- Jump to CRE Analysis: Click on any CRE ID (in blue bold) to automatically switch to the SNV to Risk CREs tab. The system will auto-select the current tissue/disease and filter for that specific CRE.
- Download: If results are available, click Download Results to save the data as a
.csvfile.
🧮 Analyze Your Single-Cell Data
Upload and score your own datasets (scRNA-seq or scATAC-seq) against our disease risk models.
Upload & Submit
- Upload File: Choose your
.h5adfile (Max 10GB). - Configuration:
-
RNA: Select Disease & Tissue from our database.
-
ATAC:
Option A:
Direct Retrieval & Auto-Harmonization (For existing tissues): Select a Tissue and Disease from the interface. The backend automatically retrieves the corresponding Risk CREs
Option B:
Download our template, create your own risk file, and upload it.
-
- Submit: Click the
Submit Analysisbutton.- You will receive a unique Job ID (e.g.,
b119c6b6...). Please copy this ID.
- You will receive a unique Job ID (e.g.,
Retrieve Results
- Task Persistence: You can close the window and come back later!
- Check Status:
- Paste your Job ID into the “Enter Task ID” box.
- If the task is
Completed, a greenDownload Resultsbutton will appear.
- Please note results not downloaded within 24 hours will be cleared by the system.
Click to view detailed ATAC Data Formatting & Workflow
⚠️ How to prepare ATAC risk data? (Step-by-step Tutorial)
The workflow is identical to the SC-VAR pipeline. You can follow the complete code tutorial here:
1. Required Input Data Formats
A. GWAS Summary Statistics
A text file containing SNP association data. Space or tab-delimited.
Required Columns: chr, pos, rsids, pval
B. Peak-to-Gene Links (scATAC Data)
A file defining the co-accessibility networks (which peaks regulate which genes). Tab-delimited.
Required Columns: gene, chr, start, end
2. Required Reference Files
You must ensure the reference genome versions match your data (hg19 or hg38).
- Gencode Annotation: Use v43 for hg19 / v44 for hg38.
- File:
gencode.vxx.annotation.gff3.gz
- File:
- MAGMA Gene Annotation:
- Original MAGMA gene body definition file (
magma0.genes.annot).
- Original MAGMA gene body definition file (
- 1000 Genomes Panel:
- PLINK binary files (
.bed/.bim/.fam) for LD calculation (e.g.,g1000_eur).
- PLINK binary files (
3. Processing Workflow
The SC-VAR pipeline consists of five key steps to generate the final risk annotation file:
Step 1: Load Data
Import GWAS statistics (scv.read_gwas). Import Peak-Gene connections (scv.get_p2g_conn or scv.load_peak_data).
Step 2: Map SNPs to Peaks
Identify GWAS SNPs that fall within the open chromatin peaks (scv.snp_peak).
Step 3: Map Peaks to Genes
Correlate peaks with gene genomic coordinates using the Gencode GFF3 file (scv.gene_corr).
Step 4: Build Annotation
Combine the overlap matrix, gene correlations, and MAGMA annotations to construct the final Peak-SNP-Gene annotation file (scv.annotate). Output Filename Example: Schizophrenia.scemagma.genes.annot
Step 5: Get risk File
When get sce-MAGMA gene results, Then use (generate_atac_risk_file) function this will give you the final candidate score risk file.
📂 Required Output Format
After running the pipeline, you need to generate a .csv (tab split)or .txt file to upload to scRiskDB.
- Header: Required.
- Columns:
CRE_ID: Genomic coordinates (format:chr:start-endorchr-start-end).ZSTAT: The calculated association score (Z-score).
Example File Content can be download at the Cell Scoring page.
❓ Frequently Asked Questions
Q: What should I do if I cannot find my rsID in the SNV search?
A: Make sure the rsID is correctly formatted (e.g., rs123456). If no result appears, the SNP may not be present in the selected tissue or disease context. And do not click Downloads bottom when there is no results,It will give you an error.
Q: Can I search for a gene instead of an rsID?
A: Yes. In the SNV-to-Gene and SNV-to-CRE modules, you can search by gene symbol (e.g., TP53) to explore disease-associated regulatory mechanisms linked to that gene.
Q: What does the ZSTAT/Weight represent in the results?
A: The ZSATA or the Weight is the Z score calculated based on P-values reflects the association strength of the SNP, gene or CRE with the selected disease under the specified tissue context.
Q: What if I lose my Task ID? A: The Task ID is unique to each submission. If lost, you may need to resubmit the analysis.
Q: My ATAC upload failed.
A: Ensure your .csv risk file follows the template format exactly (Columns: CRE_ID, ZSTAT).And your single cell data may not exceed 10 GB due to the limitations of the web server.
Q: Why do some tissues have no results for a specific disease?
A: Risk association depends on the presence of trait-relevant SNPs or active regulatory elements in that specific tissue. If no signal is found, the table will be empty.
📌 Citation
If you use scRiskDB in your research, please cite:
Gefei Zhao, Binbin Lai. SC-VAR: a computational tool for interpreting polygenic disease risks using single-cell epigenomic data. Briefings in Bioinformatics, 2025. DOI: 10.1093/bib/bbaf123
Gefei Zhao, Binbin Lai. scRiskDB: A Single-Cell Epigenomic Resource Linking Complex Traits to Regulatory Mechanisms across Human Tissues.
References
- Zhang, Hou, et al. Polygenic enrichment distinguishes disease associations of individual cells in single-cell RNA-seq data, Nature Genetics, 2022.
- Yu G, Wang L, Han Y and He Q. clusterProfiler: an R package for comparing biological themes among gene clusters. OMICS: A Journal of Integrative Biology, 2012.
- Lin SH, Thakur R, Machiela MJ. LDexpress: an online tool for integrating population-specific linkage disequilibrium patterns with tissue-specific expression data. BMC Bioinformatics. 2021;22(1):608.
- Buniello, Suveges, et al.Open Targets Platform: facilitating therapeutic hypotheses building in drug discovery, Nucleic Acids Research, Volume 53, Issue D1, 6 January 2025, Pages D1467–D1475.
Contact Us
If you have any questions or suggestions, please don't hesitate to reach out to us. Your input is valuable to us.
Contact: Gefei Zhao & Binbin Lai
Email: laib@bjmu.edu.cn
Address: Peking University Health science center, Beijing, China
Our lab: https://laiblab.github.io
Submit tissue-specific risk list
Instructions:
Please ensure your file follows the required format. You can download the template below:
Download Submission Template (.csv)