跳到主要内容
知仓学习社ZHICANG

bulk-rna-seq-deseq2-analysis-with-omicverse

Walk Claude through PyDESeq2-based differential expression, including ID mapping, DE testing, fold-change thresholding, and enrichment visualisation.

不碰外部(只输出文字)无严重或高危命中FreedomIntelligence/OpenClaw-Medical-Skills

它会碰到什么

扫了多少2 个文本文件,5 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

Bulk RNA-seq DESeq2 analysis with omicverse

Overview

Use this skill when a user wants to reproduce the DESeq2 workflow showcased in [t_deseq2.ipynb](../../omicverse_guide/docs/Tutorials-bulk/t_deseq2.ipynb). It covers loading raw featureCounts matrices, mapping Ensembl IDs to symbols, running PyDESeq2 via ov.bulk.pyDEG, and exploring downstream enrichment plots.

Instructions

  1. Import and format the expression matrix
  • Call import omicverse as ov and ov.utils.ov_plot_set() to standardise visuals.
  • Read tab-separated count data from featureCounts using ov.utils.read(..., index_col=0, header=1).
  • Strip trailing .bam from column names with [c.split('/')[-1].replace('.bam', '') for c in data.columns].
  1. Map gene identifiers
  • Ensure the appropriate mapping pair exists by running ov.utils.download_geneid_annotation_pair().
  • Replace gene_id with gene symbols using ov.bulk.Matrix_ID_mapping(data, 'genesets/pair_<GENOME>.tsv').
  1. Initialise the DEG object
  • Create dds = ov.bulk.pyDEG(data) from the mapped counts.
  • Resolve duplicate gene names with dds.drop_duplicates_index() and confirm success in logs.
  1. Define contrasts and run DESeq2
  • Collect sample labels into treatment_groups and control_groups lists that match column names exactly.
  • Execute dds.deg_analysis(treatment_groups, control_groups, method='DEseq2') to invoke PyDESeq2.
  1. Filter and tune thresholds
  • Inspect result shape (dds.result.shape) and optionally filter low-expression genes, e.g. dds.result.loc[dds.result['log2(BaseMean)'] > 1].
  • Set thresholds via dds.foldchange_set(fc_threshold=-1, pval_threshold=0.05, logp_max=6) to auto-pick fold-change cutoffs.
  1. Visualise differential genes
  • Draw volcano plots with dds.plot_volcano(...) and summarise key genes.
  • Produce per-gene boxplots: dds.plot_boxplot(genes=[...], treatment_groups=..., control_groups=..., figsize=(2, 3)).
  1. Run enrichment analyses (optional)
  • Download enrichment libraries using ov.utils.download_pathway_database() and load them through ov.utils.geneset_prepare.
  • Rank genes for GSEA with rnk = dds.ranking2gsea().
  • Instantiate gsea_obj = ov.bulk.pyGSEA(rnk, pathway_dict) and call gsea_obj.enrichment() to compute terms.
  • Plot enrichment bubble charts via gsea_obj.plot_enrichment(...) and GSEA curves with gsea_obj.plot_gsea(term_num=..., ...).
  1. Troubleshooting
  • If PyDESeq2 raises errors about size factors, remind users to provide raw counts (not log-transformed data).
  • gene_id mapping depends on species; direct them to download the correct genome pair when results look sparse.
  • Large pathway libraries may require raising recursion limits or filtering to the top N terms before plotting.

Examples

  • "Run PyDESeq2 on treated vs control replicates and highlight the top enriched WikiPathways terms."
  • "Filter DEGs to genes with log2(BaseMean) > 1, auto-select fold-change cutoffs, and create volcano and boxplots."
  • "Generate the ranked gene list for GSEA and plot the enrichment curve for the top pathway."

References

  • Tutorial notebook: [t_deseq2.ipynb](../../omicverse_guide/docs/Tutorials-bulk/t_deseq2.ipynb)
  • Sample featureCounts matrix: [sample/counts.txt](../../sample/counts.txt)
  • Quick copy/paste commands: [reference.md](reference.md)

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

它属于哪个仓库

星标★ 3,010
本站分层T1
该仓技能数897
原文件路径skills/bulk-deseq2-analysis/SKILL.md

同一个仓库里的其他技能

看这个仓库的全部 897 个技能