single-cell-rna-qc

Automated quality control for single-cell RNA-seq data following scverse best practices: MAD-based cell filtering, QC metrics, and before/after visualizations for .h5ad and .h5 files.

single-cell-rna-qc is a single agent skill from Anthropic's open-source knowledge-work-plugins repository (https://github.com/anthropics/knowledge-work-plugins), specifically from the bio-research plugin bundle. It packages a full quality-control workflow for single-cell RNA sequencing (single-cell RNA-seq, scRNA-seq) data and follows the best practices of the scverse ecosystem built around scanpy and AnnData. The authoritative description lives in SKILL.md under bio-research/skills/single-cell-rna-qc; the whole repository is Apache-2.0 licensed and openly readable.

What the skill automates

Per the provider, the skill accepts raw data as .h5ad files (AnnData from Python workflows) or .h5 files (10x Genomics Cell Ranger output) and detects the format automatically. It computes the standard per-cell QC metrics: total counts (count depth), the number of detected genes, and the fraction of mitochondrial, ribosomal, and hemoglobin genes. It then flags problematic cells, removes genes seen in only a handful of cells, and produces before/after visualizations with the applied thresholds drawn in. The outputs are a cleaned dataset ready for downstream analysis plus a copy of the original data with QC annotations attached.

MAD-based filtering instead of fixed cutoffs

The core of the skill is outlier detection via median absolute deviation (MAD) rather than rigid cutoffs such as "at least 500 genes". According to the bundled reference, fixed thresholds fail because protocol, tissue, and species produce very different value ranges. MAD-based bounds instead adapt to the distribution of the dataset at hand. The defaults are deliberately permissive so that rare cell populations are not discarded by accident; an additional hard cutoff is applied to mitochondrial percentage. All thresholds, the gene-name patterns for other species, and the MAD multipliers can be adjusted through command-line parameters.

Two ways to work

For standard cases, the convenience script scripts/qc_analysis.py runs the entire pipeline in one call. For non-standard needs – a different order of steps, metrics only without filtering, different thresholds per cell type – the skill exposes modular functions from qc_core.py and qc_plotting.py, such as calculate_qc_metrics, detect_outliers_mad, filter_cells, or plot_qc_distributions. The reference file references/scverse_qc_guidelines.md explains each metric, the rationale behind the thresholds, and how to read the plots (histograms, violin plots, and scatter plots).

Requirements and limits

The skill needs a Python environment with anndata, scanpy, scipy, matplotlib, seaborn, and numpy. Anthropic states that the repository's plugins are built primarily for Claude Cowork but explicitly names Claude Code as a compatible environment as well. As with every skill in the collection, the provider notes that actual behavior may differ and that the workflow should be tested in your own environment before production use. The skill covers quality control only; steps such as ambient RNA correction, doublet detection, normalization, and batch correction are not included and are mentioned in the reference only as typical follow-up work.

Installation

The plugin bundle is added through the Claude marketplace: first claude plugin marketplace add anthropics/knowledge-work-plugins, then claude plugin install bio-research@knowledge-work-plugins. After that the skill activates automatically whenever a request for QC, cell filtering, or scverse-style preprocessing is recognized. Alternatively, the bundle can be installed straight from Claude Cowork via claude.com/plugins.

Who it is for

single-cell-rna-qc is useful for bioinformatics teams, core facilities, and researchers who want to preprocess scRNA-seq datasets reproducibly and against recognized standards without rewriting QC code every time. Anyone with a firmly established in-house QC workflow, or with non-standard filtering logic, can reach for the modular building blocks instead of the full pipeline. It is not meant for tasks without a single-cell dimension, such as bulk RNA-seq or plain visualization of already-filtered data. The skill itself carries no cost; it is open source and free to use, with only the chosen Claude environment and local compute time to account for.

Free
Provider
Anthropic
License
Apache-2.0
Last reviewed
09.09.2026

Repository and documentation

Categories

Compatible with

Claude Code