scvi-tools
Deep-learning workflows for integration, batch correction, and multimodal single-cell analysis with scVI, scANVI, totalVI, PeakVI, MultiVI, DestVI, veloVI, and sysVI.
- Skill Road
- scvi-tools
Categories
scvi-tools is a focused skill from Anthropic's open knowledge-work-plugins repository and its bio-research plugin. It makes the principal probabilistic deep-learning models in the scvi-tools ecosystem available as a repeatable analysis guide. The official SKILL.md describes how to choose between scVI, scANVI, totalVI, PeakVI, MultiVI, DestVI, veloVI, and sysVI according to the data type and scientific question. This catalog entry follows that primary source together with the official scvi-tools documentation. The repository is released under Apache-2.0. GitHub stars are not stored for this skill because the Skill model has no field for them; popularity of the parent repository is therefore not evidence of quality for this individual skill.
Intended use
According to the provider, the skill helps researchers use deep-learning single-cell methods when datasets must be integrated, batch effects corrected, cell labels transferred, or several modalities modeled together. For unsupervised scRNA-seq, the model overview points to scVI for integration, differential expression, and latent representations. When cell labels are available, scANVI supports semi-supervised integration and label transfer. totalVI targets CITE-seq with RNA and protein measurements, PeakVI targets scATAC-seq, MultiVI targets joint RNA and ATAC multiome data, and DestVI targets spatial transcriptomics deconvolution with an scRNA reference. veloVI addresses RNA velocity and transcriptional dynamics, while sysVI is intended for system-level or cross-technology batch effects. scArches is discussed in the skill context for mapping new data to pretrained references.
Workflow and data requirements
The workflow starts by checking and preparing an AnnData dataset. A central technical boundary is that scvi-tools models require raw count data represented as integers. Normalized values must not accidentally replace counts; the official guidance recommends preserving raw data in a layer before normalization and specifying that layer when setting up the model. For integration, the documentation calls for an appropriate batch_key so the model can distinguish technical or biological groups. Highly variable gene selection is treated as an important preparation step. The skill points readers to the official documentation for exact APIs, current version information, and model-specific requirements instead of promising one permanent parameter combination.
From dataset to interpretation
The workflow guidance covers validation, preparation, model training, embeddings, clustering, differential expression, label transfer, and dataset integration. This does not automatically prove biological truth. A latent space can reduce batch structure, but it can also weaken genuine biological differences when metadata, study design, or model selection is unsuitable. Results should be checked against known markers, independent quality controls, and the study question. For spatial deconvolution, protein denoising, RNA velocity, or reference mapping, the properties of the reference data are especially important. The provider recommends validating the workflow in the user's own environment before production use.
Boundaries, safety, and operations
The skill provides analysis guidance and points to Python-based tools; it does not replace statistical advice, laboratory validation, or inspection of raw-data quality. Required packages, hardware, GPU drivers, and model versions must match the local environment. Sensitive omics data should be processed only in an approved workspace. Anyone using a hosted Claude client must also check which inputs are transmitted to the relevant model provider; local Python computation and data processing by a connected assistant are not the same thing. Unknown files, metadata, or generated prompts should not be accepted as trusted analysis instructions without review. The skill is a good fit for bioinformatics teams, core facilities, and researchers who want model selection and recurring preparation steps to be more documented and reproducible. It is less suitable when the task is exclusively bulk RNA-seq, simple exploration of already prepared plots, or a fully validated specialist pipeline that must not be adapted.
- Provider
- Anthropic
- License
- Apache-2.0
- Last reviewed
- 09.09.2026
Repository and documentation
Categories
Compatible with
Related guides
Guides and background related to this entry.
Set up Mapbox MCP Server
Set up the Mapbox MCP Server: hosted endpoint or local token, a first test, and sensible limits.
30.09.2026
Set up the Elastic Agent Builder MCP Server
Enable Agent Builder in Kibana, configure tools, and securely connect the built-in MCP endpoint to an AI client via API key or OAuth 2.1.
20.09.2026
Set up the Searchcraft MCP Server
Start Set up the Searchcraft MCP Server with verified links, minimal permissions, and a safe first test.
19.09.2026
Set up Redis MCP Server safely
Configure Redis MCP locally, scope ACL rights, and encrypt the connection.
18.09.2026