nextflow-development
Guided nf-core analyses for RNA-seq, WGS/WES, and ATAC-seq with Nextflow, reproducible test profiles, and verifiable outputs.
- Skill Road
- nextflow-development
Categories
nextflow-development is a focused skill from Anthropic's open knowledge-work-plugins repository and its bio-research plugin. The official source is https://github.com/anthropics/knowledge-work-plugins/tree/main/bio-research/skills/nextflow-development. According to the provider, it is intended for researchers and bench scientists who need to run substantial omics analyses with established nf-core pipelines without assembling every bioinformatics step from scratch. It combines Claude-guided decisions with Nextflow and the rnaseq, sarek, and atacseq pipelines. This catalog entry records the provider's stated scope; it is not a guarantee that any scientific result is accurate.
Purpose and workflow
The skill organizes an analysis into explicit, inspectable stages. For public datasets from GEO or SRA, it starts by collecting study information and deciding which subset of samples is actually needed. For local FASTQ files, that acquisition stage is skipped. It then checks the environment, selects a pipeline with the user, runs a small test profile, creates or validates a samplesheet, confirms the reference genome, and only then starts the analysis on real data. At the end, it checks important output files and the Nextflow log. This sequence reduces avoidable failed runs and exposes assumptions before substantial compute or data transfer begins.
Supported analysis classes
For RNA-seq, the source points to nf-core/rnaseq and gene-expression analysis. For whole-genome or whole-exome sequencing, nf-core/sarek is used for variant analysis. For ATAC-seq, nf-core/atacseq addresses chromatin-accessibility analysis. The source includes versioned pipeline examples while also requiring the user to confirm the pipeline, version, reference genome, and relevant options. That confirmation matters because interpretation depends on organism, experimental design, read configuration, and alignment choices being consistent. The skill can detect data types and validate samplesheet structures, but it does not replace experimental design or a domain expert's judgment.
Reproducibility and outputs
A central practice is to require a successful test profile before processing real data. Nextflow can resume work from prior execution state; the skill presents that as a recovery and reproducibility mechanism, not as evidence that a failed run produced valid science. Expected outputs include merged gene-expression and TPM tables for RNA-seq, variant files and recalibrated sequence data for Sarek, and peak calls plus coverage tracks for ATAC-seq. MultiQC reports and log messages provide useful completion checks. Research teams should archive the raw-data identifiers, samplesheet, pipeline version, reference genome, configuration, and resulting reports together so that later review can distinguish a reproducible run from an undocumented ad hoc execution.
Requirements and safety boundaries
The reference assumes a working environment containing Nextflow, Java, and a container or HPC runtime. Docker is used as one example profile, while Singularity or Apptainer may be relevant on a cluster. These requirements must be checked in the user's environment because the skill does not supply software or compute capacity. GEO and SRA data may be large and are subject to their own terms of use. Genomic datasets can contain information that requires careful handling; permissions, storage locations, network access, and any model sharing must therefore be governed by the research team. An automatically generated samplesheet and a variant call should never be accepted as a scientific conclusion without review, controls, and appropriate quality assessment.
Positioning and audience
The skill is useful for bench scientists, early-career bioinformaticians, and research groups seeking a structured and repeatable entry point for common sequencing analyses. It is a weaker fit for validated bespoke workflows, unusual reference assemblies, or advanced pipeline development. According to the provider, this is a prototype example rather than a production-ready guarantee. The nf-core pipelines, Nextflow, reference genomes, and public-data services each have their own documentation, versions, and citation requirements. Anthropic also states that this integration is not officially endorsed by or affiliated with the nf-core community. Results should therefore be supported by expert quality control, suitable experimental controls, and relevant primary literature before publication or clinical interpretation.
- Provider
- Anthropic
- License
- Apache-2.0
- Last reviewed
- 09.09.2026
Repository and documentation
Categories
Compatible with
Related guides
Guides and background related to this entry.
Set up Mapbox MCP Server
Set up the Mapbox MCP Server: hosted endpoint or local token, a first test, and sensible limits.
30.09.2026
Set up the Elastic Agent Builder MCP Server
Enable Agent Builder in Kibana, configure tools, and securely connect the built-in MCP endpoint to an AI client via API key or OAuth 2.1.
20.09.2026
Set up the Searchcraft MCP Server
Start Set up the Searchcraft MCP Server with verified links, minimal permissions, and a safe first test.
19.09.2026
Set up Redis MCP Server safely
Configure Redis MCP locally, scope ACL rights, and encrypt the connection.
18.09.2026