data-context-extractor

Extracts company-specific data knowledge and turns it into a maintainable analysis skill with reference files.

data-context-extractor is an individual skill from Anthropic's official open-source knowledge-work-plugins repository. Its direct source is https://github.com/anthropics/knowledge-work-plugins/tree/main/data/skills/data-context-extractor. The skill is not a database, an independent MCP server, or a ready-made warehouse connection. It provides a structured working method that lets Claude collaborate with data analysts to capture the company-specific knowledge behind a data warehouse and document it as a reusable analysis guide. That distinction matters: the value comes from careful interviews, verified schema details, and clear reference files, not from an automatic guarantee that every statement is true.

Purpose and operating modes

According to the provider, the instructions work in Bootstrap Mode and Iteration Mode. Bootstrap Mode creates a new data analysis skill from scratch. Typical triggers include a request to create a data context skill for a warehouse, set up data analysis for a company database, or generate a data skill for a particular organization. Iteration Mode loads an existing skill and adds missing context about a domain, metrics, tables, or terminology. This makes the entry useful for an initial knowledge-capture project and for ongoing maintenance as schemas and business definitions change.

Source and schema discovery

The workflow starts by identifying the database system in use. The official instructions name BigQuery, Snowflake, PostgreSQL or Redshift, and Databricks as common options. The next step is to use the schema and query tools available in the current client to identify datasets, schemas, and the tables that matter most to the team. The instructions explicitly ask which three to five tables analysts query most often instead of attempting to document an entire warehouse indiscriminately. SQL examples provide orientation for dialect selection. They are not a request to execute production database commands without review, authorization, or an approved client.

Company knowledge as a reference base

A central focus is knowledge that ordinary data catalogs often omit. Claude asks what terms such as user, customer, or organization mean in the business and how those entities relate. It then captures primary and business identifiers, alternate keys, important metrics, calculation formulas, and time-period conventions. Standard exclusions are equally important: analysts can document how to filter test data, internal users, fraud, or deleted records. Questions about mistakes made by new analysts surface timezone issues, NULL behavior, historical tables, ambiguous column names, and other recurring traps. The intended result does not merely describe tables; it preserves the local language and actual analytical conventions of a team.

Structure and deliverable

According to the provider, the skill creates a directory structure with a central SKILL.md file and references for entities, metrics, tables, and optional dashboards. An entity document can record its definition, primary table, identifiers, relationships, and common filters. A metrics document explains meaning, formula, source tables, and edge cases. Table documents should cover location, purpose, primary key, refresh frequency, important columns, relationships, and common queries. In Iteration Mode, new domain references are added and the navigation in the existing skill is updated. The resulting collection can then be packaged as a ZIP file for delivery.

Boundaries, privacy, and fit

This skill can collect knowledge and structure documentation, but it does not automatically validate business logic. Analysts still need to review definitions, permissions, joins, and results. A plausible metric description is not a substitute for data governance or accountable approval. Warehouse access depends on the selected client and its tools; the skill itself grants no permissions. Schema extracts, metrics, and sample queries may contain sensitive company information. Use approved data only, minimize the context, review retention and access controls, and never place credentials in prompts or reference files. Anthropic's official work on context engineering explains that context made of system instructions, tools, external data, and conversation history must be managed deliberately. That perspective complements the skill's workflow but does not replace review of an organization's policies. The repository is licensed under Apache-2.0. This description was checked against the official source on September 9, 2026 and is not evidence that any generated company documentation is correct.

Free
Provider
Anthropic
License
Apache-2.0
Last reviewed
09.09.2026

Repository and documentation

Categories

Compatible with

Claude Code