data-context-extractor
Extracts company-specific data knowledge and turns it into a maintainable analysis skill with reference files.
- Skill Road
- data-context-extractor
Categories
data-context-extractor is an individual skill from Anthropic's official open-source knowledge-work-plugins repository. Its direct source is https://github.com/anthropics/knowledge-work-plugins/tree/main/data/skills/data-context-extractor. The skill is not a database, an independent MCP server, or a ready-made warehouse connection. It provides a structured working method that lets Claude collaborate with data analysts to capture the company-specific knowledge behind a data warehouse and document it as a reusable analysis guide. That distinction matters: the value comes from careful interviews, verified schema details, and clear reference files, not from an automatic guarantee that every statement is true.
Purpose and operating modes
According to the provider, the instructions work in Bootstrap Mode and Iteration Mode. Bootstrap Mode creates a new data analysis skill from scratch. Typical triggers include a request to create a data context skill for a warehouse, set up data analysis for a company database, or generate a data skill for a particular organization. Iteration Mode loads an existing skill and adds missing context about a domain, metrics, tables, or terminology. This makes the entry useful for an initial knowledge-capture project and for ongoing maintenance as schemas and business definitions change.
Source and schema discovery
The workflow starts by identifying the database system in use. The official instructions name BigQuery, Snowflake, PostgreSQL or Redshift, and Databricks as common options. The next step is to use the schema and query tools available in the current client to identify datasets, schemas, and the tables that matter most to the team. The instructions explicitly ask which three to five tables analysts query most often instead of attempting to document an entire warehouse indiscriminately. SQL examples provide orientation for dialect selection. They are not a request to execute production database commands without review, authorization, or an approved client.
Company knowledge as a reference base
A central focus is knowledge that ordinary data catalogs often omit. Claude asks what terms such as user, customer, or organization mean in the business and how those entities relate. It then captures primary and business identifiers, alternate keys, important metrics, calculation formulas, and time-period conventions. Standard exclusions are equally important: analysts can document how to filter test data, internal users, fraud, or deleted records. Questions about mistakes made by new analysts surface timezone issues, NULL behavior, historical tables, ambiguous column names, and other recurring traps. The intended result does not merely describe tables; it preserves the local language and actual analytical conventions of a team.
Structure and deliverable
According to the provider, the skill creates a directory structure with a central SKILL.md file and references for entities, metrics, tables, and optional dashboards. An entity document can record its definition, primary table, identifiers, relationships, and common filters. A metrics document explains meaning, formula, source tables, and edge cases. Table documents should cover location, purpose, primary key, refresh frequency, important columns, relationships, and common queries. In Iteration Mode, new domain references are added and the navigation in the existing skill is updated. The resulting collection can then be packaged as a ZIP file for delivery.
Boundaries, privacy, and fit
This skill can collect knowledge and structure documentation, but it does not automatically validate business logic. Analysts still need to review definitions, permissions, joins, and results. A plausible metric description is not a substitute for data governance or accountable approval. Warehouse access depends on the selected client and its tools; the skill itself grants no permissions. Schema extracts, metrics, and sample queries may contain sensitive company information. Use approved data only, minimize the context, review retention and access controls, and never place credentials in prompts or reference files. Anthropic's official work on context engineering explains that context made of system instructions, tools, external data, and conversation history must be managed deliberately. That perspective complements the skill's workflow but does not replace review of an organization's policies. The repository is licensed under Apache-2.0. This description was checked against the official source on September 9, 2026 and is not evidence that any generated company documentation is correct.
- Provider
- Anthropic
- License
- Apache-2.0
- Last reviewed
- 09.09.2026
Repository and documentation
Categories
Compatible with
Related guides
Guides and background related to this entry.
Set up Mapbox MCP Server
Set up the Mapbox MCP Server: hosted endpoint or local token, a first test, and sensible limits.
30.09.2026
Set up the Elastic Agent Builder MCP Server
Enable Agent Builder in Kibana, configure tools, and securely connect the built-in MCP endpoint to an AI client via API key or OAuth 2.1.
20.09.2026
Set up the Searchcraft MCP Server
Start Set up the Searchcraft MCP Server with verified links, minimal permissions, and a safe first test.
19.09.2026
Set up Redis MCP Server safely
Configure Redis MCP locally, scope ACL rights, and encrypt the connection.
18.09.2026