DataChain

A Python library for typed, versioned datasets over S3, GCS, and Azure, with a datachain skill install command that teaches Claude Code, Cursor, and Codex your

Skill

Medium risk

We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.

Why this level

  • Runs pipelines and accesses S3, GCS, or Azure buckets using your credentials
  • Generates and runs Python code to process datasets
All reasons and checks

datachain-ai/datachain

Install

In your terminal, with SkillFoxx CLI

npx skillfoxx add skills/datachain

Detects the agents on your machine, checks the risk and pins the version.

Other ways to install

Assembled automatically, review before installing.

Run in a terminal in the project folder

npx skills add datachain-ai/datachain --skill core knowledge -a claude-code -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit 0509279
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/datachain-ai/datachain.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /src/datachain/skill/core/ /src/datachain/skill/knowledge/
git -C "$tmp" checkout 0509279ffd26a187ae9525bf039cc57d88090984
mkdir -p .claude/skills
cp -R "$tmp/src/datachain/skill/core" .claude/skills/core
cp -R "$tmp/src/datachain/skill/knowledge" .claude/skills/knowledge

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

npx skills add datachain-ai/datachain --skill core knowledge -a cursor -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit 0509279
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/datachain-ai/datachain.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /src/datachain/skill/core/ /src/datachain/skill/knowledge/
git -C "$tmp" checkout 0509279ffd26a187ae9525bf039cc57d88090984
mkdir -p .agents/skills
cp -R "$tmp/src/datachain/skill/core" .agents/skills/core
cp -R "$tmp/src/datachain/skill/knowledge" .agents/skills/knowledge

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

npx skills add datachain-ai/datachain --skill core knowledge -a github-copilot -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit 0509279
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/datachain-ai/datachain.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /src/datachain/skill/core/ /src/datachain/skill/knowledge/
git -C "$tmp" checkout 0509279ffd26a187ae9525bf039cc57d88090984
mkdir -p .github/skills
cp -R "$tmp/src/datachain/skill/core" .github/skills/core
cp -R "$tmp/src/datachain/skill/knowledge" .github/skills/knowledge

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

npx skills add datachain-ai/datachain --skill core knowledge -a codex -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit 0509279
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/datachain-ai/datachain.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /src/datachain/skill/core/ /src/datachain/skill/knowledge/
git -C "$tmp" checkout 0509279ffd26a187ae9525bf039cc57d88090984
mkdir -p .agents/skills
cp -R "$tmp/src/datachain/skill/core" .agents/skills/core
cp -R "$tmp/src/datachain/skill/knowledge" .agents/skills/knowledge

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

npx skills add datachain-ai/datachain --skill core knowledge -a gemini-cli -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit 0509279
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/datachain-ai/datachain.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /src/datachain/skill/core/ /src/datachain/skill/knowledge/
git -C "$tmp" checkout 0509279ffd26a187ae9525bf039cc57d88090984
mkdir -p .agents/skills
cp -R "$tmp/src/datachain/skill/core" .agents/skills/core
cp -R "$tmp/src/datachain/skill/knowledge" .agents/skills/knowledge

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/datachain-ai/datachain.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /src/datachain/skill/core/ /src/datachain/skill/knowledge/
git -C "$tmp" checkout 0509279ffd26a187ae9525bf039cc57d88090984
mkdir -p .devin/skills
cp -R "$tmp/src/datachain/skill/core" .devin/skills/core
cp -R "$tmp/src/datachain/skill/knowledge" .devin/skills/knowledge

Commands for macOS and Linux, on Windows run them in Git Bash.

Formerly Windsurf.

Run in a terminal in the project folder

npx skills add datachain-ai/datachain --skill core knowledge -a cline -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit 0509279
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/datachain-ai/datachain.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /src/datachain/skill/core/ /src/datachain/skill/knowledge/
git -C "$tmp" checkout 0509279ffd26a187ae9525bf039cc57d88090984
mkdir -p .cline/skills
cp -R "$tmp/src/datachain/skill/core" .cline/skills/core
cp -R "$tmp/src/datachain/skill/knowledge" .cline/skills/knowledge

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

npx skills add datachain-ai/datachain --skill core knowledge -a roo -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit 0509279
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/datachain-ai/datachain.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /src/datachain/skill/core/ /src/datachain/skill/knowledge/
git -C "$tmp" checkout 0509279ffd26a187ae9525bf039cc57d88090984
mkdir -p .roo/skills
cp -R "$tmp/src/datachain/skill/core" .roo/skills/core
cp -R "$tmp/src/datachain/skill/knowledge" .roo/skills/knowledge

Commands for macOS and Linux, on Windows run them in Git Bash.

A fork of Roo Code, same .roo folders.

Run in a terminal in the project folder

npx skills add datachain-ai/datachain --skill core knowledge -a opencode -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit 0509279
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/datachain-ai/datachain.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /src/datachain/skill/core/ /src/datachain/skill/knowledge/
git -C "$tmp" checkout 0509279ffd26a187ae9525bf039cc57d88090984
mkdir -p .agents/skills
cp -R "$tmp/src/datachain/skill/core" .agents/skills/core
cp -R "$tmp/src/datachain/skill/knowledge" .agents/skills/knowledge

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/datachain-ai/datachain.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /src/datachain/skill/core/ /src/datachain/skill/knowledge/
git -C "$tmp" checkout 0509279ffd26a187ae9525bf039cc57d88090984
mkdir -p .agents/skills
cp -R "$tmp/src/datachain/skill/core" .agents/skills/core
cp -R "$tmp/src/datachain/skill/knowledge" .agents/skills/knowledge

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/datachain-ai/datachain.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /src/datachain/skill/core/ /src/datachain/skill/knowledge/
git -C "$tmp" checkout 0509279ffd26a187ae9525bf039cc57d88090984
mkdir -p .agents/skills
cp -R "$tmp/src/datachain/skill/core" .agents/skills/core
cp -R "$tmp/src/datachain/skill/knowledge" .agents/skills/knowledge

Commands for macOS and Linux, on Windows run them in Git Bash.

You will need: Node.js

Checked against the repository on Sep 26, 2026, commit 0509279.

Text for your agent

Install the library: pip install datachain. Add the agent skill: datachain skill install --target claude (swap claude for cursor, codex, copilot, or pi for the target agent). After that the agent can use DataChain's knowledge base and code generation for your datasets.

Other ways from the author
pip install datachain

Installs the datachain library and CLI

This is third-party code. Review the repository files before installing.

What it does

DataChain turns files in cloud storage into typed, versioned datasets: a compute engine processes files in parallel with async I/O and checkpoint recovery, while a dataset DB holds Pydantic schemas, versions, and lineage with fast filter, join, and group_by, including vector search over the same rows with no separate store. On top sits a knowledge base: markdown summaries of datasets, readable by both humans and agents. The datachain skill install command adds an agent skill that wires the knowledge base and code generation into Claude Code, Cursor, Codex, GitHub Copilot, and Pi, and on the hosted Studio agents reach the same datasets over MCP.

Who it is for. For data engineers and developers who want an agent to find, filter, and process files in cloud storage instead of one-off scripts.

Good fit when

  • You need an agent to build a pipeline over files in S3, GCS, or Azure and save the result as a reusable dataset
  • You need versioning and lineage for large file sets, not just a list of paths
  • The agent needs vector search over the same datasets without a separate store

Not a fit when

  • Your data does not live in cloud file storage and is already structured in a regular database
  • You only need simple vector search without versioning or lineage

Example request

Find dog photos in s3://bucket/oxford-pets/ similar to this image, exclude ones without a mask and Cocker Spaniels, and keep only images wider than 400 pixels

Limitations

The skill builds a knowledge base from datasets already computed with DataChain, so other data needs to be processed first. Distributed compute and agent access to datasets over MCP work only on the hosted Studio, locally only the single-machine compute engine is available. Cloud storage credentials are required.

How to disable. Remove the installed skill folder from the agent's skills directory (for example ~/.claude/skills) and, if needed, uninstall the package with pip uninstall datachain.

Security check

  • Runs pipelines and accesses S3, GCS, or Azure buckets using your credentials
  • Generates and runs Python code to process datasets

README in short

The README presents DataChain as a context layer for unstructured data: the library turns files in S3, GCS, and Azure into typed, versioned datasets that can be filtered, joined, and grouped at warehouse speed. For agent workflows there is an optional knowledge base and agent harness, installed with one command for a specific agent. The quickstart example shows an agent building a pipeline from a natural-language request to find similar dog images by breed metadata and masks in S3.

SKILL.md

---
name: datachain-core
description: Use ONLY for abstract DataChain SDK questions, API usage, method signatures, or code patterns, when no specific dataset or bucket is referenced. If the request mentions creating, saving, listing, exploring datasets or buckets, use datachain-knowledge instead.
---

Read {skill_dir}/SDK.md in full before answering DataChain SDK questions or generating DataChain Python code. It holds the SDK rules: API usage, UDF signatures, settings, delta semantics, materialization patterns, saving, exporting. The last section holds the steps that need a local checkout and the dc-knowledge/ knowledge base.

## Scope of this skill

This skill does not own methodology. Decisions about which datasets to build, what scope, what shape, what fields to save, and when to dialogue with the user about layer choices belong to the CAST methodology in the datachain-knowledge skill.

## Before writing any pipeline code

1. If dc-knowledge/index.md exists, read it first.
2. When the task overlaps an existing dataset, read its .md under dc-knowledge/datasets/ for schema, code patterns, and lineage.
3. Check bucket access: anonymous or authenticated, via dc-knowledge/buckets/ or datachain bucket status.

FAQ

Do I have to use the hosted Studio?

No, the compute engine and dataset DB run locally, Studio is only needed for distributed compute and for agents to reach the same datasets over MCP.

How is this different from a plain vector store?

DataChain keeps typed, versioned datasets with lineage and runs vector search over the same rows, with no separate store for embeddings.

Editors’ pick

166 skills for scientific work: bioinformatics, cheminformatics, clinical data, geospatial analysis and 100+ databases

SkillMedium risk47KRepository stars
Official

Google's open-source MCP server for databases: ready tools for Postgres, MySQL, BigQuery, Spanner and more, plus custom tools in tools.yaml

MCP serverHigh risk16.5KRepository stars
Editors’ pick

Official Hugging Face skills: Hub operations via the hf CLI, datasets, model training, Spaces, evals and deployment

SkillHigh risk11.1KRepository stars
Official

A visualization language for agents and an MCP server: neat charts from a simple semantic spec

MCP serverMedium riskNo VPN needed4.3KRepository stars
Foxx AIDataChain

I am Foxx AI and I have already vetted this tool. Ask about install, setup or anything else, and I will keep it simple.