Home/Data Analysis/chdb-datastore
C

chdb-datastore

by @clickhousev
4.4(120)

chdb DataStore is a lazy, ClickHouse-backed pandas replacement that allows users to analyze large-scale data from local files, databases, and cloud storage using the same pandas code. It compiles operations into optimized SQL, significantly improving performance, especially for cross-source joins and groupby aggregations on big datasets.

chdbdatastorepandas-replacementclickhousedata-analysisGitHub
Installation
npx skills add https://github.com/clickhouse/agent-skills --skill chdb-datastore
compare_arrows

Before / After Comparison

1
Before

Users running groupby+sum on 100M rows with pandas often encounter out-of-memory errors or wait over 60 seconds, requiring manual code tweaks or data splitting.

After

This skill compiles pandas API calls into ClickHouse SQL automatically, completing the same task in 10 seconds with minimal memory usage, enabling seamless processing of even larger datasets.

SKILL.md

chdb DataStore — It's Just Faster Pandas

The Key Insight

# Change this:
import pandas as pd
# To this:
import chdb.datastore as pd
# Everything else stays the same.

DataStore is a lazy, ClickHouse-backed pandas replacement. Your existing pandas code works unchanged — but operations compile to optimized SQL and execute only when results are needed (e.g., print(), len(), iteration).

pip install chdb

Decision Tree: Pick the Right Approach

1. "I have a file/database and want to analyze it with pandas"
   → DataStore.from_file() / from_mysql() / from_s3() etc.
   → See references/connectors.md

2. "I need to join data from different sources"
   → Create DataStores from each source, use .join()
   → See examples/examples.md #3-5

3. "My pandas code is too slow"
   → import chdb.datastore as pd — change one line, keep the rest

4. "I need raw SQL queries"
   → Use the chdb-sql skill instead

Connect to Any Data Source — One Pattern

from datastore import DataStore

# Local file (auto-detects .parquet, .csv, .json, .arrow, .orc, .avro, .tsv, .xml)
ds = DataStore.from_file("sales.parquet")

# Database
ds = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")

# Cloud storage
ds = DataStore.from_s3("s3://bucket/data.parquet", nosign=True)

# URI shorthand — auto-detects source type
ds = DataStore.uri("mysql://root:pass@db:3306/shop/orders")

All 16+ sources and URI schemes → connectors.md

After Connecting — Full Pandas API

result = ds[ds["age"] > 25]                                          # filter
result = ds[["name", "city"]]                                        # select columns
result = ds.sort_values("revenue", ascending=False)                  # sort
result = ds.groupby("dept")["salary"].mean()                         # groupby
result = ds.assign(margin=lambda x: x["profit"] / x["revenue"])     # computed column
ds["name"].str.upper()                                               # string accessor
ds["date"].dt.year                                                   # datetime accessor
result = ds1.join(ds2, on="id")                                      # join
result = ds.head(10)                                                 # preview
print(ds.to_sql())                                                   # see generated SQL

209 DataFrame methods supported. Full API → api-reference.md

Cross-Source Join — The Killer Feature

from datastore import DataStore

customers = DataStore.from_mysql(host="db:3306", database="crm", table="customers", user="root", password="pass")
orders = DataStore.from_file("orders.parquet")

result = (orders
    .join(customers, left_on="customer_id", right_on="id")
    .groupby("country")
    .agg({"amount": "sum", "rating": "mean"})
    .sort_values("sum", ascending=False))
print(result)

More join examples → examples.md

Writing Data

source = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")
target = DataStore("file", path="summary.parquet", format="Parquet")

target.insert_into("category", "total", "count").select_from(
    source.groupby("category").select("category", "sum(amount) AS total", "count() AS count")
).execute()

Troubleshooting

ProblemFix
ImportError: No module named 'chdb'pip install chdb
ImportError: cannot import 'DataStore'Use from datastore import DataStore or from chdb.datastore import DataStore
Database connection timeoutInclude port in host: host="db:3306" not host="db"
Join returns empty resultCheck key types match (both int or both string); use .to_sql() to inspect
Unexpected resultsCall ds.to_sql() to see the generated SQL and debug
Environment checkRun python scripts/verify_install.py (from skill directory)

References

Note: This skill teaches how to use chdb DataStore. For raw SQL queries, use the chdb-sql skill. For contributing to chdb source code, see CLAUDE.md in the project root.


chdb DataStore

Agent skill for using chdb's pandas-compatible DataStore API — a drop-in pandas replacement backed by ClickHouse.

Installation

npx skills add clickhouse/agent-skills

What's Included

FilePurpose
SKILL.mdSkill definition and quick-start guide
references/api-reference.mdFull DataStore method signatures
references/connectors.mdAll 16+ data source connection methods
examples/examples.md11 runnable examples with expected output
scripts/verify_install.pyEnvironment verification script

Trigger Phrases

This skill activates when you:

  • "Analyze this file with pandas"
  • "Speed up my pandas code"
  • "Query this MySQL/PostgreSQL/S3 table as a DataFrame"
  • "Join data from different sources"
  • "Use DataStore to..."
  • "Import datastore as pd"

Related

  • chdb-sql — For raw ClickHouse SQL queries, use the chdb-sql skill instead
  • clickhouse-best-practices — For ClickHouse schema/query optimization

Documentation

User Reviews (0)

Write a Review

Effect
Usability
Docs
Compatibility

No reviews yet

Statistics

Installs4.9K
Rating4.4 / 5.0
Version
Updated2026年8月1日
Comparisons1

User Rating

4.4(120)
5
37%
4
43%
3
13%
2
5%
1
2%

Rate this Skill

0.0

Compatible Platforms

🤖claude-code

Timeline

Created2026年7月26日
Last Updated2026年8月1日
🎁 Agent Knowledge Cards
Survey