跳转到内容

快速开始

所有 /api/v1 调用都采用 PAT(Personal Access Token)的 Bearer 认证。可以在控制台签发,也可以用已登录会话(JWT)通过 API 签发。

Terminal window
curl -X POST https://d2b.dev/api/v1/me/tokens \
-H "Authorization: Bearer $JWT" -H "Content-Type: application/json" \
-d '{"name": "my-agent", "scopes": ["workbooks:read", "workbooks:write"]}'
# → {"token": "d2b_pat_...", ...} (the token is shown in plaintext only in this response)
  • Scope 是资源×操作:workbooks:read/write/delete,cloud-files:read(列出 / 导入 Drive / OneDrive — 登录默认包含,PAT 需显式指定)、governance:configure、workspaces:read/create/configure/delete、keys:read/mint/revoke、account:read、billing:manage。蕴含只在同一资源内(write ⊃ read);删除工作簿、恢复快照、回退版本需要 workbooks:delete
  • 令牌始终只绑定一个账户(一个令牌 = 一个账户)。从已登录会话签发时指向默认账户;传 account_id 可改为你的某个开发者账户。GET /api/v1/me 会返回 account_id / account_name
  • resource 用于收窄范围:workbook:<id>(仅该工作簿 — 把令牌交给 Agent 时的推荐做法)或 workspace:<id>(仅该工作区的工作簿;新建也固定在那里)。默认的 account 是账户下的全部。可达范围只由 resource 决定,scope 决定密钥能做什么 — 没有 workspaces:read 的密钥只要 pin 在 account,仍能在其账户内的任何工作区创建。在 workspace_id 里指定不可达的工作区,列表和创建都返回 403
  • GET /api/v1/me/workspaces 按名称列出令牌可达的工作区(is_default = 省略 workspace_id 时的创建位置)。/me 和工作簿列表也带有 workspace_name
  • 工作区创建是 workspaces:create,密钥签发是 keys:mint,计费操作是 billing:manage — 权限是相互独立的资源×操作 scope,没有伞形 scope;签发密钥时只携带所需权限(与 GitHub fine-grained PAT 同型)
  • 省略 workspace_id 的创建会落在令牌所属账户最早的工作区(默认账户则是个人工作区)。没有工作区的开发者账户会返回 409 — 请先在控制台(Workspaces)或用 workspaces:create 密钥(POST /api/control/workspaces)创建
Terminal window
pip install d2b-sdk
from d2b import D2BClient
client = D2BClient(api_key="d2b_pat_...", base_url="https://d2b.dev")
# 1) A workbook is the working container
wb = client.workbooks.create(title="monthly-sales")["id"]
# 2) Throw the messy Excel at it. wait=True has the SDK babysit the 202+job.
# After extraction and structuring you get clean, typed tables
result = client.sources.upload(wb, "sales_2026-06.xlsx", wait=True)
# 3) See what landed (the basic agent move)
for t in client.tables.list(wb):
print(t["name"], t["row_count"])
schema = client.tables.schema(wb, "sales")
# 4) Analyse with SQL (governance applied, read-only)
out = client.query.sql(wb, 'SELECT product, sum(amount) FROM "sales" GROUP BY 1')
# 5) Derived tables are transforms (lineage is preserved)
client.transforms.create(
wb, name="agg/monthly", kind="sql",
template='CREATE OR REPLACE TABLE "{{ artifact_name }}" AS '
'SELECT product, sum(amount) AS revenue FROM "{{ src }}" GROUP BY 1',
artifact_name="product_sales", args={"src": "sales"},
)
# 6) Back to humans
xlsx = client.export.tables(wb, tables=["product_sales"]) # formatted xlsx
original = client.sources.render_template(wb, "sales_2026-06.xlsx") # original formatting, values refreshed
# 7) Pin a version
client.versions.commit(wb, "2026-06")
Terminal window
BASE=https://d2b.dev; H="Authorization: Bearer $D2B_PAT"
WB=$(curl -s -X POST $BASE/api/v1/workbooks -H "$H" -H "Content-Type: application/json" \
-d '{"title": "monthly"}' | jq -r .id)
curl -s -X POST $BASE/api/v1/workbooks/$WB/sources -H "$H" -F file=@sales.xlsx -F mode=auto -F async=true
# → {"job_id": ...} → poll GET $BASE/api/v1/jobs/{job_id}
curl -s $BASE/api/v1/workbooks/$WB/tables -H "$H"
curl -s -X POST $BASE/api/v1/workbooks/$WB/query -H "$H" -H "Content-Type: application/json" \
-d '{"sql": "SELECT count(*) FROM \"sales\""}'
Terminal window
pipx install d2b-sdk # or uvx --from d2b-sdk d2b …
d2b login
d2b workbooks create --title monthly
d2b upload sales.xlsx --workbook WB --wait
d2b query 'SELECT count(*) FROM "sales"' --workbook WB

下一步: 核心概念、通过 MCP 连接、用 git 管理 workbook。

杂乱的 Excel — 何时用 auto,何时用 staged

Section titled “杂乱的 Excel — 何时用 auto,何时用 staged”

先用默认的 mode=auto(D2B 把每张工作表切分成各个表,并读取合并表头、单位行、小计行和层级来整理表)。只有当提取或结构化不符合预期时才降级到 staged:

  1. 用 mode=staged 重新上传同一文件(仅保存字节、零解释)
  2. POST .../sources/{name}/analyze → 检查并修正返回的 parse spec(表头行、数据范围、类型)
  3. POST .../sources/{name}/materialize 确定

判断流程:auto → 目视结果 → 若有偏差,用 staged 掌控 parse spec 复现。auto 的结果以独立名称保留,可对照修正。

上传响应的 structuring 字段明示结果:structured(已结构化)/ skipped(指定了 structuring=skip — 仅原始表)/ deferred(指定了 structuring=defer — 稍后换入)/ raw_fallback(请求了结构化但返回的是原始表;原因见 structuring_reason — no_credits 为额度不足,building 为构建中,重新上传即可重试,failed:… 为失败;也可以改用 staged)/ off(此服务器已关闭结构化)。CLI 在 raw_fallback 时向 stderr 输出警告。

通过 API / CLI / SDK / MCP 上传时,默认也由 D2B 进行结构化(structuring=auto):工作表被切分为各个表及其周围的标题和注释(<工作表>_補足情報),切出的表按其在工作表上的排列原样保留。在此基础上,期间横排在列上的表被整理为长表,含小计和合计的报表按层级拆分为每层一张表(<表>_階層1、<表>_階層2 …;由同一层级的行计算出的合计和差额放入 <表>_階層1_計算項目 等)以及科目树(<表>_科目)。只有一张表且周围没有其他内容的工作表不会被切分,整理后的表接管文件名。原始工作表保留在血缘中。结构化会消耗工作簿的额度(再次上传相同内容时从缓存返回,不收取结构化费用)。如果想保留原始表、用自己的 LLM 和 add_transform 整形,请指定 --no-structuring(API:structuring=skip),工作表会原样落地,约 1 秒且免费。--defer-structuring(structuring=defer)立即返回原始表,在后台进行结构化,完成后换入结构化表(换入时触发 artifact.updated)。