Sato Hub published two datasets on Hugging Face on 3 October 2026. Both are v1.0, dated 2 October, licensed CC-BY-4.0, data by satohub.ai. One tests whether a model knows how to build an onchain agent. The other teaches a model when to call our MCP tools and, just as important, when not to.
Run the bench
Onchain Agent Builder Bench is 367 multiple-choice questions, four options each, one right answer. Any harness can score it. It covers five families:
- ▸Stack selection (77): which layers and standards a build goal needs. The answer is a set of categories, never a product or a ranking.
- ▸Protocol facts (55): what ERC-8004, x402, ERC-4337, ERC-7715, MCP, A2A and Solana's SATI are for, and what they are not for.
- ▸Custody reading (54): which reading a generic package description supports: takes your key, your key leaves, can move funds on its own, or normal wallet behaviour. The wording describes. It never judges.
- ▸Glossary (148): onchain-agent vocabulary, term to definition.
- ▸Scope (33): what an assistant should decline, and what it should answer from dated evidence.
Task files ship in the dataset repository under eval/, for both lm-evaluation-harness and Inspect AI:
``
lm_eval --model hf --model_args pretrained=<model> --tasks satohub_onchain_agent_builder_bench --include_path ./eval
inspect eval eval/inspect_task.py --model <provider/model>
``
Or load it directly: load_dataset("SatoHub/onchain-agent-builder-bench", "all", split="validation"). Every row carries a rationale and a source_url pointing at the satohub.ai page it rests on, plus a canary string. If a model reproduces the canary, the file was in its training data.
Train tool use
Sato MCP Tool-Calling is 1,539 synthetic conversations in which an assistant uses Sato Hub's read-only MCP tools to answer questions about building onchain agents. The same conversations ship in three formats, so you can drop them into the pipeline you already have:
- ▸
messages: OpenAI chat format withtool_calls. - ▸
hermes: the NousResearch Hermes function-calling columns. - ▸
xlam: the Salesforce xlam columns, first call or calls only.
Every tool call was checked against the live parameter schema (required fields, no unknown keys, enums, types) before an example was kept. Answers that came from Sato Hub end with Source: <satohub.ai page>.
The negatives
263 of the 1,539 conversations (17%) are cases where calling Sato Hub is the wrong move:
- ▸Token prices, live balances, transaction status: use a market-data, balance or status tool instead. The examples offer one, with invented values labelled as such.
- ▸Trading or investment advice: decline, no tool.
- ▸General crypto explainers: answer directly, no tool.
- ▸Verdicts on a package: decline, and say what to check.
A model that calls a tool for every crypto question is not using tools well. The negatives are there to teach the other half.
The held-out split
The bench has a held-out test split of 91 questions. It is not published, so no model can be tuned on it. What is published is its SHA-256:
58a1b8d460445a6fc72739c588a0f17ed09911cc155c780ffed512c59d56c1a0
If the split ever changes, that hash will no longer match. That is the check. New versions are cut deliberately, tagged, and old ones stay available.
What is not in the datasets
No Sato Score, tiers or score inputs. No history or snapshots. No install steps or deploy specs. No Sato Check readings of what code does with keys. No success rates from our checks. No figures about Sato Hub's own usage. The stack questions never rank or recommend a project, and the custody scenarios are generic: they name no real package. Tool outputs in the tool-calling set are frozen, thin fixtures, a seeded sample and not the engine's ranking. For the current directory, use the live site and the MCP server. Sato Hub has no token.
Baseline results
We ran five models closed-book on the 367 public questions on 3 October 2026: zero-shot, no tools, no web search, each answer one letter. An answer that is not a single letter counts as wrong.
- ▸Grok 4.7: 99% overall, 97% on stack selection (1 answer not a single letter)
- ▸Gemini 3.8 Flash: 98% overall, 92% on stack selection
- ▸Claude Sonnet 5.5: 96% overall, 84% on stack selection (12 answers began with an explanation instead of a letter and count as wrong)
- ▸Claude Haiku 4.5: 92% overall, 79% on stack selection
- ▸GPT-5 mini: 85% overall, 55% on stack selection
The per-family table, with the exact model ids each API reported, is on the dataset card. Two things stand out. Frontier models are near the ceiling on protocol facts, custody reading, glossary and scope. Stack selection, which asks for the full set of layers and standards a build needs, is where they separate, by more than 40 points from top to bottom. These questions are public, so later models may have seen them; the held-out split is the check on that.
Trivial baselines, read straight off the v1.0 release. Chance is 25% on every family. Accuracy if you always pick the longest option, or always the shortest:
- ▸Stack selection: longest 5%, shortest 35%
- ▸Protocol facts: longest 51%, shortest 7%
- ▸Custody reading: longest 44%, shortest 19%
- ▸Glossary: longest 27%, shortest 25%
- ▸Scope: longest 52%, shortest 2%
Read a model's score against these. Protocol facts, custody reading and scope can be partly solved by option length alone.
Corrections
To report a wrong question or answer, use https://satohub.ai/disputes. Both datasets:
- ▸https://huggingface.co/datasets/SatoHub/onchain-agent-builder-bench
- ▸https://huggingface.co/datasets/SatoHub/sato-mcp-tool-calling