DAIJIN / PROJECT BRAIN
CURRENT STATUS · 23 AUG 2026現在のステータス · 2026年8月23日

DAIJIN

Daijin builds a measured knowledge brain for any repository, distilled from the repo's own code and history. It serves that brain to coding agents over MCP, and runs a certification gym where a student model attempts real exams mined from the repo's history, a teacher model grades them against five cited criteria, and harvested lessons feed back into the brain. An active RAG loop that measures what it serves, and improves it.

Daijin はどんなリポジトリにも、そのコードと履歴から蒸留した「プロジェクトブレイン」を構築するツールです。計測済みのブレインを MCP 経由でコーディングエージェントに提供し、さらに認定ジムでは、リポジトリの実履歴から採掘した試験に学生モデルが挑み、教師モデルが5つの根拠付き基準で採点し、収穫した教訓をブレインに書き戻します。提供する知識を計測し、改善し続ける能動的な RAG ループです。

WHAT IS SERVED何が提供されるかbrain.search, brain.conventions, brain.impact_of, brain.relevant_adrs, served over MCP once retrieval clears the measured floor.brain.search / brain.conventions / brain.impact_of / brain.relevant_adrs の4ツール。検索品質が計測済みフロアを超えたときだけ MCP で提供されます。
SPEND GATE支出ゲートBlocked by default. Retrieval never calls a paid API, and every paid gym run re-blocks the gate the moment it ends.既定でブロック。検索経路は有料 API を一切呼ばず、ジムの有料実行はどれも終了と同時にゲートを再ブロックします。
1 ·LIVE REPO MEASUREMENT実リポジトリでの計測

THE FLOOR, MEASURED ON A REAL REPO実リポジトリで測ったフロア

A live layer-1 init on the owner's TokaiHub repository, measured 2026-08-22 at zero spend. The floor is reported in count form, never as a bare rounded percentage.

オーナーの TokaiHub リポジトリでの実測(2026-08-22、レイヤー1 init、支出ゼロ)。フロアは常に件数表記で報告され、丸めたパーセントだけで語ることはありません。

MCP UNLOCK THRESHOLDMCP 解放しきい値
75%CASE RATE 0.75ケース率 0.75
MEASURED CASE RATE実測ケース率
96%24 OF 25 · TOKAIHUB · 22 AUG 202625件中24件 · TOKAIHUB · 2026-08-22
GOLD CASES FOUNDゴールドケース正答24 / 25ABOVE THE 0.75 UNLOCK解放しきい値 0.75 超えRetrieval answered 24 of the 25 mined gold cases (96%)採掘した25件のゴールドケースのうち24件に正答(96%)
MRR0.777RECORDED, NOT FLOORED記録のみ、フロア対象外Mean reciprocal rank, tracked for movement only平均逆順位。推移の観測にだけ使います
MUST-NOT VIOLATIONS禁止違反0ENFORCED FLOOR強制フロアNo forbidden document surfaced for any gold caseどのケースでも禁止文書は一件も浮上せず
SERVING BUDGET提供トークン予算3000TOKENS, SWEEP-CHOSENトークン、スイープで選定Smallest budget within one case of the best score最良スコアから1ケース以内で最小の予算
MCP SERVINGMCP 提供UNLOCKED解放済みSNIPPET AVAILABLEスニペット発行可The paste-ready MCP config is issued above the floorフロア超えで貼り付け可能な MCP 設定を発行
GOLD-SET FLOOR SWEEP · CLEAN FIXTUREゴールドセット・フロアスイープ · クリーンフィクスチャ25 of 25 at every budget全予算で25件中25件THE BUDGET SWEEP RUNS AT INIT TIME: 3000 / 4000 / 6000 / 8000 TOKENS初期化時に 3000 / 4000 / 6000 / 8000 トークンでスイープ
258025/2525/2525/2525/253000400060008000TOKEN BUDGET PER ANSWER回答あたりトークン予算
A flat curve on the fixture: the sweep picks the smallest budget within one case of the best score, so serving defaults to 3000 tokens.フィクスチャでは曲線がフラット。スイープは最良スコアから1ケース以内で最小の予算を選ぶため、提供予算は 3000 トークンになります。
2 ·VERIFICATION検証

TWO SUITES, ZERO FAILURES2つのテストスイート、失敗ゼロ

Both halves of the product carry their own test suite, run at zero network and zero spend. Counts are from the last verification, 2026-08-23.

エンジンと TUI はそれぞれ独立したテストスイートを持ち、ネットワークも支出もゼロで走ります。件数は最終検証(2026-08-23)時点のものです。

ENGINE (NODE 22, ESM)エンジン(NODE 22, ESM)
802TESTS PASS · 0 FAILテスト成功 · 失敗0
TUI (PYTHON TEXTUAL)TUI(PYTHON TEXTUAL)
453TESTS PASS · 0 FAILテスト成功 · 失敗0
TOTAL TESTSテスト総数1255802 engine + 453 TUI, both suites greenエンジン802 + TUI 453、両スイートともグリーン
FAILURES失敗0Zero failures on either side at last verification最終検証時点で両側とも失敗ゼロ
NETWORK CALLS IN TESTSテスト中の外部通信0The suites run with no network and no provider spendスイートは外部通信なし・プロバイダ支出なしで実行
FROZEN CONTRACTS凍結された契約2store.d.ts and the RPC surface (methods.md v5); a shape gate can fail the buildstore.d.ts と RPC 仕様(methods.md v5)。シェイプゲートがビルドを落とせます
3 ·END TO ENDエンドツーエンド

FIVE LIVE CHAINS, EXIT 0ライブ実行5チェーン、すべて正常終了

The full drive, run live against real providers: mine exams, promote, run the gym, grade, harvest, apply. Five chains across four repos, all exit 0; mining validated 3 of 3 exams on the clean fixture.

試験の採掘 → 昇格 → ジム実行 → 採点 → 収穫 → 適用までの全工程を、実プロバイダ相手にライブで実行。4リポジトリで5チェーン、すべて exit 0。クリーンフィクスチャでは採掘した試験3件中3件が検証を通過しました。

LIVE CHAINSライブチェーン5 exit 0MINE → GRADE → APPLY採掘 → 採点 → 適用Every chain ran the whole loop against live providers全チェーンがライブのプロバイダ相手にループ全体を完走
REPOS DRIVEN対象リポジトリ4NODE, PYTHON, FIXTURES, DAIJIN ITSELFNODE、PYTHON、フィクスチャ、DAIJIN 自身Including a Python repo with a unittest gate and a clone of daijin itselfunittest ゲート持ちの Python リポジトリや daijin 自身のクローンを含む
EXAMS VALIDATED検証済み試験3 / 3ON THE CLEAN FIXTUREクリーンフィクスチャにてMined from real commits, checked in a per-exam worktree at base実コミットから採掘し、試験ごとのワークツリーで基準時点を検証
STUDENT学生GLM-5.3 (zai). Attempts each exam in a sandboxed copy of the repo through a strict JSON action protocol: real edits, real tokens, a real cap.GLM-5.3(zai)。リポジトリのサンドボックスコピー内で、厳格な JSON アクションプロトコルにより試験に挑戦。実際の編集、実際のトークン、実際の上限つき。Never grades its own work; the harness owns running gates.自分の成果は採点しない。ゲート実行はハーネスの権限。
TEACHER教師Claude Opus. Grades each attempt on five cited axes; every citation is bound by digests, and the verdict is capped mechanically.Claude Opus。5つの根拠付き軸で採点。引用はダイジェストで束縛され、判定は機械的に上限が課されます。Never sees the gold commit; never grades work it authored.正解コミットは見えない。自分が作った課題は採点しない。
EXAM COMMITTEE / AUDITOR試験委員 / 監査役Claude Fable. Mines and authors exams from the repo's real commit history, and triages watcher findings with written reasoning.Claude Fable。リポジトリの実コミット履歴から試験を採掘・作成し、ウォッチャーの指摘を根拠付きでトリアージします。Proposes; mechanics accept. No model promotes its own exam.提案するだけで、採否は機構が決める。モデルは自分の試験を昇格できない。
WATCHERウォッチャーGLM-5.3. Verifies each finding in one cheap generation and rides its note ahead of the auditor's.GLM-5.3。各指摘を1回の安価な生成で検証し、監査役より先に所見を添えます。Read-only: assesses, never fixes, never spends on its own.読み取り専用。評価はするが、修正も自発的な支出もしない。

HONEST BOUND: in the live chains the teacher refused to invent a knowledge gap five times, so no live batch carried an accepted lesson yet. That is the design refusing to teach a lie, not a gap in the wiring; the write path is pinned by tests against a real brain.

正直な注記: ライブ実行では教師が「知識ギャップの捏造」を5回拒否したため、承認済み教訓を含むバッチはまだありません。これは配線の欠陥ではなく、嘘を教えることを設計が拒んだ結果です。書き込み経路は実ブレインに対するテストで固定済みです。

4 ·WHAT YOU CAN DO WITH ITできること

ONE TOOL, FIVE ACTS1つのツール、5つの動作

Everything below the gym runs at zero spend; the gym itself sits behind an owner-only gate that is blocked by default.

ジム以外はすべて支出ゼロで動作します。ジム自体はオーナー専用ゲートの背後にあり、既定でブロックされています。

  • Build a brain. Attach any git repo; daijin analyzes it, scaffolds durable markdown knowledge (architecture, conventions, decisions, lessons) in .daijin/brain, and indexes it locally. The index is disposable; the markdown is the truth.
  • ブレインを構築する。git リポジトリを接続すると、daijin が解析し、恒久的な markdown 知識(アーキテクチャ・規約・意思決定・教訓)を .daijin/brain に生成してローカルでインデックス化します。インデックスは使い捨て、markdown が真実です。
  • Measure it. A gold set of retrieval cases is mined from the repo's real commit history, passes its own integrity gates, and scores the brain in count form with a permuted control beside it.
  • 計測する。検索ケースのゴールドセットを実コミット履歴から採掘し、それ自体が整合性ゲートを通過した上で、ブレインを件数表記で採点。並べ替え対照群も常に併記されます。
  • Serve it over MCP. Above the measured floor, daijin hands you a paste-ready MCP snippet: brain.search, brain.conventions, brain.impact_of, brain.relevant_adrs for your coding agents.
  • MCP で提供する。計測済みフロアを超えると、貼り付けるだけの MCP スニペットが発行されます。コーディングエージェントから brain.search / brain.conventions / brain.impact_of / brain.relevant_adrs が使えます。
  • Run the gym. Owner-gated: a student model attempts exams mined from your history in a sandbox, a teacher model grades five cited axes, and an append-only ledger records everything.
  • ジムを走らせる。オーナーゲート制。学生モデルが履歴から採掘した試験にサンドボックス内で挑み、教師モデルが5つの根拠付き軸で採点。追記専用台帳がすべてを記録します。
  • Harvest lessons. Graded gaps become reviewed lesson proposals; applying a batch writes real lesson documents into the brain and reindexes, so retrieval serves what was learned.
  • 教訓を収穫する。採点で見つかったギャップは審査済みの教訓提案になり、適用するとブレインに教訓文書が書き込まれ再インデックス。学んだことが検索で提供されるようになります。
5 ·INSTALLインストール

TWO COMMANDS, ONE PREFIXコマンド2つ、プレフィックス1つ

Everything lands in one self-contained prefix with a single daijin command on your PATH. Installing twice lands in the same state as once.

すべてが自己完結のプレフィックスに収まり、PATH には daijin コマンドが1つ増えるだけ。2回インストールしても1回と同じ状態になります。

FROM SOURCE · ソースから · github.com/MohamedFuad16/daijin
git clone https://github.com/MohamedFuad16/daijin.git
cd daijin
bash install/install.sh
NODE >= 22PYTHON >= 3.10OLLAMA SERVING bge-m3 (EMBEDDINGS)OLLAMA で bge-m3 を配信(埋め込み用)

Then: daijin /path/to/your/repo to connect a real repo, or daijin . --mock to explore the full UI with no engine and no Ollama.

その後は daijin /path/to/your/repo で実リポジトリに接続、または daijin . --mock でエンジンも Ollama もなしに UI 全体を試せます。

6 ·HISTORY経緯

HOW THE STUDENT MODEL WAS CHOSEN学生モデルはこうして選ばれた

A historical benchmark from the predecessor project (July 2026): candidate student models ran the same real-repo exam bank under the five-axis teacher rubric, on equal footing. GLM won on quality per cost, and its successor GLM-5.3 is the student daijin ships with today. The records below are kept verbatim, in their original English.

前身プロジェクトでの歴史的ベンチマーク(2026年7月)。学生モデルの候補たちが、同一の実リポジトリ試験バンクを5軸の教師ルーブリックのもと、同条件で受験しました。コストあたり品質で GLM が勝ち、その後継である GLM-5.3 が現在の daijin の学生です。以下の記録は当時のまま、原文(英語)で保存しています。

TEST 02MODEL COMPARISON · 23 JUL 2026Coding model capability comparisonVIEW RESULTS
MODEL COMPARISONCoding model capability comparison
23 JUL 202627 ranked attempts4 settings3 models
27 saved results · no current winnerOld labels: 1 pass · 5 partial · 20 fail · corrected total $12.14

AUDIT HOLD. These saved results used a grading rule that treated equivalent work differently. They remain visible for history but are not ranked until deterministic regrading.

MODELSETTINGSOUTCOMESTASK SIZESTEACHER AVERAGE5 QUALITY SCORESGRADEDENGINEER USAGEGRADING + TOTAL COST
Kimi K2.7 Code
Kimi K2.7 Codemoonshotai/kimi-k2.7-code
HISTORICAL LABEL · PENDING REGRADE
Reasoning levelProvider defaultPrecisionHost-dependentOutput cap32K tokens / replyTime limit300s
1 PASS1 PARTIAL7 FAIL
SMALL2/4 usefulMEDIUM0/2 usefulLARGE0/3 useful
AVERAGE SCORE2.17 / 51 is weak · 5 is excellent
CORRECT1.89 / 5STYLE2.22 / 5DECISIONS2.33 / 5REASONING2.11 / 5SCOPE2.33 / 5
9/9complete
2.3M total tokensWork used2.3MThinking used32K · 64.6%Answer generated18KWork allowanceby task size · 300K–600K at the time of the run / examThinking allowanceshared with work budget / examAverage time1.7 min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops0Thinking-cap stops0Tool-flow errors16Unknown calls0
$2.81 grading
$3.57 total
$0.40 per task
0 unknown teacher calls
GLM-5.2
GLM-5.2z-ai/glm-5.2
HISTORICAL LABEL · PENDING REGRADE
Reasoning levelThinking on · Max effort (provider default)PrecisionHost-dependentOutput cap32K tokens / replyTime limit300s
0 PASS2 PARTIAL7 FAIL
SMALL2/4 usefulMEDIUM0/2 usefulLARGE0/3 useful
AVERAGE SCORE2.71 / 51 is weak · 5 is excellent
CORRECT1.78 / 5STYLE2.44 / 5DECISIONS3.33 / 5REASONING2.67 / 5SCOPE3.33 / 5
9/9complete
2M total tokensWork used2MThinking used101K · 82.0%Answer generated22KWork allowanceby task size · 300K–600K at the time of the run / examThinking allowanceshared with work budget / examAverage time5.4 min/taskReply-limit stops0Timeout truncations0Provider-error tasks22.2%Work-cap stops0Thinking-cap stops0Tool-flow errors19Unknown calls2
$2.87 grading
$3.71 total
$0.41 per task
0 unknown teacher calls
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
HISTORICAL LABEL · PENDING REGRADE
Reasoning levelProvider defaultPrecisionHost-dependentOutput cap12K tokens / replyTime limit300s
0 PASS2 PARTIAL6 FAIL
SMALL1/4 usefulMEDIUM0/2 usefulLARGE1/3 useful
AVERAGE SCORE2.08 / 51 is weak · 5 is excellent
CORRECT1.63 / 5STYLE2.38 / 5DECISIONS2.38 / 5REASONING2.13 / 5SCOPE1.88 / 5
8/91 ungraded
2.4M total tokensWork used2.4MThinking used27K · 45.4%Answer generated32KWork allowanceby task size · 300K–600K at the time of the run / examThinking allowanceshared with work budget / examAverage time0.6 min/taskReply-limit stops0Timeout truncations0Provider-error tasks22.2%Work-cap stops1Thinking-cap stops0Tool-flow errors22Unknown calls2
$2.20 grading
$2.80 total
$0.31 per task
0 unknown teacher calls
GLM-5.2
GLM-5.2z-ai/glm-5.2
CONTEXT ONLY · NOT CAPABILITY RANKED
Reasoning levelThinking on · Max effort (provider default)PrecisionHost-dependentOutput cap12K tokens / replyTime limit300s
0 PASS0 PARTIAL2 FAIL
SMALL0/2 usefulMEDIUM0/2 usefulLARGE0/1 useful
AVERAGE SCORE1.40 / 51 is weak · 5 is excellent
CORRECT1.50 / 5STYLE2.00 / 5DECISIONS1.00 / 5REASONING1.50 / 5SCOPE1.00 / 5
2/53 ungraded
1.3M total tokensWork used1.3MThinking used132K · 99.6%Answer generated2KWork allowanceby task size · 300K–600K at the time of the run / examThinking allowanceshared with work budget / examAverage time— min/taskReply-limit stops11Timeout truncations0Provider-error tasks0.0%Work-cap stops3Thinking-cap stops0Tool-flow errors15Unknown calls0
$0.46 grading
$1.15 total
$0.23 per task
0 unknown teacher calls
$0.90 aborted/no verdict · 362K tokens
DITHERED MODEL EVIDENCEFour views of the same graded results
Historical labels only. The deterministic regrade will replace these outcome rankings.
BAR · OUTCOME-CHARTOutcome distributionSame 9 exams per setting
6 useful26 graded
03692 usefulKimi9 / 9 gradedKimi K2.7 Code: 1 pass, 1 partial, 7 fail; 9 of 9 graded2 usefulGLM-5.29 / 9 gradedGLM-5.2: 0 pass, 2 partial, 7 fail; 9 of 9 graded2 usefulDeepSeek8 / 9 gradedDeepSeek V4 Pro: 0 pass, 2 partial, 6 fail; 8 of 9 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135Kimi K2.7 Code: Correctness 1.89 / 5, Convention 2.22 / 5, Decisions 2.33 / 5, Reasoning 2.11 / 5, Scope 2.33 / 5GLM-5.2: Correctness 1.78 / 5, Convention 2.44 / 5, Decisions 3.33 / 5, Reasoning 2.67 / 5, Scope 3.33 / 5DeepSeek V4 Pro: Correctness 1.63 / 5, Convention 2.38 / 5, Decisions 2.38 / 5, Reasoning 2.13 / 5, Scope 1.88 / 5CorrectnessConventionDecisionsReasoningScope
Kimi K2.7 CodeGLM-5.2DeepSeek V4 Pro
RADAR · RADAR-CHARTQuality profileFarther from the center is better
Kimi K2.7 Code: Correctness 1.89 / 5, Convention 2.22 / 5, Decisions 2.33 / 5, Reasoning 2.11 / 5, Scope 2.33 / 5GLM-5.2: Correctness 1.78 / 5, Convention 2.44 / 5, Decisions 3.33 / 5, Reasoning 2.67 / 5, Scope 3.33 / 5DeepSeek V4 Pro: Correctness 1.63 / 5, Convention 2.38 / 5, Decisions 2.38 / 5, Reasoning 2.13 / 5, Scope 1.88 / 5CorrectnessConventionDecisionsReasoningScope
Kimi K2.7 CodeGLM-5.2DeepSeek V4 Pro
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Pass: 1 of 27Partial: 5 of 27Fail: 20 of 27Not graded: 1 of 276 useful26 graded
Pass · 1Partial · 5Fail · 20Not graded · 1

THE OTHER TRIALS OF THE SAME CAMPAIGN同キャンペーンのその他の試行

Summary only; the full charts for these runs live in the predecessor project's records. Candidates throughout: deepseek-v4-pro, glm-5.2, kimi-k2.7-code.

ここでは要約のみ。これらの実行の完全なチャートは前身プロジェクトの記録にあります。候補は一貫して deepseek-v4-pro、glm-5.2、kimi-k2.7-code の3つでした。

TRIAL試行DATE日付WHAT IT TESTED検証内容MODELSモデル数OUTCOME結果
TEST 0324 JUL 2026Maximum-reasoning model comparison最大推論設定でのモデル比較36 / 27 useful · 1 pass, 5 partial27中6が有効 · 合格1、部分5
TEST 0424 JUL 2026Provider-default reasoning baseline, deterministically regradedプロバイダ既定推論のベースライン(決定的に再採点)35 / 27 useful · 0 pass, 5 partial27中5が有効 · 合格0、部分5
TEST 0524 JUL 2026GLM and Kimi head-to-head at equal work and thinking limitsGLM と Kimi の同一予算での直接対決216 of 48 planned attempts · no current winner計画48回中16回実施 · 勝者未確定
TEST 1414 AUG 2026Late 2-model recheck at higher difficulty高難度での2モデル再検証(後期)20 / 6 useful · the run that paused the predecessor's training6中0が有効 · 前身の訓練停止の契機となった実行

Only trials that actually ran appear here. A configuration that was measured and rejected never reads as history.

実際に走った試行だけを載せています。計測の末に不採用となった設定が歴史として残ることはありません。