DAIJIN
Daijin builds a measured knowledge brain for any repository, distilled from the repo's own code and history. It serves that brain to coding agents over MCP, and runs a certification gym where a student model attempts real exams mined from the repo's history, a teacher model grades them against five cited criteria, and harvested lessons feed back into the brain. An active RAG loop that measures what it serves, and improves it.
Daijin はどんなリポジトリにも、そのコードと履歴から蒸留した「プロジェクトブレイン」を構築するツールです。計測済みのブレインを MCP 経由でコーディングエージェントに提供し、さらに認定ジムでは、リポジトリの実履歴から採掘した試験に学生モデルが挑み、教師モデルが5つの根拠付き基準で採点し、収穫した教訓をブレインに書き戻します。提供する知識を計測し、改善し続ける能動的な RAG ループです。
THE FLOOR, MEASURED ON A REAL REPO実リポジトリで測ったフロア
A live layer-1 init on the owner's TokaiHub repository, measured 2026-08-22 at zero spend. The floor is reported in count form, never as a bare rounded percentage.
オーナーの TokaiHub リポジトリでの実測(2026-08-22、レイヤー1 init、支出ゼロ)。フロアは常に件数表記で報告され、丸めたパーセントだけで語ることはありません。
TWO SUITES, ZERO FAILURES2つのテストスイート、失敗ゼロ
Both halves of the product carry their own test suite, run at zero network and zero spend. Counts are from the last verification, 2026-08-23.
エンジンと TUI はそれぞれ独立したテストスイートを持ち、ネットワークも支出もゼロで走ります。件数は最終検証(2026-08-23)時点のものです。
FIVE LIVE CHAINS, EXIT 0ライブ実行5チェーン、すべて正常終了
The full drive, run live against real providers: mine exams, promote, run the gym, grade, harvest, apply. Five chains across four repos, all exit 0; mining validated 3 of 3 exams on the clean fixture.
試験の採掘 → 昇格 → ジム実行 → 採点 → 収穫 → 適用までの全工程を、実プロバイダ相手にライブで実行。4リポジトリで5チェーン、すべて exit 0。クリーンフィクスチャでは採掘した試験3件中3件が検証を通過しました。
HONEST BOUND: in the live chains the teacher refused to invent a knowledge gap five times, so no live batch carried an accepted lesson yet. That is the design refusing to teach a lie, not a gap in the wiring; the write path is pinned by tests against a real brain.
正直な注記: ライブ実行では教師が「知識ギャップの捏造」を5回拒否したため、承認済み教訓を含むバッチはまだありません。これは配線の欠陥ではなく、嘘を教えることを設計が拒んだ結果です。書き込み経路は実ブレインに対するテストで固定済みです。
ONE TOOL, FIVE ACTS1つのツール、5つの動作
Everything below the gym runs at zero spend; the gym itself sits behind an owner-only gate that is blocked by default.
ジム以外はすべて支出ゼロで動作します。ジム自体はオーナー専用ゲートの背後にあり、既定でブロックされています。
- Build a brain. Attach any git repo; daijin analyzes it, scaffolds durable markdown knowledge (architecture, conventions, decisions, lessons) in .daijin/brain, and indexes it locally. The index is disposable; the markdown is the truth.
- ブレインを構築する。git リポジトリを接続すると、daijin が解析し、恒久的な markdown 知識(アーキテクチャ・規約・意思決定・教訓)を .daijin/brain に生成してローカルでインデックス化します。インデックスは使い捨て、markdown が真実です。
- Measure it. A gold set of retrieval cases is mined from the repo's real commit history, passes its own integrity gates, and scores the brain in count form with a permuted control beside it.
- 計測する。検索ケースのゴールドセットを実コミット履歴から採掘し、それ自体が整合性ゲートを通過した上で、ブレインを件数表記で採点。並べ替え対照群も常に併記されます。
- Serve it over MCP. Above the measured floor, daijin hands you a paste-ready MCP snippet: brain.search, brain.conventions, brain.impact_of, brain.relevant_adrs for your coding agents.
- MCP で提供する。計測済みフロアを超えると、貼り付けるだけの MCP スニペットが発行されます。コーディングエージェントから brain.search / brain.conventions / brain.impact_of / brain.relevant_adrs が使えます。
- Run the gym. Owner-gated: a student model attempts exams mined from your history in a sandbox, a teacher model grades five cited axes, and an append-only ledger records everything.
- ジムを走らせる。オーナーゲート制。学生モデルが履歴から採掘した試験にサンドボックス内で挑み、教師モデルが5つの根拠付き軸で採点。追記専用台帳がすべてを記録します。
- Harvest lessons. Graded gaps become reviewed lesson proposals; applying a batch writes real lesson documents into the brain and reindexes, so retrieval serves what was learned.
- 教訓を収穫する。採点で見つかったギャップは審査済みの教訓提案になり、適用するとブレインに教訓文書が書き込まれ再インデックス。学んだことが検索で提供されるようになります。
TWO COMMANDS, ONE PREFIXコマンド2つ、プレフィックス1つ
Everything lands in one self-contained prefix with a single daijin command on your PATH. Installing twice lands in the same state as once.
すべてが自己完結のプレフィックスに収まり、PATH には daijin コマンドが1つ増えるだけ。2回インストールしても1回と同じ状態になります。
git clone https://github.com/MohamedFuad16/daijin.git cd daijin bash install/install.sh
Then: daijin /path/to/your/repo to connect a real repo, or daijin . --mock to explore the full UI with no engine and no Ollama.
その後は daijin /path/to/your/repo で実リポジトリに接続、または daijin . --mock でエンジンも Ollama もなしに UI 全体を試せます。
HOW THE STUDENT MODEL WAS CHOSEN学生モデルはこうして選ばれた
A historical benchmark from the predecessor project (July 2026): candidate student models ran the same real-repo exam bank under the five-axis teacher rubric, on equal footing. GLM won on quality per cost, and its successor GLM-5.3 is the student daijin ships with today. The records below are kept verbatim, in their original English.
前身プロジェクトでの歴史的ベンチマーク(2026年7月)。学生モデルの候補たちが、同一の実リポジトリ試験バンクを5軸の教師ルーブリックのもと、同条件で受験しました。コストあたり品質で GLM が勝ち、その後継である GLM-5.3 が現在の daijin の学生です。以下の記録は当時のまま、原文(英語)で保存しています。
TEST 02MODEL COMPARISON · 23 JUL 2026Coding model capability comparisonVIEW RESULTS
AUDIT HOLD. These saved results used a grading rule that treated equivalent work differently. They remain visible for history but are not ranked until deterministic regrading.
| MODEL | SETTINGS | OUTCOMES | TASK SIZES | TEACHER AVERAGE | 5 QUALITY SCORES | GRADED | ENGINEER USAGE | GRADING + TOTAL COST |
|---|---|---|---|---|---|---|---|---|
Kimi K2.7 Code Kimi K2.7 Codemoonshotai/kimi-k2.7-code | HISTORICAL LABEL · PENDING REGRADE Reasoning levelProvider defaultPrecisionHost-dependentOutput cap32K tokens / replyTime limit300s | 1 PASS1 PARTIAL7 FAIL | SMALL2/4 usefulMEDIUM0/2 usefulLARGE0/3 useful | AVERAGE SCORE2.17 / 51 is weak · 5 is excellent | CORRECT1.89 / 5STYLE2.22 / 5DECISIONS2.33 / 5REASONING2.11 / 5SCOPE2.33 / 5 | 9/9complete | 2.3M total tokensWork used2.3MThinking used32K · 64.6%Answer generated18KWork allowanceby task size · 300K–600K at the time of the run / examThinking allowanceshared with work budget / examAverage time1.7 min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops0Thinking-cap stops0Tool-flow errors16Unknown calls0 | $2.81 grading $3.57 total $0.40 per task 0 unknown teacher calls |
GLM-5.2 GLM-5.2z-ai/glm-5.2 | HISTORICAL LABEL · PENDING REGRADE Reasoning levelThinking on · Max effort (provider default)PrecisionHost-dependentOutput cap32K tokens / replyTime limit300s | 0 PASS2 PARTIAL7 FAIL | SMALL2/4 usefulMEDIUM0/2 usefulLARGE0/3 useful | AVERAGE SCORE2.71 / 51 is weak · 5 is excellent | CORRECT1.78 / 5STYLE2.44 / 5DECISIONS3.33 / 5REASONING2.67 / 5SCOPE3.33 / 5 | 9/9complete | 2M total tokensWork used2MThinking used101K · 82.0%Answer generated22KWork allowanceby task size · 300K–600K at the time of the run / examThinking allowanceshared with work budget / examAverage time5.4 min/taskReply-limit stops0Timeout truncations0Provider-error tasks22.2%Work-cap stops0Thinking-cap stops0Tool-flow errors19Unknown calls2 | $2.87 grading $3.71 total $0.41 per task 0 unknown teacher calls |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | HISTORICAL LABEL · PENDING REGRADE Reasoning levelProvider defaultPrecisionHost-dependentOutput cap12K tokens / replyTime limit300s | 0 PASS2 PARTIAL6 FAIL | SMALL1/4 usefulMEDIUM0/2 usefulLARGE1/3 useful | AVERAGE SCORE2.08 / 51 is weak · 5 is excellent | CORRECT1.63 / 5STYLE2.38 / 5DECISIONS2.38 / 5REASONING2.13 / 5SCOPE1.88 / 5 | 8/91 ungraded | 2.4M total tokensWork used2.4MThinking used27K · 45.4%Answer generated32KWork allowanceby task size · 300K–600K at the time of the run / examThinking allowanceshared with work budget / examAverage time0.6 min/taskReply-limit stops0Timeout truncations0Provider-error tasks22.2%Work-cap stops1Thinking-cap stops0Tool-flow errors22Unknown calls2 | $2.20 grading $2.80 total $0.31 per task 0 unknown teacher calls |
GLM-5.2 GLM-5.2z-ai/glm-5.2 | CONTEXT ONLY · NOT CAPABILITY RANKED Reasoning levelThinking on · Max effort (provider default)PrecisionHost-dependentOutput cap12K tokens / replyTime limit300s | 0 PASS0 PARTIAL2 FAIL | SMALL0/2 usefulMEDIUM0/2 usefulLARGE0/1 useful | AVERAGE SCORE1.40 / 51 is weak · 5 is excellent | CORRECT1.50 / 5STYLE2.00 / 5DECISIONS1.00 / 5REASONING1.50 / 5SCOPE1.00 / 5 | 2/53 ungraded | 1.3M total tokensWork used1.3MThinking used132K · 99.6%Answer generated2KWork allowanceby task size · 300K–600K at the time of the run / examThinking allowanceshared with work budget / examAverage time— min/taskReply-limit stops11Timeout truncations0Provider-error tasks0.0%Work-cap stops3Thinking-cap stops0Tool-flow errors15Unknown calls0 | $0.46 grading $1.15 total $0.23 per task 0 unknown teacher calls $0.90 aborted/no verdict · 362K tokens |
THE OTHER TRIALS OF THE SAME CAMPAIGN同キャンペーンのその他の試行
Summary only; the full charts for these runs live in the predecessor project's records. Candidates throughout: deepseek-v4-pro, glm-5.2, kimi-k2.7-code.
ここでは要約のみ。これらの実行の完全なチャートは前身プロジェクトの記録にあります。候補は一貫して deepseek-v4-pro、glm-5.2、kimi-k2.7-code の3つでした。
| TRIAL試行 | DATE日付 | WHAT IT TESTED検証内容 | MODELSモデル数 | OUTCOME結果 |
|---|---|---|---|---|
| TEST 03 | 24 JUL 2026 | Maximum-reasoning model comparison最大推論設定でのモデル比較 | 3 | 6 / 27 useful · 1 pass, 5 partial27中6が有効 · 合格1、部分5 |
| TEST 04 | 24 JUL 2026 | Provider-default reasoning baseline, deterministically regradedプロバイダ既定推論のベースライン(決定的に再採点) | 3 | 5 / 27 useful · 0 pass, 5 partial27中5が有効 · 合格0、部分5 |
| TEST 05 | 24 JUL 2026 | GLM and Kimi head-to-head at equal work and thinking limitsGLM と Kimi の同一予算での直接対決 | 2 | 16 of 48 planned attempts · no current winner計画48回中16回実施 · 勝者未確定 |
| TEST 14 | 14 AUG 2026 | Late 2-model recheck at higher difficulty高難度での2モデル再検証(後期) | 2 | 0 / 6 useful · the run that paused the predecessor's training6中0が有効 · 前身の訓練停止の契機となった実行 |
Only trials that actually ran appear here. A configuration that was measured and rejected never reads as history.
実際に走った試行だけを載せています。計測の末に不採用となった設定が歴史として残ることはありません。