DAIJIN / HOW IT WORKS
PROJECT ARCHITECTUREプロジェクトアーキテクチャ

How Daijin worksDaijin の仕組み

Daijin is a project brain for any repository: a Python Textual TUI over a Node engine daemon, speaking a frozen JSON-RPC contract. The sections below walk the six parts: the engine, the brain store, the retrieval floor, gates, the gym, and the TUI.

Daijin はどんなリポジトリにも適用できる「プロジェクトブレイン」です。Node のエンジンデーモンの上に Python Textual の TUI が載り、凍結された JSON-RPC 契約で会話します。以下では、エンジン、ブレインストア、検索フロア、ゲート、ジム、TUI の6つの部品を順に説明します。

The whole product is terminal-native and local-first. The engine daemon owns the store, the init pipeline, retrieval, MCP serving, and the gym; the TUI is a pure client that renders from the RPC surface plus an event stream. Nothing in the retrieval path ever calls a paid API, and everything that could spend sits behind an owner-only gate that the engine can only ever write as blocked.

プロダクト全体がターミナルネイティブかつローカルファーストです。エンジンデーモンがストア、初期化パイプライン、検索、MCP 提供、ジムを所有し、TUI は RPC サーフェスとイベントストリームだけから描画する純粋なクライアントです。検索経路は有料 API を一切呼ばず、支出し得るものはすべて、エンジン自身は「ブロック」としか書けないオーナー専用ゲートの背後にあります。

THE PIPELINE, IN ORDERパイプライン、順番どおりに
01ATTACH接続State lives under ~/.daijin,never in your tree状態は ~/.daijin に保存されあなたのツリーには書かないLOCALローカル 02INIT初期化Analyze, scaffold brain, mine goldset, discover gates, index解析、ブレイン生成、ゴールドセット採掘、ゲート発見、索引化ZERO SPEND支出ゼロ 03FLOORフロアScored against the gold set,permuted control beside itゴールドセットで採点し、並べ替え対照群を常に併記MEASURED計測済み 04SERVE提供A paste-ready MCP snippetfor your coding agents貼り付けるだけの MCP スニペットをエージェントにUNLOCK 0.75解放 0.75 05GYMジムGraded exams, harvestedlessons, back into the brain試験を採点し、教訓を収穫してブレインへ還元OWNER-GATEDオーナー制

The engine daemon: a contract that can fail the buildエンジンデーモン: ビルドを落とせる契約

The engine is Node 22, pure ESM, no framework. Clients speak JSON-RPC 2.0 over stdio or a Unix socket, reusing the MCP stdio transport framing. The RPC surface is a frozen, versioned document, engine/src/rpc/methods.md, currently contract v5: the first call on connect is a handshake that returns the contract version, and a mismatch renders an upgrade screen rather than a method error.

エンジンは Node 22、純粋な ESM、フレームワークなし。クライアントは MCP の stdio トランスポートのフレーミングを再利用した JSON-RPC 2.0(stdio または Unix ソケット経由)で会話します。RPC サーフェスは凍結・版管理された文書 engine/src/rpc/methods.md で、現在は契約 v5。接続時の最初の呼び出しはハンドシェイクで契約バージョンを返し、不一致はメソッドエラーではなくアップグレード画面として描画されます。

  • A shape gate enforces the contract in both directions. What the engine emits must match what the document describes, and vice versa, with the uncovered set printed every run. A contract change requires the leader; the gate can break the build.
  • シェイプゲートが契約を双方向に強制します。エンジンの出力は文書の記述と一致しなければならず、その逆も同様。未カバー集合は毎回出力されます。契約変更にはリーダーの承認が必要で、ゲートはビルドを落とせます。
  • Long work is a job with a live stream. Init, mining, gym cycles and the goal loop stream step events ({ ts, jobId, phase, step, detail }); every job ends with exactly one terminal event whose level says whether it ended well, was cancelled, or broke, and a failure always carries its reason.
  • 長い処理はジョブとしてライブストリーム化。初期化、採掘、ジムのサイクル、ゴールループはステップイベント({ ts, jobId, phase, step, detail })を流します。どのジョブも終端イベントをちょうど1つ発行し、その level が正常終了・キャンセル・故障を区別。失敗には必ず理由が付きます。
  • Spend-touching methods are enumerated exhaustively. Everything else is zero-spend by construction, and adding a spend path to any other method is a contract change, not an implementation detail.
  • 支出に触れるメソッドは網羅的に列挙。それ以外は構造上支出ゼロであり、他のメソッドに支出経路を足すことは実装詳細ではなく契約変更です。

The brain store: markdown is the truth, the index is disposableブレインストア: markdown が真実、インデックスは使い捨て

The brain is durable, evidence-cited markdown in .daijin/brain/: architecture, conventions, decisions, lessons. It is hand-editable; a generated brain can be curated into a hand-written one with no format migration. The index derives from those files and lives outside the repo under ~/.daijin/, so it can be deleted and regenerated at any time.

ブレインは .daijin/brain/ にある、根拠引用付きの恒久的な markdown です。アーキテクチャ、規約、意思決定、教訓。手で編集でき、生成されたブレインはフォーマット移行なしでキュレーション済みのものへ育てられます。インデックスはそれらのファイルから導出され、リポジトリ外の ~/.daijin/ に置かれるため、いつでも削除して再生成できます。

  • Local-first hybrid retrieval, zero spend. bge-m3 embeddings (1024-dim) via local Ollama, stored in sqlite-vec, beside FTS5 full-text search (porter unicode61), fused with reciprocal rank fusion. Nothing in this path calls a paid API, ever.
  • ローカルファーストのハイブリッド検索、支出ゼロ。ローカル Ollama による bge-m3 埋め込み(1024次元)を sqlite-vec に格納し、FTS5 全文検索(porter unicode61)と並べて、逆順位融合(RRF)で統合。この経路が有料 API を呼ぶことは決してありません。
  • Three layers of memory. The agent contract (never indexed, always loaded), the brain (durable markdown with citations), and the index (throwaway). A repo that already keeps a curated knowledge folder can be adopted as-is: the source folder is never written.
  • 3層のメモリ構造。エージェント契約(索引化されず常時ロード)、ブレイン(引用付きの恒久 markdown)、インデックス(使い捨て)。既にキュレーション済みの知識フォルダを持つリポジトリはそのまま取り込め、元フォルダには一切書き込みません。

The retrieval floor: a floor, not a vibe検索フロア: 雰囲気ではなく、フロア

A brain that cannot answer questions about its own repo does not get recommended to your tools. Init mines a gold set of retrieval cases from commit archaeology and structural self-reference: queries whose answers are artifacts (a file, a symbol, a document id), never prose. The gold set must pass its own integrity gates before it is allowed to measure anything: existence, leakage, staleness, provenance, and a diversity floor.

自分のリポジトリについての質問に答えられないブレインは、ツールに推薦されません。初期化時に、コミット考古学と構造的自己参照から検索ケースのゴールドセットを採掘します。答えは常に成果物(ファイル、シンボル、文書 ID)であり、散文ではありません。ゴールドセット自体が、存在・リーク・陳腐化・出所・多様性フロアという整合性ゲートを通過して初めて、計測に使うことを許されます。

  • Count form, never a bare percentage. The case rate is reported as 24 of 25, with must-not violations enforced at zero. MRR is recorded for movement only and deliberately not floored.
  • 常に件数表記、裸のパーセントは使わない。ケース率は 25件中24件 の形で報告され、禁止違反はゼロが強制されます。MRR は推移の観測のためだけに記録され、意図的にフロア化しません。
  • A permuted control beside every floor. The same store, embedder and budget are re-scored with permuted answers, so a saturated score cannot masquerade as a meaningful one: the report carries the discriminating range, not just the number.
  • どのフロアにも並べ替え対照群を併記。同じストア・埋め込み・予算で答えを並べ替えて再採点するため、飽和したスコアが意味のあるスコアのふりをできません。レポートには数値だけでなく判別レンジが載ります。
  • The serving budget is measured, not guessed. An init-time sweep runs 3000 / 4000 / 6000 / 8000 tokens and picks the smallest budget within one case of the best score, with the rationale recorded in one line.
  • 提供予算は推測ではなく計測で決める。初期化時のスイープが 3000 / 4000 / 6000 / 8000 トークンを走らせ、最良スコアから1ケース以内で最小の予算を選び、根拠を一行で記録します。
  • MCP unlocks at 0.75. Above the threshold, with zero violations, mcpSnippet hands you the paste-ready config serving brain.search, brain.conventions, brain.impact_of and brain.relevant_adrs. Below it, the refusal explains itself in the words the measurement used.
  • MCP はケース率 0.75 で解放。しきい値超えかつ違反ゼロで、mcpSnippet が brain.search / brain.conventions / brain.impact_of / brain.relevant_adrs を提供する貼り付け可能な設定を発行します。下回った場合の拒否文は、計測が使った言葉そのままで理由を説明します。

Gates: the repo's own checks, classified liveゲート: リポジトリ自身のチェックを、生きているかで分類

CI commands are discovered from the repo itself: package.json scripts, CI configs, Makefiles. Each candidate is run against a measured baseline and classified live, measured, pre-broken, or unavailable. Only gates that can fail on a bad edit and pass on a good one count as carrying signal; a gate that fails either way cannot tell a good edit from a bad one, and counting it is how a repo reports checks it does not have.

CI コマンドはリポジトリ自身から発見されます。package.json のスクリプト、CI 設定、Makefile。各候補は計測済みベースラインに対して実行され、live(生きている)、measured(計測基準)、pre-broken(元から壊れている)、unavailable(実行不能)に分類されます。悪い編集で落ち、良い編集で通るゲートだけがシグナルを持つとみなされます。どちらでも落ちるゲートは良し悪しを区別できず、それを数に入れることは「持っていないチェックを持っていると報告する」ことだからです。

  • The result is data the user owns. Discovery writes gates.yaml with liveness evidence attached; the engine never overwrites a user's edit, and if discovery loses the race it keeps your version and says so on the stream.
  • 結果はユーザーが所有するデータ。発見処理は生存性の証拠付きで gates.yaml を書きます。エンジンはユーザーの編集を決して上書きせず、競合に負けた場合はあなたの版を残し、その旨をストリームで告げます。

The gym: an active RAG loopジム: 能動的な RAG ループ

On top of the measured brain sits a certification gym. Exams are mined from the repo's real commit history; a student model attempts them in a sandboxed copy; a teacher model grades five cited axes; harvested lessons feed back into the brain. Every attempt, grade and harvest lands in an append-only ledger, and only evaluation-mode runs touch the scored record.

計測済みブレインの上に認定ジムが載ります。試験はリポジトリの実コミット履歴から採掘され、学生モデルがサンドボックスコピーで挑戦し、教師モデルが5つの根拠付き軸で採点し、収穫された教訓がブレインへ還流します。すべての挑戦・採点・収穫は追記専用台帳に記録され、採点対象の記録に触れるのは evaluation モードの実行だけです。

THE LEARNING LOOP, END TO END学習ループ、端から端まで
exam drawn試験を抽選 diff + gates差分とゲート結果 five cited axes5軸の根拠付き採点 lesson proposals教訓の提案のみ owner appliesオーナーが適用 serves the next cycle次サイクルに供給 01MINE採掘The auditor authors exams fromreal commit history監査役が実コミット履歴から試験を作成する 02ATTEMPT挑戦The student edits a sandboxedcopy under a token cap学生がトークン上限つきでサンドボックスコピーを編集 03GRADE採点The teacher cites evidence;digests bind the rubric教師が証拠を引用し、ダイジェストがルーブリックを束縛 04HARVEST収穫Graded gaps become reviewedproposals; nothing is written採点済みギャップを審査付き提案に。この段階では何も書かない 05APPLY適用The owner's separate act writesaccepted lessons into the brainオーナーの明示操作で、承認済み教訓をブレインへ書き込む 06BRAINブレインReindexed, so retrieval serveswhat was learned再索引され、学んだことが検索で提供されるようになる
  • The student works in a sandbox. Each attempt runs in a sandboxed copy of the repo: file tools only for CLI-driven roles, a strict JSON action protocol (read / edit / check / submit) for API roles, with every path confined to the sandbox and the harness owning gate runs. Per-attempt tokens are accounted against a real cap.
  • 学生はサンドボックスの中で作業。各挑戦はリポジトリのサンドボックスコピーで実行されます。CLI 駆動のロールはファイルツールのみ、API ロールは厳格な JSON アクションプロトコル(read / edit / check / submit)で、パスはすべてサンドボックス内に閉じ込められ、ゲート実行はハーネスの権限です。挑戦ごとのトークンは実上限に対して精算されます。
  • Grading is cited or it is refused. The teacher reads the task, the diff, the gate results and the shown document ids, never the gold commit. Five axes, each with a citation; run and packet digests are bound mechanically; an invented citation, a withheld-document reference, or a grader that authored the exam dies on the refusal list. A refusal leaves the attempt pending rather than failing the cycle.
  • 採点は根拠付きでなければ拒否される。教師が読むのは課題、差分、ゲート結果、提示された文書 ID であり、正解コミットは決して見ません。5軸それぞれに引用が必要で、実行とパケットのダイジェストは機械的に束縛されます。捏造した引用、非提示文書への参照、自分が作った試験の採点は拒否リストで即死。拒否は挑戦を保留にするだけで、サイクル自体は失敗しません。
  • Harvest proposes; apply is the owner's separate act. gymHarvest asks the teacher one question per graded gap and records a proposal-only batch; zero accepted proposals is a recorded, measured outcome. gymHarvestApply then writes accepted evaluation lessons into .daijin/brain/lessons/ with provenance in the body, and reindexes through the same derivation init uses. A batch applies at most once, and an unapplied batch is visibly unapplied.
  • 収穫は提案まで。適用はオーナーの別操作。gymHarvest は採点済みギャップごとに教師へ1問ずつ尋ね、提案のみのバッチを記録します。承認ゼロも記録される計測結果です。gymHarvestApply が承認済みの evaluation 教訓を出所情報付きで .daijin/brain/lessons/ に書き込み、初期化と同じ導出で再索引します。バッチの適用は最大1回、未適用のバッチは未適用と見えるようになっています。
  • The spend gate re-blocks itself. Opening it takes a scope, a written reason of at least 20 characters, and explicit confirmation; every job that spends re-blocks the gate in its finally, so an authorization lives exactly as long as the run it authorized. The engine never opens a gate on its own initiative.
  • 支出ゲートは自動で再ブロックされる。開くにはスコープ、20文字以上の理由、明示的な確認が必要です。支出するジョブはすべて finally でゲートを再ブロックするため、認可はそれが認可した実行と同じ寿命しか持ちません。エンジンが自発的にゲートを開くことは決してありません。

Four roles, four boundaries4つのロール、4つの境界

AUDITOR / EXAM COMMITTEE監査役 / 試験委員Mines and authors exams from commit history; triages watcher findings with written reasoning.コミット履歴から試験を採掘・作成し、ウォッチャーの指摘を根拠付きでトリアージ。Proposes; mechanics accept. No model promotes its own exam into the scored record.提案するだけで採否は機構が決める。モデルは自分の試験を採点記録に昇格できない。
STUDENT / ENGINEER学生 / エンジニアThe model under test. Solves one exam inside the sandbox, respecting gates and the token cap, and submits a minimal diff.試験対象のモデル。サンドボックス内で1つの試験を解き、ゲートとトークン上限を守って最小の差分を提出。Never grades; never sees the gold commit.採点はしない。正解コミットも見えない。
TEACHER教師Grades each attempt against the five-axis rubric with evidence-cited scores; answers one harvest question per graded gap.5軸ルーブリックで根拠引用付きの採点を行い、収穫では採点済みギャップごとに1問に答える。Never grades work it authored; refuses to invent a knowledge gap.自作の課題は採点しない。知識ギャップの捏造は拒否する。
WATCHERウォッチャーMonitors runs and engine health, verifies each finding in one cheap generation, reports to the auditor.実行とエンジンの健全性を監視し、各指摘を1回の安価な生成で検証して監査役に報告。Read-only: never edits, never grades, never spends on its own.読み取り専用。編集も採点も自発的な支出もしない。

The TUI: eight screens, one contractTUI: 8画面、1つの契約

The client is Python Textual: repos, a live init activity feed, a brain browser with a retrieval tester, gates, gym, exams, board, and settings. Charts use a textured vocabulary - pattern and color as two channels - so they stay readable without color; motion has three modes (full, reduced, off); and a full mock mode runs the entire UI with no engine and no Ollama at all.

クライアントは Python Textual 製です。リポジトリ一覧、初期化のライブフィード、検索テスター付きブレインブラウザ、ゲート、ジム、試験、ボード、設定の8画面。チャートはパターンと色の2チャネルを使うテクスチャ語彙で、色がなくても読めます。モーションはフル・省・オフの3モード。さらに完全なモックモードがあり、エンジンも Ollama もなしに UI 全体が動きます。

  • Every screen renders from the contract. The TUI is a pure client of methods.md v5 plus the event stream; the mock and the live engine answer the same shapes, which is what the shape gate holds in place.
  • すべての画面は契約から描画される。TUI は methods.md v5 とイベントストリームの純粋なクライアントです。モックと実エンジンは同じシェイプで答え、それを保持しているのがシェイプゲートです。

HONEST BOUND正直な注記

In the live end-to-end chains, every graded cycle attributed its gaps to measured-only causes, and the teacher refused to invent a knowledge gap five times. So no live batch has carried an accepted lesson yet: the write half of the loop is pinned by tests against a real brain, and the live transition awaits a run with a genuine knowledge gap. That is the design refusing to teach a lie, not a gap in the wiring.

ライブのエンドツーエンド実行では、採点された全サイクルがギャップを計測由来の原因にのみ帰属させ、教師は知識ギャップの捏造を5回拒否しました。そのため承認済み教訓を含むライブバッチはまだありません。ループの書き込み側は実ブレインに対するテストで固定済みで、ライブでの遷移は本物の知識ギャップを持つ実行を待っています。これは配線の欠陥ではなく、嘘を教えることを設計が拒んだ結果です。