AI BRAIN / OWNER DASHBOARD
CURRENT STATUS · 29 JUL 2026

PAID WORK RESTRICTED

WHY TRAINING IS PAUSEDCycles 32 and 33 completed; await a new owner-authorized cohort.
WHAT HAPPENS NEXTCheck the repository gate before starting any model call.
1 ·TEST FOUNDATION

EVALUATION BANK READINESS

Before any score means anything, the exam bank itself has to be sound. These are counts of tasks, not results.

EVALUATION EXAMS5849 daily-practice + 9 independent
INDEPENDENT EXAMS9Reserved exams that never teach the Brain
EXCLUDED EXAMS5Removed because the task and known-good answer did not match reliably
TASK DESCRIPTIONS APPROVED Medium and large exams carry a written description of required behaviour, integration points, and acceptance criteria. Only an approved description is used by the runner. A draft is written but deliberately not approved: approving one appends it to that exam prompt and changes the retrieval query, so the drafts are held back to keep the current baseline comparable; a rejected one quarantines its exam. Small exams do not need a description.4 / 4435 HELD IN DRAFT ON PURPOSE5 REJECTED19 SMALL · NOT NEEDED
INDEPENDENT BASELINEREADYExact exam wording is saved
2 ·THE LEARNING LOOP

TRAINING CYCLES One cycle is one five-exam practice batch: the engineer attempts the exams, the teacher grades them, and accepted lessons are written to the Brain. Read down the cycles to see whether learning is helping.

The loop itself: each cycle gives the student five exams, grades them, and writes back what it learned. The five most recent are here; open one to see what it scored and what it cost.

CORRECTNESS +2.20 OF 5 SINCE THE FIRST CYCLETraining cycles on GLM-5.2
28 JUL 20269 cycles54 graded exams21 total on record
09nightly · 28 JUL 2026Cycle 33VIEW RESULTS
MODELRUN SETTINGSOUTCOMESTASK SIZESTASK AVERAGE5 QUALITY SCORESUSEFUL RESULTSENGINEER USAGEBRAIN GROWTH
GLM-5.2
GLM-5.2glm-5.2
ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget200,000 / exam
2 PASS0 PARTIAL3 FAIL
SMALL1/2 usefulMEDIUM1/3 usefulLARGE0/0 useful
AVERAGE SCORE3.52 / 51 is weak · 5 is excellent
CORRECT3.20 / 5STYLE3.60 / 5DECISIONS3.80 / 5REASONING3.60 / 5SCOPE3.40 / 5
2 / 540.0% partial or pass
1.1M total tokensThinking used87K · 87.6%Answer generated12KGraded5/5Provider-error tasks0.0%
+0 lessons
0 docs / 0 chunks
flat-rate engineer
DITHERED CYCLE EVIDENCEFour views of Cycle 33
One cycle only. The consolidated trend across every cycle is in the next section.
BAR · OUTCOME-CHARTOutcome distributionSame 5 exams per setting
2 useful5 graded
02452 usefulCycle 335 / 5 gradedCycle 33: 2 pass, 0 partial, 3 fail; 5 of 5 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135Cycle 33: Correctness 3.20 / 5, Convention 3.60 / 5, Decisions 3.80 / 5, Reasoning 3.60 / 5, Scope 3.40 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 33
RADAR · RADAR-CHARTQuality profileFarther from the center is better
Cycle 33: Correctness 3.20 / 5, Convention 3.60 / 5, Decisions 3.80 / 5, Reasoning 3.60 / 5, Scope 3.40 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 33
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Pass: 2 of 5Fail: 3 of 52 useful5 graded
Pass · 2Partial · 0Fail · 3Not graded · 0
08nightly · 28 JUL 2026Cycle 32VIEW RESULTS
MODELRUN SETTINGSOUTCOMESTASK SIZESTASK AVERAGE5 QUALITY SCORESUSEFUL RESULTSENGINEER USAGEBRAIN GROWTH
GLM-5.2
GLM-5.2glm-5.2
ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam
1 PASS1 PARTIAL3 FAIL
SMALL1/1 usefulMEDIUM0/1 usefulLARGE1/3 useful
AVERAGE SCORE3.24 / 51 is weak · 5 is excellent
CORRECT3.00 / 5STYLE3.40 / 5DECISIONS3.60 / 5REASONING3.40 / 5SCOPE2.80 / 5
2 / 540.0% partial or pass
1.7M total tokensThinking used150K · 91.3%Answer generated14KGraded5/5Provider-error tasks0.0%
+0 lessons
0 docs / 0 chunks
flat-rate engineer
DITHERED CYCLE EVIDENCEFour views of Cycle 32
One cycle only. The consolidated trend across every cycle is in the next section.
BAR · OUTCOME-CHARTOutcome distributionSame 5 exams per setting
2 useful5 graded
02452 usefulCycle 325 / 5 gradedCycle 32: 1 pass, 1 partial, 3 fail; 5 of 5 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135Cycle 32: Correctness 3.00 / 5, Convention 3.40 / 5, Decisions 3.60 / 5, Reasoning 3.40 / 5, Scope 2.80 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 32
RADAR · RADAR-CHARTQuality profileFarther from the center is better
Cycle 32: Correctness 3.00 / 5, Convention 3.40 / 5, Decisions 3.60 / 5, Reasoning 3.40 / 5, Scope 2.80 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 32
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Pass: 1 of 5Partial: 1 of 5Fail: 3 of 52 useful5 graded
Pass · 1Partial · 1Fail · 3Not graded · 0
07weekly · 28 JUL 2026Cycle 31VIEW RESULTS
MODELRUN SETTINGSOUTCOMESTASK SIZESTASK AVERAGE5 QUALITY SCORESUSEFUL RESULTSENGINEER USAGEBRAIN GROWTH
GLM-5.2
GLM-5.2glm-5.2
ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam
0 PASS4 PARTIAL10 FAIL
SMALL4/8 usefulMEDIUM0/1 usefulLARGE0/5 useful
AVERAGE SCORE2.05 / 51 is weak · 5 is excellent
CORRECT2.00 / 5STYLE2.42 / 5DECISIONS1.83 / 5REASONING2.33 / 5SCOPE1.67 / 5
4 / 1428.6% partial or pass
4.3M total tokensThinking used448K · 94.3%Answer generated27KGraded12/14Provider-error tasks0.0%
+0 lessons
0 docs / 0 chunks
flat-rate engineer
DITHERED CYCLE EVIDENCEFour views of Cycle 31
One cycle only. The consolidated trend across every cycle is in the next section.
BAR · OUTCOME-CHARTOutcome distributionSame 14 exams per setting
4 useful14 graded
0510144 usefulCycle 3114 / 14 gradedCycle 31: 0 pass, 4 partial, 10 fail; 14 of 14 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135Cycle 31: Correctness 2.00 / 5, Convention 2.42 / 5, Decisions 1.83 / 5, Reasoning 2.33 / 5, Scope 1.67 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 31
RADAR · RADAR-CHARTQuality profileFarther from the center is better
Cycle 31: Correctness 2.00 / 5, Convention 2.42 / 5, Decisions 1.83 / 5, Reasoning 2.33 / 5, Scope 1.67 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 31
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Partial: 4 of 14Fail: 10 of 144 useful14 graded
Pass · 0Partial · 4Fail · 10Not graded · 0
06nightly · 27 JUL 2026Cycle 30VIEW RESULTS
MODELRUN SETTINGSOUTCOMESTASK SIZESTASK AVERAGE5 QUALITY SCORESUSEFUL RESULTSENGINEER USAGEBRAIN GROWTH
GLM-5.2
GLM-5.2glm-5.2
ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam
0 PASS0 PARTIAL5 FAIL
SMALL0/2 usefulMEDIUM0/1 usefulLARGE0/2 useful
AVERAGE SCORE1.52 / 51 is weak · 5 is excellent
CORRECT1.20 / 5STYLE1.80 / 5DECISIONS1.60 / 5REASONING1.80 / 5SCOPE1.20 / 5
0 / 50.0% partial or pass
1.5M total tokensThinking used142K · 92.7%Answer generated11KGraded5/5Provider-error tasks0.0%
+0 lessons
0 docs / 0 chunks
flat-rate engineer
DITHERED CYCLE EVIDENCEFour views of Cycle 30
One cycle only. The consolidated trend across every cycle is in the next section.
BAR · OUTCOME-CHARTOutcome distributionSame 5 exams per setting
0 useful5 graded
02450 usefulCycle 305 / 5 gradedCycle 30: 0 pass, 0 partial, 5 fail; 5 of 5 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135Cycle 30: Correctness 1.20 / 5, Convention 1.80 / 5, Decisions 1.60 / 5, Reasoning 1.80 / 5, Scope 1.20 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 30
RADAR · RADAR-CHARTQuality profileFarther from the center is better
Cycle 30: Correctness 1.20 / 5, Convention 1.80 / 5, Decisions 1.60 / 5, Reasoning 1.80 / 5, Scope 1.20 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 30
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Fail: 5 of 50 useful5 graded
Pass · 0Partial · 0Fail · 5Not graded · 0
05nightly · 25 JUL 2026Cycle 29VIEW RESULTS
MODELRUN SETTINGSOUTCOMESTASK SIZESTASK AVERAGE5 QUALITY SCORESUSEFUL RESULTSENGINEER USAGEBRAIN GROWTH
GLM-5.2
GLM-5.2glm-5.2
ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam
1 PASS0 PARTIAL4 FAIL
SMALL1/1 usefulMEDIUM0/3 usefulLARGE0/1 useful
AVERAGE SCORE2.76 / 51 is weak · 5 is excellent
CORRECT2.40 / 5STYLE3.20 / 5DECISIONS3.20 / 5REASONING2.60 / 5SCOPE2.40 / 5
1 / 520.0% partial or pass
1.7M total tokensThinking used168K · 92.6%Answer generated13KGraded5/5Provider-error tasks0.0%
+8 lessons
8 docs / 8 chunks
flat-rate engineer
DITHERED CYCLE EVIDENCEFour views of Cycle 29
One cycle only. The consolidated trend across every cycle is in the next section.
BAR · OUTCOME-CHARTOutcome distributionSame 5 exams per setting
1 useful5 graded
02451 usefulCycle 295 / 5 gradedCycle 29: 1 pass, 0 partial, 4 fail; 5 of 5 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135Cycle 29: Correctness 2.40 / 5, Convention 3.20 / 5, Decisions 3.20 / 5, Reasoning 2.60 / 5, Scope 2.40 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 29
RADAR · RADAR-CHARTQuality profileFarther from the center is better
Cycle 29: Correctness 2.40 / 5, Convention 3.20 / 5, Decisions 3.20 / 5, Reasoning 2.60 / 5, Scope 2.40 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 29
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Pass: 1 of 5Fail: 4 of 51 useful5 graded
Pass · 1Partial · 0Fail · 4Not graded · 0
OLDER CYCLES · 4 MORE ON THIS ENGINEER

Same engineer and the same grading, just further back. Kept so a trend can be checked, folded away so the recent result leads.

04nightly · 25 JUL 2026Cycle 28VIEW RESULTS
MODELRUN SETTINGSOUTCOMESTASK SIZESTASK AVERAGE5 QUALITY SCORESUSEFUL RESULTSENGINEER USAGEBRAIN GROWTH
GLM-5.2
GLM-5.2glm-5.2
ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam
0 PASS2 PARTIAL3 FAIL
SMALL1/1 usefulMEDIUM1/4 usefulLARGE0/0 useful
AVERAGE SCORE2.90 / 51 is weak · 5 is excellent
CORRECT3.00 / 5STYLE2.25 / 5DECISIONS3.25 / 5REASONING3.00 / 5SCOPE3.00 / 5
2 / 540.0% partial or pass
1.3M total tokensThinking used149K · 93.5%Answer generated10KGraded4/5Provider-error tasks0.0%
+8 lessons
8 docs / 8 chunks
flat-rate engineer
DITHERED CYCLE EVIDENCEFour views of Cycle 28
One cycle only. The consolidated trend across every cycle is in the next section.
BAR · OUTCOME-CHARTOutcome distributionSame 5 exams per setting
2 useful5 graded
02452 usefulCycle 285 / 5 gradedCycle 28: 0 pass, 2 partial, 3 fail; 5 of 5 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135Cycle 28: Correctness 3.00 / 5, Convention 2.25 / 5, Decisions 3.25 / 5, Reasoning 3.00 / 5, Scope 3.00 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 28
RADAR · RADAR-CHARTQuality profileFarther from the center is better
Cycle 28: Correctness 3.00 / 5, Convention 2.25 / 5, Decisions 3.25 / 5, Reasoning 3.00 / 5, Scope 3.00 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 28
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Partial: 2 of 5Fail: 3 of 52 useful5 graded
Pass · 0Partial · 2Fail · 3Not graded · 0
03nightly · 25 JUL 2026Cycle 27VIEW RESULTS
MODELRUN SETTINGSOUTCOMESTASK SIZESTASK AVERAGE5 QUALITY SCORESUSEFUL RESULTSENGINEER USAGEBRAIN GROWTH
GLM-5.2
GLM-5.2glm-5.2
ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam
0 PASS1 PARTIAL4 FAIL
SMALL0/0 usefulMEDIUM0/0 usefulLARGE1/5 useful
AVERAGE SCORE2.35 / 51 is weak · 5 is excellent
CORRECT1.75 / 5STYLE2.50 / 5DECISIONS3.00 / 5REASONING2.50 / 5SCOPE2.00 / 5
1 / 520.0% partial or pass
2.1M total tokensThinking used245K · 96.1%Answer generated10KGraded4/5Provider-error tasks0.0%
+6 lessons
9 docs / 9 chunks
flat-rate engineer
DITHERED CYCLE EVIDENCEFour views of Cycle 27
One cycle only. The consolidated trend across every cycle is in the next section.
BAR · OUTCOME-CHARTOutcome distributionSame 5 exams per setting
1 useful5 graded
02451 usefulCycle 275 / 5 gradedCycle 27: 0 pass, 1 partial, 4 fail; 5 of 5 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135Cycle 27: Correctness 1.75 / 5, Convention 2.50 / 5, Decisions 3.00 / 5, Reasoning 2.50 / 5, Scope 2.00 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 27
RADAR · RADAR-CHARTQuality profileFarther from the center is better
Cycle 27: Correctness 1.75 / 5, Convention 2.50 / 5, Decisions 3.00 / 5, Reasoning 2.50 / 5, Scope 2.00 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 27
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Partial: 1 of 5Fail: 4 of 51 useful5 graded
Pass · 0Partial · 1Fail · 4Not graded · 0
02nightly · 25 JUL 2026Cycle 26VIEW RESULTS
MODELRUN SETTINGSOUTCOMESTASK SIZESTASK AVERAGE5 QUALITY SCORESUSEFUL RESULTSENGINEER USAGEBRAIN GROWTH
GLM-5.2
GLM-5.2glm-5.2
ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam
1 PASS1 PARTIAL3 FAIL
SMALL1/2 usefulMEDIUM1/2 usefulLARGE0/1 useful
AVERAGE SCORE2.64 / 51 is weak · 5 is excellent
CORRECT2.40 / 5STYLE3.20 / 5DECISIONS2.80 / 5REASONING2.40 / 5SCOPE2.40 / 5
2 / 540.0% partial or pass
1.6M total tokensThinking used142K · 89.3%Answer generated17KGraded5/5Provider-error tasks0.0%
+9 lessons
8 docs / 8 chunks
flat-rate engineer
DITHERED CYCLE EVIDENCEFour views of Cycle 26
One cycle only. The consolidated trend across every cycle is in the next section.
BAR · OUTCOME-CHARTOutcome distributionSame 5 exams per setting
2 useful5 graded
02452 usefulCycle 265 / 5 gradedCycle 26: 1 pass, 1 partial, 3 fail; 5 of 5 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135Cycle 26: Correctness 2.40 / 5, Convention 3.20 / 5, Decisions 2.80 / 5, Reasoning 2.40 / 5, Scope 2.40 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 26
RADAR · RADAR-CHARTQuality profileFarther from the center is better
Cycle 26: Correctness 2.40 / 5, Convention 3.20 / 5, Decisions 2.80 / 5, Reasoning 2.40 / 5, Scope 2.40 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 26
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Pass: 1 of 5Partial: 1 of 5Fail: 3 of 52 useful5 graded
Pass · 1Partial · 1Fail · 3Not graded · 0
01nightly · 25 JUL 2026Cycle 25 · baselineVIEW RESULTS
MODELRUN SETTINGSOUTCOMESTASK SIZESTASK AVERAGE5 QUALITY SCORESUSEFUL RESULTSENGINEER USAGEBRAIN GROWTH
GLM-5.2
GLM-5.2glm-5.2
BASELINE
ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam
0 PASS0 PARTIAL5 FAIL
SMALL0/2 usefulMEDIUM0/2 usefulLARGE0/1 useful
AVERAGE SCORE1.92 / 51 is weak · 5 is excellent
CORRECT1.00 / 5STYLE2.00 / 5DECISIONS2.80 / 5REASONING2.00 / 5SCOPE1.80 / 5
0 / 50.0% partial or pass
1.7M total tokensThinking used160K · 91.8%Answer generated14KGraded5/5Provider-error tasks0.0%
+10 lessons
11 docs / 11 chunks
flat-rate engineer
DITHERED CYCLE EVIDENCEFour views of Cycle 25 · baseline
One cycle only. The consolidated trend across every cycle is in the next section.
BAR · OUTCOME-CHARTOutcome distributionSame 5 exams per setting
0 useful5 graded
02450 usefulCycle 25 · baseline5 / 5 gradedCycle 25 · baseline: 0 pass, 0 partial, 5 fail; 5 of 5 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135Cycle 25 · baseline: Correctness 1.00 / 5, Convention 2.00 / 5, Decisions 2.80 / 5, Reasoning 2.00 / 5, Scope 1.80 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 25 · baseline
RADAR · RADAR-CHARTQuality profileFarther from the center is better
Cycle 25 · baseline: Correctness 1.00 / 5, Convention 2.00 / 5, Decisions 2.80 / 5, Reasoning 2.00 / 5, Scope 1.80 / 5CorrectnessConventionDecisionsReasoningScope
Cycle 25 · baseline
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Fail: 5 of 50 useful5 graded
Pass · 0Partial · 0Fail · 5Not graded · 0
EARLIER CYCLES · 12 ON A PREVIOUS ENGINEER MODEL

These ran on DeepSeek V4 Pro and, for the oldest of them, a harness that stopped a solution after its first edit. They are kept for history and are deliberately excluded from the charts above.

MODELRUN SETTINGSOUTCOMESTASK SIZESTASK AVERAGE5 QUALITY SCORESUSEFUL RESULTSENGINEER USAGEBRAIN GROWTH
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap
0 PASS0 PARTIAL5 FAIL
SMALL0/1 usefulMEDIUM0/1 usefulLARGE0/3 useful
AVERAGE SCORE1.80 / 51 is weak · 5 is excellent
CORRECT1.20 / 5STYLE1.80 / 5DECISIONS2.20 / 5REASONING2.00 / 5SCOPE1.80 / 5
0 / 50.0% partial or pass
1.4M total tokensThinking used0 · 0.0%Answer generated0Graded5/5Provider-error tasks0.0%
+9 lessons
0 docs / 0 chunks
$1.39
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap
0 PASS0 PARTIAL5 FAIL
SMALL0/1 usefulMEDIUM0/0 usefulLARGE0/4 useful
AVERAGE SCORE1.88 / 51 is weak · 5 is excellent
CORRECT1.20 / 5STYLE2.40 / 5DECISIONS1.80 / 5REASONING2.00 / 5SCOPE2.00 / 5
0 / 50.0% partial or pass
1.8M total tokensThinking used0 · 0.0%Answer generated0Graded5/5Provider-error tasks0.0%
+12 lessons
9 docs / 9 chunks
$1.25
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap
0 PASS0 PARTIAL5 FAIL
SMALL0/2 usefulMEDIUM0/1 usefulLARGE0/2 useful
AVERAGE SCORE2.20 / 51 is weak · 5 is excellent
CORRECT1.75 / 5STYLE2.75 / 5DECISIONS2.25 / 5REASONING2.25 / 5SCOPE2.00 / 5
0 / 50.0% partial or pass
1.6M total tokensThinking used0 · 0.0%Answer generated0Graded4/5Provider-error tasks0.0%
+8 lessons
9 docs / 9 chunks
$0.94
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap
0 PASS2 PARTIAL16 FAIL
SMALL2/4 usefulMEDIUM0/4 usefulLARGE0/10 useful
AVERAGE SCORE2.25 / 51 is weak · 5 is excellent
CORRECT1.53 / 5STYLE2.59 / 5DECISIONS2.71 / 5REASONING2.24 / 5SCOPE2.18 / 5
2 / 1811.1% partial or pass
5.7M total tokensThinking used0 · 0.0%Answer generated0Graded17/18Provider-error tasks0.0%
+7 lessons
1 docs / 1 chunks
$6.81
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap
0 PASS1 PARTIAL4 FAIL
SMALL0/1 usefulMEDIUM0/2 usefulLARGE1/2 useful
AVERAGE SCORE2.20 / 51 is weak · 5 is excellent
CORRECT1.75 / 5STYLE2.75 / 5DECISIONS2.50 / 5REASONING2.00 / 5SCOPE2.00 / 5
1 / 520.0% partial or pass
1.7M total tokensThinking used0 · 0.0%Answer generated0Graded4/5Provider-error tasks0.0%
+8 lessons
9 docs / 9 chunks
$1.47
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap
0 PASS0 PARTIAL5 FAIL
SMALL0/3 usefulMEDIUM0/2 usefulLARGE0/0 useful
AVERAGE SCORE2.20 / 51 is weak · 5 is excellent
CORRECT1.50 / 5STYLE3.00 / 5DECISIONS2.50 / 5REASONING2.00 / 5SCOPE2.00 / 5
0 / 50.0% partial or pass
1.3M total tokensThinking used0 · 0.0%Answer generated0Graded2/5Provider-error tasks0.0%
+5 lessons
9 docs / 9 chunks
$0.67
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap
0 PASS1 PARTIAL4 FAIL
SMALL1/1 usefulMEDIUM0/2 usefulLARGE0/2 useful
AVERAGE SCORE1.84 / 51 is weak · 5 is excellent
CORRECT1.40 / 5STYLE2.80 / 5DECISIONS2.00 / 5REASONING1.60 / 5SCOPE1.40 / 5
1 / 520.0% partial or pass
1.8M total tokensThinking used0 · 0.0%Answer generated0Graded5/5Provider-error tasks0.0%
+9 lessons
10 docs / 10 chunks
$1.93
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap
0 PASS1 PARTIAL4 FAIL
SMALL1/2 usefulMEDIUM0/2 usefulLARGE0/1 useful
AVERAGE SCORE2.24 / 51 is weak · 5 is excellent
CORRECT1.80 / 5STYLE2.60 / 5DECISIONS2.40 / 5REASONING2.00 / 5SCOPE2.40 / 5
1 / 520.0% partial or pass
1.1M total tokensThinking used0 · 0.0%Answer generated0Graded5/5Provider-error tasks0.0%
+10 lessons
5 docs / 5 chunks
$1.49
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap
0 PASS0 PARTIAL5 FAIL
SMALL0/1 usefulMEDIUM0/1 usefulLARGE0/3 useful
AVERAGE SCORE1.72 / 51 is weak · 5 is excellent
CORRECT1.00 / 5STYLE1.60 / 5DECISIONS2.00 / 5REASONING2.00 / 5SCOPE2.00 / 5
0 / 50.0% partial or pass
1.8M total tokensThinking used0 · 0.0%Answer generated0Graded5/5Provider-error tasks0.0%
+11 lessons
11 docs / 11 chunks
$2.22
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap
0 PASS0 PARTIAL5 FAIL
SMALL0/4 usefulMEDIUM0/1 usefulLARGE0/0 useful
AVERAGE SCORE2.16 / 51 is weak · 5 is excellent
CORRECT1.20 / 5STYLE3.00 / 5DECISIONS2.40 / 5REASONING2.00 / 5SCOPE2.20 / 5
0 / 50.0% partial or pass
1M total tokensThinking used0 · 0.0%Answer generated0Graded5/5Provider-error tasks0.0%
+4 lessons
6 docs / 6 chunks
$1.26
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap
0 PASS0 PARTIAL5 FAIL
SMALL0/0 usefulMEDIUM0/1 usefulLARGE0/4 useful
AVERAGE SCORE1.52 / 51 is weak · 5 is excellent
CORRECT1.00 / 5STYLE1.40 / 5DECISIONS2.00 / 5REASONING1.60 / 5SCOPE1.60 / 5
0 / 50.0% partial or pass
1.9M total tokensThinking used0 · 0.0%Answer generated0Graded5/5Provider-error tasks40.0%
+7 lessons
7 docs / 7 chunks
$1.84
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap
0 PASS4 PARTIAL14 FAIL
SMALL3/5 usefulMEDIUM0/3 usefulLARGE1/10 useful
AVERAGE SCORE2.48 / 51 is weak · 5 is excellent
CORRECT1.88 / 5STYLE2.63 / 5DECISIONS2.94 / 5REASONING2.56 / 5SCOPE2.38 / 5
4 / 1822.2% partial or pass
5.8M total tokensThinking used0 · 0.0%Answer generated0Graded16/18Provider-error tasks16.7%
+8 lessons
20 docs / 20 chunks
$8.26
3 ·WHY THIS ENGINE

HOW THE STUDENT MODEL WAS CHOSEN

One-off trials between candidate models, each sitting the same saved exams and graded the same way. They chose the engine and taught the Brain nothing. Open one to compare what worked, code quality, time and cost.

TEST 10CONFIGURATION TEST · 25 JUL 2026GLM-5.2 native Z.AI integration smoke · one examVIEW RESULTS
CONFIGURATION TESTGLM-5.2 native Z.AI integration smoke · one exam
25 JUL 20261 ranked attempts1 settings1 model
1 / 1 useful results0 pass · 1 partial · 0 fail

CURRENT TEST WORDING VERIFIED

MODELSETTINGSOUTCOMESTASK SIZESTEACHER AVERAGE5 QUALITY SCORESGRADEDENGINEER USAGEGRADING + TOTAL COST
GLM-5.2
GLM-5.2glm-5.2
LEADING RESULT
Reasoning levelThinking on · Max effort (provider default)PrecisionFP8Output cap33K tokens / replyTime limit900s
0 PASS1 PARTIAL0 FAIL
SMALL1/1 usefulMEDIUM0/0 usefulLARGE0/0 useful
AVERAGE SCORE3.20 / 51 is weak · 5 is excellent
CORRECT3.00 / 5STYLE3.00 / 5DECISIONS4.00 / 5REASONING3.00 / 5SCOPE3.00 / 5
1/1complete
212K total tokensWork used212KThinking used27K · 92.1%Answer generated2KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time7.6 min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops1Thinking-cap stops0Tool-flow errors1Unknown calls0
$0.33 grading
$0.33 total
$0.33 per task
0 unknown teacher calls
DITHERED MODEL EVIDENCEFour views of the same graded results
Legacy GLM 12K context is excluded because it was not a fair capability run.
BAR · OUTCOME-CHARTOutcome distributionSame 1 exams per setting
1 useful1 graded
011 usefulGLM-5.21 / 1 gradedGLM-5.2: 0 pass, 1 partial, 0 fail; 1 of 1 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135GLM-5.2: Correctness 3.00 / 5, Convention 3.00 / 5, Decisions 4.00 / 5, Reasoning 3.00 / 5, Scope 3.00 / 5CorrectnessConventionDecisionsReasoningScope
GLM-5.2
RADAR · RADAR-CHARTQuality profileFarther from the center is better
GLM-5.2: Correctness 3.00 / 5, Convention 3.00 / 5, Decisions 4.00 / 5, Reasoning 3.00 / 5, Scope 3.00 / 5CorrectnessConventionDecisionsReasoningScope
GLM-5.2
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Partial: 1 of 11 useful1 graded
Pass · 0Partial · 1Fail · 0Not graded · 0
TEST 09CONFIGURATION TEST · 24 JUL 2026GLM-5.2 fixed-harness final validation · Z.AI FP8VIEW RESULTS
CONFIGURATION TESTGLM-5.2 fixed-harness final validation · Z.AI FP8
24 JUL 20264 ranked attempts1 settings1 model
1 / 4 useful results1 pass · 0 partial · 3 fail

CURRENT TEST WORDING VERIFIED

MODELSETTINGSOUTCOMESTASK SIZESTEACHER AVERAGE5 QUALITY SCORESGRADEDENGINEER USAGEGRADING + TOTAL COST
GLM-5.2
GLM-5.2z-ai/glm-5.2
OUTCOME LEADER
Reasoning levelhigh reasoningPrecisionFP8Output cap33K tokens / replyTime limit900s
1 PASS0 PARTIAL3 FAIL
SMALL1/4 usefulMEDIUM0/0 usefulLARGE0/0 useful
AVERAGE SCORE3.10 / 51 is weak · 5 is excellent
CORRECT2.50 / 5STYLE4.25 / 5DECISIONS3.00 / 5REASONING2.75 / 5SCOPE3.00 / 5
4/4complete
927K total tokensWork used872KThinking used55K · 88.8%Answer generated7KWork allowance200K / examThinking allowance150K / examAverage time5.0 min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops4Thinking-cap stops0Tool-flow errors10Unknown calls0
$1.11 grading
$1.81 total
$0.45 per task
0 unknown teacher calls
DITHERED MODEL EVIDENCEFour views of the same graded results
Legacy GLM 12K context is excluded because it was not a fair capability run.
BAR · OUTCOME-CHARTOutcome distributionSame 4 exams per setting
1 useful4 graded
02341 usefulGLM-5.24 / 4 gradedGLM-5.2: 1 pass, 0 partial, 3 fail; 4 of 4 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135GLM-5.2: Correctness 2.50 / 5, Convention 4.25 / 5, Decisions 3.00 / 5, Reasoning 2.75 / 5, Scope 3.00 / 5CorrectnessConventionDecisionsReasoningScope
GLM-5.2
RADAR · RADAR-CHARTQuality profileFarther from the center is better
GLM-5.2: Correctness 2.50 / 5, Convention 4.25 / 5, Decisions 3.00 / 5, Reasoning 2.75 / 5, Scope 3.00 / 5CorrectnessConventionDecisionsReasoningScope
GLM-5.2
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Pass: 1 of 4Fail: 3 of 41 useful4 graded
Pass · 1Partial · 0Fail · 3Not graded · 0
TEST 08CONFIGURATION TEST · 24 JUL 2026GLM-5.2 reasoning and token-cap testVIEW RESULTS
CONFIGURATION TESTGLM-5.2 reasoning and token-cap test
24 JUL 20262 ranked attempts1 settings1 model
0 / 2 useful results0 pass · 0 partial · 2 fail

CURRENT TEST WORDING VERIFIED

MODELSETTINGSOUTCOMESTASK SIZESTEACHER AVERAGE5 QUALITY SCORESGRADEDENGINEER USAGEGRADING + TOTAL COST
GLM-5.2
GLM-5.2z-ai/glm-5.2
HIGHEST SCORE
Reasoning levelhigh reasoningPrecisionFP8Output cap33K tokens / replyTime limit900s
0 PASS0 PARTIAL2 FAIL
SMALL0/0 usefulMEDIUM0/0 usefulLARGE0/2 useful
AVERAGE SCORE2.40 / 51 is weak · 5 is excellent
CORRECT2.00 / 5STYLE2.50 / 5DECISIONS3.00 / 5REASONING2.50 / 5SCOPE2.00 / 5
2/2complete
920K total tokensWork used831KThinking used88K · 85.3%Answer generated15KWork allowance400K / examThinking allowance150K / examAverage time11.1 min/taskReply-limit stops1Timeout truncations0Provider-error tasks0.0%Work-cap stops2Thinking-cap stops0Tool-flow errors8Unknown calls0
$1.17 grading
$1.99 total
$0.99 per task
0 unknown teacher calls
DITHERED MODEL EVIDENCEFour views of the same graded results
Legacy GLM 12K context is excluded because it was not a fair capability run.
BAR · OUTCOME-CHARTOutcome distributionSame 2 exams per setting
0 useful2 graded
0120 usefulGLM-5.22 / 2 gradedGLM-5.2: 0 pass, 0 partial, 2 fail; 2 of 2 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135GLM-5.2: Correctness 2.00 / 5, Convention 2.50 / 5, Decisions 3.00 / 5, Reasoning 2.50 / 5, Scope 2.00 / 5CorrectnessConventionDecisionsReasoningScope
GLM-5.2
RADAR · RADAR-CHARTQuality profileFarther from the center is better
GLM-5.2: Correctness 2.00 / 5, Convention 2.50 / 5, Decisions 3.00 / 5, Reasoning 2.50 / 5, Scope 2.00 / 5CorrectnessConventionDecisionsReasoningScope
GLM-5.2
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Fail: 2 of 20 useful2 graded
Pass · 0Partial · 0Fail · 2Not graded · 0
TEST 07CONFIGURATION TEST · 24 JUL 2026GLM-5.2 quality probe v3 · Z.AI FP8VIEW RESULTS
CONFIGURATION TESTGLM-5.2 quality probe v3 · Z.AI FP8
24 JUL 20263 ranked attempts1 settings1 model
1 / 3 useful results0 pass · 1 partial · 2 fail

CURRENT TEST WORDING VERIFIED

MODELSETTINGSOUTCOMESTASK SIZESTEACHER AVERAGE5 QUALITY SCORESGRADEDENGINEER USAGEGRADING + TOTAL COST
GLM-5.2
GLM-5.2z-ai/glm-5.2
LEADING RESULT
Reasoning levelhigh reasoningPrecisionFP8Output cap33K tokens / replyTime limit900s
0 PASS1 PARTIAL2 FAIL
SMALL1/1 usefulMEDIUM0/0 usefulLARGE0/2 useful
AVERAGE SCORE2.13 / 51 is weak · 5 is excellent
CORRECT2.00 / 5STYLE2.33 / 5DECISIONS2.33 / 5REASONING2.00 / 5SCOPE2.00 / 5
3/3complete
1.1M total tokensWork used1MThinking used111K · 85.4%Answer generated19KWork allowance200K / 400K / examThinking allowance150K / examAverage time9.5 min/taskReply-limit stops1Timeout truncations0Provider-error tasks0.0%Work-cap stops2Thinking-cap stops0Tool-flow errors11Unknown calls0
$1.17 grading
$2.18 total
$0.73 per task
0 unknown teacher calls
DITHERED MODEL EVIDENCEFour views of the same graded results
Legacy GLM 12K context is excluded because it was not a fair capability run.
BAR · OUTCOME-CHARTOutcome distributionSame 3 exams per setting
1 useful3 graded
01231 usefulGLM-5.23 / 3 gradedGLM-5.2: 0 pass, 1 partial, 2 fail; 3 of 3 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135GLM-5.2: Correctness 2.00 / 5, Convention 2.33 / 5, Decisions 2.33 / 5, Reasoning 2.00 / 5, Scope 2.00 / 5CorrectnessConventionDecisionsReasoningScope
GLM-5.2
RADAR · RADAR-CHARTQuality profileFarther from the center is better
GLM-5.2: Correctness 2.00 / 5, Convention 2.33 / 5, Decisions 2.33 / 5, Reasoning 2.00 / 5, Scope 2.00 / 5CorrectnessConventionDecisionsReasoningScope
GLM-5.2
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Partial: 1 of 3Fail: 2 of 31 useful3 graded
Pass · 0Partial · 1Fail · 2Not graded · 0
TEST 06CONFIGURATION TEST · 24 JUL 2026GLM-5.2 quality probe · Z.AI FP8VIEW RESULTS
CONFIGURATION TESTGLM-5.2 quality probe · Z.AI FP8
24 JUL 20263 ranked attempts1 settings1 model
0 / 3 useful results0 pass · 0 partial · 3 fail

CURRENT TEST WORDING VERIFIED

MODELSETTINGSOUTCOMESTASK SIZESTEACHER AVERAGE5 QUALITY SCORESGRADEDENGINEER USAGEGRADING + TOTAL COST
GLM-5.2
GLM-5.2z-ai/glm-5.2
HIGHEST SCORE
Reasoning levelhigh reasoningPrecisionFP8Output cap33K tokens / replyTime limit900s
0 PASS0 PARTIAL3 FAIL
SMALL0/1 usefulMEDIUM0/0 usefulLARGE0/2 useful
AVERAGE SCORE2.00 / 51 is weak · 5 is excellent
CORRECT1.00 / 5STYLE1.67 / 5DECISIONS2.67 / 5REASONING2.00 / 5SCOPE2.67 / 5
3/3complete
916K total tokensWork used817KThinking used98K · 86.2%Answer generated16KWork allowance200K / 400K / examThinking allowance150K / examAverage time8.5 min/taskReply-limit stops1Timeout truncations0Provider-error tasks0.0%Work-cap stops0Thinking-cap stops0Tool-flow errors17Unknown calls0
$0.98 grading
$1.86 total
$0.62 per task
0 unknown teacher calls
DITHERED MODEL EVIDENCEFour views of the same graded results
Legacy GLM 12K context is excluded because it was not a fair capability run.
BAR · OUTCOME-CHARTOutcome distributionSame 3 exams per setting
0 useful3 graded
01230 usefulGLM-5.23 / 3 gradedGLM-5.2: 0 pass, 0 partial, 3 fail; 3 of 3 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135GLM-5.2: Correctness 1.00 / 5, Convention 1.67 / 5, Decisions 2.67 / 5, Reasoning 2.00 / 5, Scope 2.67 / 5CorrectnessConventionDecisionsReasoningScope
GLM-5.2
RADAR · RADAR-CHARTQuality profileFarther from the center is better
GLM-5.2: Correctness 1.00 / 5, Convention 1.67 / 5, Decisions 2.67 / 5, Reasoning 2.00 / 5, Scope 2.67 / 5CorrectnessConventionDecisionsReasoningScope
GLM-5.2
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Fail: 3 of 30 useful3 graded
Pass · 0Partial · 0Fail · 3Not graded · 0
OLDER COMPARISONS · 5 MORE ON RECORD

Earlier model and setting trials, same exams and same grading. Kept for the record, folded away because the choice they informed is already made.

TEST 05FAIR-BUDGET HEAD-TO-HEAD · 24 JUL 2026GLM and Kimi · equal work and thinking limitsVIEW RESULTS
FAIR-BUDGET HEAD-TO-HEADGLM and Kimi · equal work and thinking limits
24 JUL 202616 of 48 planned attempts2 settings1 model
16 / 48 attempts · no current winner0 pass · 3 partial · 13 fail
MODELSETTINGSOUTCOMESTASK SIZESTEACHER AVERAGE5 QUALITY SCORESGRADEDENGINEER USAGEGRADING + TOTAL COST
GLM-5.2
GLM-5.2z-ai/glm-5.2
PARTIAL DATA
Reasoning levelThinking on · Max effort (provider default)PrecisionHost-dependentOutput cap66K tokens / replyTime limit900s
0 PASS3 PARTIAL9 FAIL
SMALL2/7 usefulMEDIUM0/2 usefulLARGE1/3 useful
AVERAGE SCORE2.28 / 51 is weak · 5 is excellent
CORRECT1.90 / 5STYLE2.90 / 5DECISIONS2.40 / 5REASONING2.20 / 5SCOPE2.00 / 5
10/122 ungraded
3.3M total tokensWork used2.7MThinking used507K · 96.3%Answer generated20KWork allowance200K / 400K / examThinking allowance150K / examAverage time17.0 min/taskReply-limit stops5Timeout truncations0Provider-error tasks41.7%Work-cap stops5Thinking-cap stops2Tool-flow errors36Unknown calls7
$2.80 grading
$5.01 total
$0.42 per task
0 unknown teacher calls
GLM-5.2
GLM-5.2z-ai/glm-5.2
PARTIAL DATA
Reasoning levelxhigh reasoningPrecisionHost-dependentOutput cap66K tokens / replyTime limit900s
0 PASS0 PARTIAL4 FAIL
SMALL0/2 usefulMEDIUM0/2 usefulLARGE0/0 useful
AVERAGE SCORE1.85 / 51 is weak · 5 is excellent
CORRECT1.25 / 5STYLE2.25 / 5DECISIONS2.25 / 5REASONING1.75 / 5SCOPE1.75 / 5
4/4complete
876K total tokensWork used784KThinking used92K · 91.6%Answer generated8KWork allowance200K / examThinking allowance150K / examAverage time13.1 min/taskReply-limit stops0Timeout truncations0Provider-error tasks50.0%Work-cap stops2Thinking-cap stops0Tool-flow errors13Unknown calls2
$1.07 grading
$1.59 total
$0.40 per task
0 unknown teacher calls
DITHERED MODEL EVIDENCEFour views of the same graded results
Partial results — charts update as more tasks are graded.
BAR · OUTCOME-CHARTOutcome distributionAvailable graded attempts only
3 useful16 graded
048123 usefulGLM-5.2 default12 / 12 gradedGLM-5.2 · provider-default reasoning: 0 pass, 3 partial, 9 fail; 12 of 12 graded0 usefulGLM-5.2 max4 / 12 gradedGLM-5.2 · maximum reasoning: 0 pass, 0 partial, 4 fail; 4 of 12 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135GLM-5.2 · provider-default reasoning: Correctness 1.90 / 5, Convention 2.90 / 5, Decisions 2.40 / 5, Reasoning 2.20 / 5, Scope 2.00 / 5GLM-5.2 · maximum reasoning: Correctness 1.25 / 5, Convention 2.25 / 5, Decisions 2.25 / 5, Reasoning 1.75 / 5, Scope 1.75 / 5CorrectnessConventionDecisionsReasoningScope
GLM-5.2 · provider-default reasoningGLM-5.2 · maximum reasoning
RADAR · RADAR-CHARTQuality profileFarther from the center is better
GLM-5.2 · provider-default reasoning: Correctness 1.90 / 5, Convention 2.90 / 5, Decisions 2.40 / 5, Reasoning 2.20 / 5, Scope 2.00 / 5GLM-5.2 · maximum reasoning: Correctness 1.25 / 5, Convention 2.25 / 5, Decisions 2.25 / 5, Reasoning 1.75 / 5, Scope 1.75 / 5CorrectnessConventionDecisionsReasoningScope
GLM-5.2 · provider-default reasoningGLM-5.2 · maximum reasoning
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Partial: 3 of 24Fail: 13 of 24Not graded: 8 of 243 useful16 graded
Pass · 0Partial · 3Fail · 13Not graded · 8
TEST 04DEFAULT REASONING BASELINE · 24 JUL 2026Provider-default baseline · deterministically regradedVIEW RESULTS
DEFAULT REASONING BASELINEProvider-default baseline · deterministically regraded
24 JUL 202627 ranked attempts3 settings3 models
5 / 27 useful results0 pass · 5 partial · 22 fail

CURRENT TEST WORDING VERIFIED

MODELSETTINGSOUTCOMESTASK SIZESTEACHER AVERAGE5 QUALITY SCORESGRADEDENGINEER USAGEGRADING + TOTAL COST
GLM-5.2
GLM-5.2z-ai/glm-5.2
LEADING RESULT
Reasoning levelprovider-default reasoningPrecisionHost-dependentOutput cap32K tokens / replyTime limit300s
0 PASS2 PARTIAL7 FAIL
SMALL1/4 usefulMEDIUM0/2 usefulLARGE1/3 useful
AVERAGE SCORE2.56 / 51 is weak · 5 is excellent
CORRECT1.89 / 5STYLE2.33 / 5DECISIONS3.11 / 5REASONING2.56 / 5SCOPE2.89 / 5
9/9complete
2M total tokensWork used2MThinking used101K · 82.0%Answer generated22KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time5.4 min/taskReply-limit stops0Timeout truncations0Provider-error tasks22.2%Work-cap stops0Thinking-cap stops0Tool-flow errors19Unknown calls2
$2.78 grading
$3.63 total
$0.40 per task
0 unknown teacher calls
Kimi K2.7 Code
Kimi K2.7 Codemoonshotai/kimi-k2.7-code
RANK 2
Reasoning levelprovider-default reasoningPrecisionHost-dependentOutput cap32K tokens / replyTime limit300s
0 PASS2 PARTIAL7 FAIL
SMALL2/4 usefulMEDIUM0/2 usefulLARGE0/3 useful
AVERAGE SCORE2.13 / 51 is weak · 5 is excellent
CORRECT1.89 / 5STYLE2.44 / 5DECISIONS2.33 / 5REASONING2.11 / 5SCOPE1.89 / 5
9/9complete
2.3M total tokensWork used2.3MThinking used32K · 64.6%Answer generated18KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time1.7 min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops0Thinking-cap stops0Tool-flow errors16Unknown calls0
$2.60 grading
$3.37 total
$0.37 per task
0 unknown teacher calls
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
RANK 3
Reasoning levelprovider-default reasoningPrecisionHost-dependentOutput cap12K tokens / replyTime limit300s
0 PASS1 PARTIAL8 FAIL
SMALL1/4 usefulMEDIUM0/2 usefulLARGE0/3 useful
AVERAGE SCORE2.05 / 51 is weak · 5 is excellent
CORRECT1.50 / 5STYLE2.38 / 5DECISIONS2.50 / 5REASONING1.88 / 5SCOPE2.00 / 5
8/91 ungraded
2.4M total tokensWork used2.4MThinking used27K · 45.4%Answer generated32KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time0.6 min/taskReply-limit stops0Timeout truncations0Provider-error tasks22.2%Work-cap stops1Thinking-cap stops0Tool-flow errors4Unknown calls2
$2.30 grading
$2.90 total
$0.32 per task
0 unknown teacher calls
DITHERED MODEL EVIDENCEFour views of the same graded results
Legacy GLM 12K context is excluded because it was not a fair capability run.
BAR · OUTCOME-CHARTOutcome distributionSame 9 exams per setting
5 useful27 graded
03692 usefulGLM-5.29 / 9 gradedGLM-5.2: 0 pass, 2 partial, 7 fail; 9 of 9 graded2 usefulKimi9 / 9 gradedKimi K2.7 Code: 0 pass, 2 partial, 7 fail; 9 of 9 graded1 usefulDeepSeek9 / 9 gradedDeepSeek V4 Pro: 0 pass, 1 partial, 8 fail; 9 of 9 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135GLM-5.2: Correctness 1.89 / 5, Convention 2.33 / 5, Decisions 3.11 / 5, Reasoning 2.56 / 5, Scope 2.89 / 5Kimi K2.7 Code: Correctness 1.89 / 5, Convention 2.44 / 5, Decisions 2.33 / 5, Reasoning 2.11 / 5, Scope 1.89 / 5DeepSeek V4 Pro: Correctness 1.50 / 5, Convention 2.38 / 5, Decisions 2.50 / 5, Reasoning 1.88 / 5, Scope 2.00 / 5CorrectnessConventionDecisionsReasoningScope
GLM-5.2Kimi K2.7 CodeDeepSeek V4 Pro
RADAR · RADAR-CHARTQuality profileFarther from the center is better
GLM-5.2: Correctness 1.89 / 5, Convention 2.33 / 5, Decisions 3.11 / 5, Reasoning 2.56 / 5, Scope 2.89 / 5Kimi K2.7 Code: Correctness 1.89 / 5, Convention 2.44 / 5, Decisions 2.33 / 5, Reasoning 2.11 / 5, Scope 1.89 / 5DeepSeek V4 Pro: Correctness 1.50 / 5, Convention 2.38 / 5, Decisions 2.50 / 5, Reasoning 1.88 / 5, Scope 2.00 / 5CorrectnessConventionDecisionsReasoningScope
GLM-5.2Kimi K2.7 CodeDeepSeek V4 Pro
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Partial: 5 of 27Fail: 22 of 275 useful27 graded
Pass · 0Partial · 5Fail · 22Not graded · 0
TEST 03MAXIMUM REASONING TEST · 24 JUL 2026Maximum-reasoning model comparisonVIEW RESULTS
MAXIMUM REASONING TESTMaximum-reasoning model comparison
24 JUL 202627 ranked attempts3 settings3 models
6 / 27 useful results1 pass · 5 partial · 21 fail

CURRENT TEST WORDING VERIFIED

TWO DIFFERENT LEADERS — Kimi is the outcome leader because this ranking counts full passes first, and Kimi earned the only pass. GLM is the average-score leader (2.26 vs 2.00). One pass in nine tasks does not prove Kimi is generally better.

MODELSETTINGSOUTCOMESTASK SIZESTEACHER AVERAGE5 QUALITY SCORESGRADEDENGINEER USAGEGRADING + TOTAL COST
Kimi K2.7 Code
Kimi K2.7 Codemoonshotai/kimi-k2.7-code
OUTCOME LEADER
Reasoning levelxhigh reasoningPrecisionHost-dependentOutput cap66K tokens / replyTime limit900s
1 PASS1 PARTIAL7 FAIL
SMALL2/4 usefulMEDIUM0/2 usefulLARGE0/3 useful
AVERAGE SCORE2.00 / 51 is weak · 5 is excellent
CORRECT1.78 / 5STYLE2.11 / 5DECISIONS2.22 / 5REASONING2.00 / 5SCOPE1.89 / 5
9/9complete
2.4M total tokensWork used2.4MThinking used32K · 66.3%Answer generated17KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time1.6 min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops0Thinking-cap stops0Tool-flow errors18Unknown calls0
$2.31 grading
$3.16 total
$0.35 per task
0 unknown teacher calls
GLM-5.2
GLM-5.2z-ai/glm-5.2
RANK 2
Reasoning levelxhigh reasoningPrecisionHost-dependentOutput cap32K / 66K tokens / replyTime limit300 / 900s
0 PASS2 PARTIAL7 FAIL
SMALL1/4 usefulMEDIUM0/2 usefulLARGE1/3 useful
AVERAGE SCORE2.26 / 51 is weak · 5 is excellent
CORRECT1.89 / 5STYLE2.22 / 5DECISIONS2.67 / 5REASONING2.22 / 5SCOPE2.33 / 5
9/9complete
2.6M total tokensWork used2.4MThinking used328K · 90.7%Answer generated34KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time11.8 min/taskReply-limit stops1Timeout truncations2Provider-error tasks11.1%Work-cap stops0Thinking-cap stops0Tool-flow errors39Unknown calls1
$2.79 grading
$4.50 total
$0.50 per task
0 unknown teacher calls
$0.08 aborted/no verdict · 140K tokens
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
RANK 3
Reasoning levelxhigh reasoningPrecisionHost-dependentOutput cap32K tokens / replyTime limit300s
0 PASS2 PARTIAL7 FAIL
SMALL1/4 usefulMEDIUM0/2 usefulLARGE1/3 useful
AVERAGE SCORE2.06 / 51 is weak · 5 is excellent
CORRECT1.83 / 5STYLE2.67 / 5DECISIONS2.00 / 5REASONING2.00 / 5SCOPE1.83 / 5
6/93 ungraded
2.5M total tokensWork used2.5MThinking used42K · 59.6%Answer generated29KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time2.6 min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops3Thinking-cap stops0Tool-flow errors24Unknown calls0
$1.80 grading
$2.45 total
$0.27 per task
0 unknown teacher calls
DITHERED MODEL EVIDENCEFour views of the same graded results
Legacy GLM 12K context is excluded because it was not a fair capability run.
BAR · OUTCOME-CHARTOutcome distributionSame 9 exams per setting
6 useful27 graded
03692 usefulKimi9 / 9 gradedKimi K2.7 Code: 1 pass, 1 partial, 7 fail; 9 of 9 graded2 usefulGLM-5.29 / 9 gradedGLM-5.2: 0 pass, 2 partial, 7 fail; 9 of 9 graded2 usefulDeepSeek9 / 9 gradedDeepSeek V4 Pro: 0 pass, 2 partial, 7 fail; 9 of 9 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135Kimi K2.7 Code: Correctness 1.78 / 5, Convention 2.11 / 5, Decisions 2.22 / 5, Reasoning 2.00 / 5, Scope 1.89 / 5GLM-5.2: Correctness 1.89 / 5, Convention 2.22 / 5, Decisions 2.67 / 5, Reasoning 2.22 / 5, Scope 2.33 / 5DeepSeek V4 Pro: Correctness 1.83 / 5, Convention 2.67 / 5, Decisions 2.00 / 5, Reasoning 2.00 / 5, Scope 1.83 / 5CorrectnessConventionDecisionsReasoningScope
Kimi K2.7 CodeGLM-5.2DeepSeek V4 Pro
RADAR · RADAR-CHARTQuality profileFarther from the center is better
Kimi K2.7 Code: Correctness 1.78 / 5, Convention 2.11 / 5, Decisions 2.22 / 5, Reasoning 2.00 / 5, Scope 1.89 / 5GLM-5.2: Correctness 1.89 / 5, Convention 2.22 / 5, Decisions 2.67 / 5, Reasoning 2.22 / 5, Scope 2.33 / 5DeepSeek V4 Pro: Correctness 1.83 / 5, Convention 2.67 / 5, Decisions 2.00 / 5, Reasoning 2.00 / 5, Scope 1.83 / 5CorrectnessConventionDecisionsReasoningScope
Kimi K2.7 CodeGLM-5.2DeepSeek V4 Pro
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Pass: 1 of 27Partial: 5 of 27Fail: 21 of 276 useful27 graded
Pass · 1Partial · 5Fail · 21Not graded · 0
TEST 02MODEL COMPARISON · 23 JUL 2026Coding model capability comparisonVIEW RESULTS
MODEL COMPARISONCoding model capability comparison
23 JUL 202627 ranked attempts4 settings3 models
27 saved results · no current winnerOld labels: 1 pass · 5 partial · 21 fail · corrected total $12.14

AUDIT HOLD — These saved results used a grading rule that treated equivalent work differently. They remain visible for history but are not ranked until deterministic regrading.

MODELSETTINGSOUTCOMESTASK SIZESTEACHER AVERAGE5 QUALITY SCORESGRADEDENGINEER USAGEGRADING + TOTAL COST
Kimi K2.7 Code
Kimi K2.7 Codemoonshotai/kimi-k2.7-code
HISTORICAL LABEL · PENDING REGRADE
Reasoning levelProvider defaultPrecisionHost-dependentOutput cap32K tokens / replyTime limit300s
1 PASS1 PARTIAL7 FAIL
SMALL2/4 usefulMEDIUM0/2 usefulLARGE0/3 useful
AVERAGE SCORE2.17 / 51 is weak · 5 is excellent
CORRECT1.89 / 5STYLE2.22 / 5DECISIONS2.33 / 5REASONING2.11 / 5SCOPE2.33 / 5
9/9complete
2.3M total tokensWork used2.3MThinking used32K · 64.6%Answer generated18KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time1.7 min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops0Thinking-cap stops0Tool-flow errors16Unknown calls0
$2.81 grading
$3.57 total
$0.40 per task
0 unknown teacher calls
GLM-5.2
GLM-5.2z-ai/glm-5.2
HISTORICAL LABEL · PENDING REGRADE
Reasoning levelThinking on · Max effort (provider default)PrecisionHost-dependentOutput cap32K tokens / replyTime limit300s
0 PASS2 PARTIAL7 FAIL
SMALL2/4 usefulMEDIUM0/2 usefulLARGE0/3 useful
AVERAGE SCORE2.71 / 51 is weak · 5 is excellent
CORRECT1.78 / 5STYLE2.44 / 5DECISIONS3.33 / 5REASONING2.67 / 5SCOPE3.33 / 5
9/9complete
2M total tokensWork used2MThinking used101K · 82.0%Answer generated22KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time5.4 min/taskReply-limit stops0Timeout truncations0Provider-error tasks22.2%Work-cap stops0Thinking-cap stops0Tool-flow errors19Unknown calls2
$2.87 grading
$3.71 total
$0.41 per task
0 unknown teacher calls
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
HISTORICAL LABEL · PENDING REGRADE
Reasoning levelProvider defaultPrecisionHost-dependentOutput cap12K tokens / replyTime limit300s
0 PASS2 PARTIAL7 FAIL
SMALL1/4 usefulMEDIUM0/2 usefulLARGE1/3 useful
AVERAGE SCORE2.08 / 51 is weak · 5 is excellent
CORRECT1.63 / 5STYLE2.38 / 5DECISIONS2.38 / 5REASONING2.13 / 5SCOPE1.88 / 5
8/91 ungraded
2.4M total tokensWork used2.4MThinking used27K · 45.4%Answer generated32KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time0.6 min/taskReply-limit stops0Timeout truncations0Provider-error tasks22.2%Work-cap stops1Thinking-cap stops0Tool-flow errors22Unknown calls2
$2.20 grading
$2.80 total
$0.31 per task
0 unknown teacher calls
GLM-5.2
GLM-5.2z-ai/glm-5.2
CONTEXT ONLY · NOT CAPABILITY RANKED
Reasoning levelThinking on · Max effort (provider default)PrecisionHost-dependentOutput cap12K tokens / replyTime limit300s
0 PASS0 PARTIAL5 FAIL
SMALL0/2 usefulMEDIUM0/2 usefulLARGE0/1 useful
AVERAGE SCORE1.40 / 51 is weak · 5 is excellent
CORRECT1.50 / 5STYLE2.00 / 5DECISIONS1.00 / 5REASONING1.50 / 5SCOPE1.00 / 5
2/53 ungraded
1.3M total tokensWork used1.3MThinking used132K · 99.6%Answer generated2KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time— min/taskReply-limit stops11Timeout truncations0Provider-error tasks0.0%Work-cap stops3Thinking-cap stops0Tool-flow errors15Unknown calls0
$0.46 grading
$1.15 total
$0.23 per task
0 unknown teacher calls
$0.90 aborted/no verdict · 362K tokens
DITHERED MODEL EVIDENCEFour views of the same graded results
Historical labels only — the deterministic regrade will replace these outcome rankings.
BAR · OUTCOME-CHARTOutcome distributionSame 9 exams per setting
6 useful27 graded
03692 usefulKimi9 / 9 gradedKimi K2.7 Code: 1 pass, 1 partial, 7 fail; 9 of 9 graded2 usefulGLM-5.29 / 9 gradedGLM-5.2: 0 pass, 2 partial, 7 fail; 9 of 9 graded2 usefulDeepSeek9 / 9 gradedDeepSeek V4 Pro: 0 pass, 2 partial, 7 fail; 9 of 9 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135Kimi K2.7 Code: Correctness 1.89 / 5, Convention 2.22 / 5, Decisions 2.33 / 5, Reasoning 2.11 / 5, Scope 2.33 / 5GLM-5.2: Correctness 1.78 / 5, Convention 2.44 / 5, Decisions 3.33 / 5, Reasoning 2.67 / 5, Scope 3.33 / 5DeepSeek V4 Pro: Correctness 1.63 / 5, Convention 2.38 / 5, Decisions 2.38 / 5, Reasoning 2.13 / 5, Scope 1.88 / 5CorrectnessConventionDecisionsReasoningScope
Kimi K2.7 CodeGLM-5.2DeepSeek V4 Pro
RADAR · RADAR-CHARTQuality profileFarther from the center is better
Kimi K2.7 Code: Correctness 1.89 / 5, Convention 2.22 / 5, Decisions 2.33 / 5, Reasoning 2.11 / 5, Scope 2.33 / 5GLM-5.2: Correctness 1.78 / 5, Convention 2.44 / 5, Decisions 3.33 / 5, Reasoning 2.67 / 5, Scope 3.33 / 5DeepSeek V4 Pro: Correctness 1.63 / 5, Convention 2.38 / 5, Decisions 2.38 / 5, Reasoning 2.13 / 5, Scope 1.88 / 5CorrectnessConventionDecisionsReasoningScope
Kimi K2.7 CodeGLM-5.2DeepSeek V4 Pro
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Pass: 1 of 27Partial: 5 of 27Fail: 21 of 276 useful27 graded
Pass · 1Partial · 5Fail · 21Not graded · 0
TEST 01CONFIGURATION TEST · 23 JUL 2026DeepSeek V4 Pro reasoning and token-cap testVIEW RESULTS
CONFIGURATION TESTDeepSeek V4 Pro reasoning and token-cap test
23 JUL 202615 ranked attempts3 settings1 model
0 / 15 useful results0 pass · 0 partial · 15 fail

OLDER EXPERIMENT — The exact task wording was not saved with these results. Use them to understand this older test, not to rank new models against it.

MODELSETTINGSOUTCOMESTASK SIZESTEACHER AVERAGE5 QUALITY SCORESGRADEDENGINEER USAGEGRADING + TOTAL COST
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
HIGHEST SCORE
Reasoning levelProvider defaultPrecisionHost-dependentOutput cap12K tokens / replyTime limit300s
0 PASS0 PARTIAL5 FAIL
SMALL0/2 usefulMEDIUM0/2 usefulLARGE0/1 useful
AVERAGE SCORE2.30 / 51 is weak · 5 is excellent
CORRECT1.25 / 5STYLE2.25 / 5DECISIONS2.75 / 5REASONING2.00 / 5SCOPE3.25 / 5
4/51 ungraded
2.2M total tokensWork used2.2MThinking used10K · 38.3%Answer generated17KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time— min/taskReply-limit stops0Timeout truncations0Provider-error tasks20.0%Work-cap stops1Thinking-cap stops0Tool-flow errors0Unknown calls0
$1.17 grading
$1.47 total
$0.29 per task
1 unknown teacher calls
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
RANK 2
Reasoning levelhigh reasoningPrecisionHost-dependentOutput cap12K tokens / replyTime limit300s
0 PASS0 PARTIAL5 FAIL
SMALL0/2 usefulMEDIUM0/2 usefulLARGE0/1 useful
AVERAGE SCORE2.08 / 51 is weak · 5 is excellent
CORRECT1.20 / 5STYLE2.00 / 5DECISIONS2.60 / 5REASONING2.20 / 5SCOPE2.40 / 5
5/5complete
2.1M total tokensWork used2.1MThinking used12K · 33.4%Answer generated24KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time— min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops0Thinking-cap stops0Tool-flow errors0Unknown calls0
$1.47 grading
$1.77 total
$0.35 per task
0 unknown teacher calls
DeepSeek V4 Pro
DeepSeek V4 Prodeepseek/deepseek-v4-pro
RANK 3
Reasoning levelhigh reasoningPrecisionHost-dependentOutput cap12K tokens / replyTime limit300s
0 PASS0 PARTIAL5 FAIL
SMALL0/2 usefulMEDIUM0/2 usefulLARGE0/1 useful
AVERAGE SCORE1.40 / 51 is weak · 5 is excellent
CORRECT1.00 / 5STYLE2.00 / 5DECISIONS1.50 / 5REASONING1.50 / 5SCOPE1.00 / 5
4/51 ungraded
1.3M total tokensWork used1.3MThinking used7K · 26.9%Answer generated19KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time— min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops1Thinking-cap stops0Tool-flow errors0Unknown calls0
$1.09 grading
$1.32 total
$0.26 per task
0 unknown teacher calls
DITHERED MODEL EVIDENCEFour views of the same graded results
Legacy GLM 12K context is excluded because it was not a fair capability run.
BAR · OUTCOME-CHARTOutcome distributionSame 5 exams per setting
0 useful15 graded
02450 usefulDeepSeek5 / 5 gradedDeepSeek V4 Pro: 0 pass, 0 partial, 5 fail; 5 of 5 graded0 usefulDeepSeek5 / 5 gradedDeepSeek V4 Pro: 0 pass, 0 partial, 5 fail; 5 of 5 graded0 usefulDeepSeek5 / 5 gradedDeepSeek V4 Pro: 0 pass, 0 partial, 5 fail; 5 of 5 gradedPassPartialFailNot graded
LINE · AREA-CHARTQuality across the five checksTeacher score · 1 is weak · 5 is excellent
135DeepSeek V4 Pro: Correctness 1.25 / 5, Convention 2.25 / 5, Decisions 2.75 / 5, Reasoning 2.00 / 5, Scope 3.25 / 5DeepSeek V4 Pro: Correctness 1.20 / 5, Convention 2.00 / 5, Decisions 2.60 / 5, Reasoning 2.20 / 5, Scope 2.40 / 5DeepSeek V4 Pro: Correctness 1.00 / 5, Convention 2.00 / 5, Decisions 1.50 / 5, Reasoning 1.50 / 5, Scope 1.00 / 5CorrectnessConventionDecisionsReasoningScope
DeepSeek V4 ProDeepSeek V4 ProDeepSeek V4 Pro
RADAR · RADAR-CHARTQuality profileFarther from the center is better
DeepSeek V4 Pro: Correctness 1.25 / 5, Convention 2.25 / 5, Decisions 2.75 / 5, Reasoning 2.00 / 5, Scope 3.25 / 5DeepSeek V4 Pro: Correctness 1.20 / 5, Convention 2.00 / 5, Decisions 2.60 / 5, Reasoning 2.20 / 5, Scope 2.40 / 5DeepSeek V4 Pro: Correctness 1.00 / 5, Convention 2.00 / 5, Decisions 1.50 / 5, Reasoning 1.50 / 5, Scope 1.00 / 5CorrectnessConventionDecisionsReasoningScope
DeepSeek V4 ProDeepSeek V4 ProDeepSeek V4 Pro
PIE · PIE-CHARTOverall result mixAll planned attempts in this comparison
Fail: 15 of 150 useful15 graded
Pass · 0Partial · 0Fail · 15Not graded · 0
4 ·CURRENT STATE

HOW IS THE STUDENT DOING?

How the student is doing right now. Only results whose exact task wording and reference answer were saved are counted, so two versions of an exam are never compared as if they were the same.

INDEPENDENT EXAMS
0.0%0 OF 9 RESERVED TASKS PASSED
TARGET NOT MET
AUTOMATED CHECKS
96.2%25 OF 26 AUTOMATED CHECKS PASSED
TARGET MET
PROJECT FIT
3.6 / 572% OF THE 5-POINT MAXIMUM
TARGET NOT MET
BRAIN HELP FOUND
NO FIXED-PANEL TASK IN THIS CYCLE
NO COMPARABLE CURRENT RESULT
CONFLICTING GUIDANCE
00 CONFLICTS FOUND
TARGET MET
KNOWLEDGE UNITS370Documents the Brain can retrieve from
SEARCHABLE PIECES522 / 522Chunks with a vector embedding
CONNECTIONS356Links between units used for structural recall
LESSONS LEARNED+139Applied across 21 completed cycles
USEFUL RESULTS23 / 140Partial or pass, every cycle on record
DITHERED BRAIN EVIDENCEEverything the Brain has learned, and what it scored
Bars are the running total of curated lessons; the line is graded correctness. Both are drawn across every completed cycle.
12345Cycle 3: 9 lessons in the Brain (+9 this cycle)Cycle 10: 21 lessons in the Brain (+12 this cycle)Cycle 12: 29 lessons in the Brain (+8 this cycle)Cycle 13: 36 lessons in the Brain (+7 this cycle)Cycle 14: 44 lessons in the Brain (+8 this cycle)Cycle 15: 49 lessons in the Brain (+5 this cycle)Cycle 16: 58 lessons in the Brain (+9 this cycle)Cycle 17: 68 lessons in the Brain (+10 this cycle)Cycle 18: 79 lessons in the Brain (+11 this cycle)Cycle 19: 83 lessons in the Brain (+4 this cycle)Cycle 21: 90 lessons in the Brain (+7 this cycle)Cycle 22: 98 lessons in the Brain (+8 this cycle)Cycle 25: 108 lessons in the Brain (+10 this cycle)Cycle 26: 117 lessons in the Brain (+9 this cycle)Cycle 27: 123 lessons in the Brain (+6 this cycle)Cycle 28: 131 lessons in the Brain (+8 this cycle)Cycle 29: 139 lessons in the Brain (+8 this cycle)Cycle 30: 139 lessons in the Brain (+0 this cycle)Cycle 31: 139 lessons in the Brain (+0 this cycle)Cycle 32: 139 lessons in the Brain (+0 this cycle)Cycle 33: 139 lessons in the Brain (+0 this cycle)Cycle 3: correctness 1.20 / 5Cycle 10: correctness 1.20 / 5Cycle 12: correctness 1.75 / 5Cycle 13: correctness 1.53 / 5Cycle 14: correctness 1.75 / 5Cycle 15: correctness 1.50 / 5Cycle 16: correctness 1.40 / 5Cycle 17: correctness 1.80 / 5Cycle 18: correctness 1.00 / 5Cycle 19: correctness 1.20 / 5Cycle 21: correctness 1.00 / 5Cycle 22: correctness 1.88 / 5Cycle 25: correctness 1.00 / 5Cycle 26: correctness 2.40 / 5Cycle 27: correctness 1.75 / 5Cycle 28: correctness 3.00 / 5Cycle 29: correctness 2.40 / 5Cycle 30: correctness 1.20 / 5Cycle 31: correctness 2.00 / 5Cycle 32: correctness 3.00 / 5Cycle 33: correctness 3.20 / 5◀ DeepSeek V4 ProGLM-5.2 ▶07013931012131415161718192122252627282930313233LESSONSSCORE · 1-5CYCLE
Lessons in the Brain (running total)Correctness on GLM-5.2Correctness before the model change
The number to watch is whether the Brain reaches the engineer. On the locked guide list it surfaced the right unit 28.1% of the time in the latest cycle. The Brain keeps growing; getting the right unit in front of the engineer at the right moment is the open problem, and it is what this phase exists to fix.
5 ·RETRIEVAL QUALITY

CAN THE BRAIN FIND ITS OWN ANSWERS? A fixed list of questions whose right answers are already known. Retrieval runs each one against the live index and is marked on whether the right documents came back, and how high. It costs nothing to run, so a ranking change can be judged immediately.

Everything above depends on search working. If the right guide never reaches the engineer, no amount of grading will fix the answer. This section measures search on its own, against known-correct answers, so a ranking change can be judged the moment it lands instead of waiting for the next paid cycle.

A fixed list of 34 questions with known right answers, every one taken from something that actually happened in this project: a shipped decision, a recorded bug, or a retrieval fault found in review. It runs against the live index without spending anything, so a ranking change can be judged the moment it lands instead of waiting for the next graded cycle. This is the isolated companion to BRAIN HELP FOUND above, which measures the same thing from inside real graded work.

ANSWER FOUND COMPLETE79.4%+8.8 ptsTARGET 85%34 curated queries · every required document returned
DOCUMENTS FOUND80.5%+7.3 ptsTARGET 90%counts partial credit where a query wanted more than one document
RANK QUALITY0.606+0.037TARGET 0.75mean reciprocal rank — 1.000 would put the right document first every time
EXACT-TOKEN QUERIES80.0%+40.0 ptsTARGET 90%5 queries for a bare identifier such as STATUS_RANK
WRONG-ANSWER-ON-TOP0TARGET METa document recorded as wrong for a query that outranked the right one
RANKING ALONE75.9%new measureTARGET 85%excludes the 5 queries answered by an always-pinned standing rule

HOW SEARCH HAS IMPROVED

Each row is one setting of the search itself, scored on the same 34 questions. Reading down the rows is the changelog of retrieval getting better.

SEARCH SETTING
the retrieval configuration in that row
ALL DOCS FOUND
questions where every required document came back
DOCS FOUND
the same, with partial credit for finding some of them
RANK QUALITY
how near the top the right document landed, 1.000 is always first
EXACT-TOKEN
queries for a bare identifier such as STATUS_RANK
WRONG ON TOP
a known-wrong document that outranked the right one
RUNSEARCH SETTINGALL DOCS FOUNDDOCS FOUNDRANK QUALITYEXACT-TOKENWRONG ON TOP
2026-07-26 16:13baseline-vector-only70.6%73.2%0.56940.0%0
2026-07-26 16:17lexical+rrf (gold set g010 not yet corrected)73.5%75.6%0.57360.0%0
2026-07-26 16:19lexical+rrf76.5%78.0%0.60280.0%0
2026-07-26 16:20lexical+rrf76.5%78.0%0.60280.0%0
2026-07-26 16:22lexical+rrf76.5%78.0%0.60280.0%0
2026-07-27 03:20standing pins for global.*79.4%80.5%0.60680.0%0

Only shipped settings appear here. A configuration that was measured and rejected never reads as history.

6 ·TECHNICAL RECORDS

WHAT EACH RUN COST AND DID

Spend, waiting times, and every individual exam. Folded away because it answers what happened rather than how is it going.

WORKFLOW HISTORY · 33 CYCLES

One workflow cycle is one complete five-exam batch: engineer attempts → grading → learning review → Brain update when allowed → saved records. Completed means the workflow finished; the answers may still have failed. Times are Japan Standard Time. Unknown calls are engineer / teacher / harvest.

CYCLETIMEMODE / STARTSTATUSTOKENSSPENDWAITSUNKNOWN E/T/HBRAIN GROWTHDETAIL
#332026-07-28 22:28 JSTnightly / manualcompleted1,106,668FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher/harvest
none0/0/0+0 units
0 docs · 0 chunks
5 exams · 2 passed · 0 partial · 3 failed · 0 improvements applied
#322026-07-28 21:56 JSTnightly / manualcompleted1,675,407FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher/harvest
none0/0/0+0 units
0 docs · 0 chunks
5 exams · 1 passed · 1 partial · 3 failed · 0 improvements applied
#312026-07-28 16:51 JSTweekly / manualcompleted4,315,651
5H START: 75 REQUESTS · 1,929,006 TOKENS
FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher/harvest
11634ms, 58680ms, 59257ms0/4/0+0 units
0 docs · 0 chunks
14 exams · 0 passed · 4 partial · 10 failed · 0 improvements applied
#302026-07-27 18:47 JSTnightly / manualcompleted1,471,516FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher/harvest
none0/0/0+0 units
0 docs · 0 chunks
5 exams · 0 passed · 0 partial · 5 failed · 0 improvements applied
#292026-07-25 13:39 JSTnightly / manualcompleted1,969,521
5H START: 248 REQUESTS · 6,808,731 TOKENS
FLAT-RATE ENGINEER
TOKENS ONLY · $2.5492 teacher/harvest
none0/0/0+8 units
8 docs · 8 chunks
5 exams · 1 passed · 0 partial · 4 failed · 8 improvements applied
#282026-07-25 12:45 JSTnightly / manualcompleted1,489,201
5H START: 195 REQUESTS · 5,528,418 TOKENS
FLAT-RATE ENGINEER
TOKENS ONLY · $2.1260 teacher/harvest
none0/0/0+8 units
8 docs · 8 chunks
5 exams · 0 passed · 2 partial · 3 failed · 8 improvements applied
#272026-07-25 11:55 JSTnightly / manualcompleted2,378,990
5H START: 135 REQUESTS · 3,457,884 TOKENS
FLAT-RATE ENGINEER
TOKENS ONLY · $2.7434 teacher/harvest
none0/0/0+6 units
9 docs · 9 chunks
5 exams · 0 passed · 1 partial · 4 failed · 6 improvements applied
#262026-07-25 10:45 JSTnightly / manualcompleted1,917,048
5H START: 72 REQUESTS · 1,840,168 TOKENS
FLAT-RATE ENGINEER
TOKENS ONLY · $2.9183 teacher/harvest
none0/0/0+9 units
8 docs · 8 chunks
5 exams · 1 passed · 1 partial · 3 failed · 9 improvements applied
#252026-07-25 09:10 JSTnightly / manualcompleted2,138,388FLAT-RATE ENGINEER
TOKENS ONLY · $3.0286 teacher/harvest
none1/0/0+10 units
11 docs · 11 chunks
5 exams · 0 passed · 0 partial · 5 failed · 10 improvements applied
#242026-07-25 02:36 JSTnightly / manualfailed1,202,600$0.0000nonefetch failed
#232026-07-25 02:02 JSTnightly / manualfailed314,695$0.0000noneGym shutdown requested: SIGINT
#222026-07-23 01:36 JSTweekly / manualcompleted6,597,005$9.041141709ms, 42948ms1/2/0+8 units
20 docs · 20 chunks
18 exams · 0 passed · 4 partial · 14 failed · 8 improvements applied
EXAM HISTORY · LATEST 30

One row is one individual evaluation exam inside a workflow cycle. Times are Japan Standard Time. Engineer and teacher spend are persisted separately.

EXAMTIMEPROJECT AREAVERDICTTOKENSSPENDBRAIN CALLS
exam-00572026-07-28 22:28 JSTwebfail209,782FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00272026-07-28 22:20 JSTwebpass216,218FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00602026-07-28 22:17 JSTwebfail235,100FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00562026-07-28 22:11 JSTwebfail217,909FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
1
exam-00532026-07-28 22:04 JSTwebpass227,659FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
1
exam-00522026-07-28 21:56 JSTwebpass232,105FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00512026-07-28 21:44 JSTwebpartial420,372FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00502026-07-28 21:33 JSTwebfail408,751FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
1
exam-00492026-07-28 19:33 JSTwebfail211,662FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00482026-07-28 19:03 JSTwebfail402,517FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00632026-07-28 16:36 JSTwebfail422,018FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00442026-07-28 16:32 JSTwebfail425,791FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00332026-07-28 16:17 JSTwebpartial188,595FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00212026-07-28 16:13 JSTwebfail234,277FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00202026-07-28 16:09 JSTwebfail440,316FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
2
exam-00192026-07-28 15:53 JSTwebfail421,830FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00182026-07-28 15:47 JSTiosfail417,188FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00062026-07-28 15:31 JSTiosfail236,536FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00012026-07-28 15:16 JSTiosfail424,512FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00472026-07-28 15:06 JSTwebfail239,860FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00162026-07-28 14:52 JSTiosfail214,196FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00462026-07-28 14:44 JSTwebpartial221,021FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00102026-07-28 14:41 JSTiospartial212,860FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00412026-07-28 14:36 JSTwebpartial216,651FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
1
exam-00452026-07-28 14:26 JSTwebfail407,033FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00052026-07-28 14:21 JSTiospartial369,525FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00432026-07-28 14:13 JSTwebfail400,950FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00042026-07-28 14:02 JSTiosfail440,986FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00422026-07-28 13:13 JSTwebfail310,512FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
0
exam-00402026-07-27 18:47 JSTwebfail226,079FLAT-RATE ENGINEER
TOKENS ONLY · $0.0000 teacher
1