PAID WORK RESTRICTED
EVALUATION BANK READINESS
Before any score means anything, the exam bank itself has to be sound. These are counts of tasks, not results.
TRAINING CYCLES One cycle is one five-exam practice batch: the engineer attempts the exams, the teacher grades them, and accepted lessons are written to the Brain. Read down the cycles to see whether learning is helping.
The loop itself: each cycle gives the student five exams, grades them, and writes back what it learned. The five most recent are here; open one to see what it scored and what it cost.
09nightly · 28 JUL 2026Cycle 33VIEW RESULTS
| MODEL | RUN SETTINGS | OUTCOMES | TASK SIZES | TASK AVERAGE | 5 QUALITY SCORES | USEFUL RESULTS | ENGINEER USAGE | BRAIN GROWTH |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2glm-5.2 | ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget200,000 / exam | 2 PASS0 PARTIAL3 FAIL | SMALL1/2 usefulMEDIUM1/3 usefulLARGE0/0 useful | AVERAGE SCORE3.52 / 51 is weak · 5 is excellent | CORRECT3.20 / 5STYLE3.60 / 5DECISIONS3.80 / 5REASONING3.60 / 5SCOPE3.40 / 5 | 2 / 540.0% partial or pass | 1.1M total tokensThinking used87K · 87.6%Answer generated12KGraded5/5Provider-error tasks0.0% | +0 lessons 0 docs / 0 chunks flat-rate engineer |
08nightly · 28 JUL 2026Cycle 32VIEW RESULTS
| MODEL | RUN SETTINGS | OUTCOMES | TASK SIZES | TASK AVERAGE | 5 QUALITY SCORES | USEFUL RESULTS | ENGINEER USAGE | BRAIN GROWTH |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2glm-5.2 | ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam | 1 PASS1 PARTIAL3 FAIL | SMALL1/1 usefulMEDIUM0/1 usefulLARGE1/3 useful | AVERAGE SCORE3.24 / 51 is weak · 5 is excellent | CORRECT3.00 / 5STYLE3.40 / 5DECISIONS3.60 / 5REASONING3.40 / 5SCOPE2.80 / 5 | 2 / 540.0% partial or pass | 1.7M total tokensThinking used150K · 91.3%Answer generated14KGraded5/5Provider-error tasks0.0% | +0 lessons 0 docs / 0 chunks flat-rate engineer |
07weekly · 28 JUL 2026Cycle 31VIEW RESULTS
| MODEL | RUN SETTINGS | OUTCOMES | TASK SIZES | TASK AVERAGE | 5 QUALITY SCORES | USEFUL RESULTS | ENGINEER USAGE | BRAIN GROWTH |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2glm-5.2 | ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam | 0 PASS4 PARTIAL10 FAIL | SMALL4/8 usefulMEDIUM0/1 usefulLARGE0/5 useful | AVERAGE SCORE2.05 / 51 is weak · 5 is excellent | CORRECT2.00 / 5STYLE2.42 / 5DECISIONS1.83 / 5REASONING2.33 / 5SCOPE1.67 / 5 | 4 / 1428.6% partial or pass | 4.3M total tokensThinking used448K · 94.3%Answer generated27KGraded12/14Provider-error tasks0.0% | +0 lessons 0 docs / 0 chunks flat-rate engineer |
06nightly · 27 JUL 2026Cycle 30VIEW RESULTS
| MODEL | RUN SETTINGS | OUTCOMES | TASK SIZES | TASK AVERAGE | 5 QUALITY SCORES | USEFUL RESULTS | ENGINEER USAGE | BRAIN GROWTH |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2glm-5.2 | ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam | 0 PASS0 PARTIAL5 FAIL | SMALL0/2 usefulMEDIUM0/1 usefulLARGE0/2 useful | AVERAGE SCORE1.52 / 51 is weak · 5 is excellent | CORRECT1.20 / 5STYLE1.80 / 5DECISIONS1.60 / 5REASONING1.80 / 5SCOPE1.20 / 5 | 0 / 50.0% partial or pass | 1.5M total tokensThinking used142K · 92.7%Answer generated11KGraded5/5Provider-error tasks0.0% | +0 lessons 0 docs / 0 chunks flat-rate engineer |
05nightly · 25 JUL 2026Cycle 29VIEW RESULTS
| MODEL | RUN SETTINGS | OUTCOMES | TASK SIZES | TASK AVERAGE | 5 QUALITY SCORES | USEFUL RESULTS | ENGINEER USAGE | BRAIN GROWTH |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2glm-5.2 | ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam | 1 PASS0 PARTIAL4 FAIL | SMALL1/1 usefulMEDIUM0/3 usefulLARGE0/1 useful | AVERAGE SCORE2.76 / 51 is weak · 5 is excellent | CORRECT2.40 / 5STYLE3.20 / 5DECISIONS3.20 / 5REASONING2.60 / 5SCOPE2.40 / 5 | 1 / 520.0% partial or pass | 1.7M total tokensThinking used168K · 92.6%Answer generated13KGraded5/5Provider-error tasks0.0% | +8 lessons 8 docs / 8 chunks flat-rate engineer |
OLDER CYCLES · 4 MORE ON THIS ENGINEER
Same engineer and the same grading, just further back. Kept so a trend can be checked, folded away so the recent result leads.
04nightly · 25 JUL 2026Cycle 28VIEW RESULTS
| MODEL | RUN SETTINGS | OUTCOMES | TASK SIZES | TASK AVERAGE | 5 QUALITY SCORES | USEFUL RESULTS | ENGINEER USAGE | BRAIN GROWTH |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2glm-5.2 | ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam | 0 PASS2 PARTIAL3 FAIL | SMALL1/1 usefulMEDIUM1/4 usefulLARGE0/0 useful | AVERAGE SCORE2.90 / 51 is weak · 5 is excellent | CORRECT3.00 / 5STYLE2.25 / 5DECISIONS3.25 / 5REASONING3.00 / 5SCOPE3.00 / 5 | 2 / 540.0% partial or pass | 1.3M total tokensThinking used149K · 93.5%Answer generated10KGraded4/5Provider-error tasks0.0% | +8 lessons 8 docs / 8 chunks flat-rate engineer |
03nightly · 25 JUL 2026Cycle 27VIEW RESULTS
| MODEL | RUN SETTINGS | OUTCOMES | TASK SIZES | TASK AVERAGE | 5 QUALITY SCORES | USEFUL RESULTS | ENGINEER USAGE | BRAIN GROWTH |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2glm-5.2 | ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam | 0 PASS1 PARTIAL4 FAIL | SMALL0/0 usefulMEDIUM0/0 usefulLARGE1/5 useful | AVERAGE SCORE2.35 / 51 is weak · 5 is excellent | CORRECT1.75 / 5STYLE2.50 / 5DECISIONS3.00 / 5REASONING2.50 / 5SCOPE2.00 / 5 | 1 / 520.0% partial or pass | 2.1M total tokensThinking used245K · 96.1%Answer generated10KGraded4/5Provider-error tasks0.0% | +6 lessons 9 docs / 9 chunks flat-rate engineer |
02nightly · 25 JUL 2026Cycle 26VIEW RESULTS
| MODEL | RUN SETTINGS | OUTCOMES | TASK SIZES | TASK AVERAGE | 5 QUALITY SCORES | USEFUL RESULTS | ENGINEER USAGE | BRAIN GROWTH |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2glm-5.2 | ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam | 1 PASS1 PARTIAL3 FAIL | SMALL1/2 usefulMEDIUM1/2 usefulLARGE0/1 useful | AVERAGE SCORE2.64 / 51 is weak · 5 is excellent | CORRECT2.40 / 5STYLE3.20 / 5DECISIONS2.80 / 5REASONING2.40 / 5SCOPE2.40 / 5 | 2 / 540.0% partial or pass | 1.6M total tokensThinking used142K · 89.3%Answer generated17KGraded5/5Provider-error tasks0.0% | +9 lessons 8 docs / 8 chunks flat-rate engineer |
01nightly · 25 JUL 2026Cycle 25 · baselineVIEW RESULTS
| MODEL | RUN SETTINGS | OUTCOMES | TASK SIZES | TASK AVERAGE | 5 QUALITY SCORES | USEFUL RESULTS | ENGINEER USAGE | BRAIN GROWTH |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2glm-5.2 | BASELINE ReasoningMax effortPrecisionFP8Reply cap32,768 / callTime limit900 s / callToken budget400,000 / exam | 0 PASS0 PARTIAL5 FAIL | SMALL0/2 usefulMEDIUM0/2 usefulLARGE0/1 useful | AVERAGE SCORE1.92 / 51 is weak · 5 is excellent | CORRECT1.00 / 5STYLE2.00 / 5DECISIONS2.80 / 5REASONING2.00 / 5SCOPE1.80 / 5 | 0 / 50.0% partial or pass | 1.7M total tokensThinking used160K · 91.8%Answer generated14KGraded5/5Provider-error tasks0.0% | +10 lessons 11 docs / 11 chunks flat-rate engineer |
EARLIER CYCLES · 12 ON A PREVIOUS ENGINEER MODEL
These ran on DeepSeek V4 Pro and, for the oldest of them, a harness that stopped a solution after its first edit. They are kept for history and are deliberately excluded from the charts above.
| MODEL | RUN SETTINGS | OUTCOMES | TASK SIZES | TASK AVERAGE | 5 QUALITY SCORES | USEFUL RESULTS | ENGINEER USAGE | BRAIN GROWTH |
|---|---|---|---|---|---|---|---|---|
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap | 0 PASS0 PARTIAL5 FAIL | SMALL0/1 usefulMEDIUM0/1 usefulLARGE0/3 useful | AVERAGE SCORE1.80 / 51 is weak · 5 is excellent | CORRECT1.20 / 5STYLE1.80 / 5DECISIONS2.20 / 5REASONING2.00 / 5SCOPE1.80 / 5 | 0 / 50.0% partial or pass | 1.4M total tokensThinking used0 · 0.0%Answer generated0Graded5/5Provider-error tasks0.0% | +9 lessons 0 docs / 0 chunks $1.39 |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap | 0 PASS0 PARTIAL5 FAIL | SMALL0/1 usefulMEDIUM0/0 usefulLARGE0/4 useful | AVERAGE SCORE1.88 / 51 is weak · 5 is excellent | CORRECT1.20 / 5STYLE2.40 / 5DECISIONS1.80 / 5REASONING2.00 / 5SCOPE2.00 / 5 | 0 / 50.0% partial or pass | 1.8M total tokensThinking used0 · 0.0%Answer generated0Graded5/5Provider-error tasks0.0% | +12 lessons 9 docs / 9 chunks $1.25 |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap | 0 PASS0 PARTIAL5 FAIL | SMALL0/2 usefulMEDIUM0/1 usefulLARGE0/2 useful | AVERAGE SCORE2.20 / 51 is weak · 5 is excellent | CORRECT1.75 / 5STYLE2.75 / 5DECISIONS2.25 / 5REASONING2.25 / 5SCOPE2.00 / 5 | 0 / 50.0% partial or pass | 1.6M total tokensThinking used0 · 0.0%Answer generated0Graded4/5Provider-error tasks0.0% | +8 lessons 9 docs / 9 chunks $0.94 |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap | 0 PASS2 PARTIAL16 FAIL | SMALL2/4 usefulMEDIUM0/4 usefulLARGE0/10 useful | AVERAGE SCORE2.25 / 51 is weak · 5 is excellent | CORRECT1.53 / 5STYLE2.59 / 5DECISIONS2.71 / 5REASONING2.24 / 5SCOPE2.18 / 5 | 2 / 1811.1% partial or pass | 5.7M total tokensThinking used0 · 0.0%Answer generated0Graded17/18Provider-error tasks0.0% | +7 lessons 1 docs / 1 chunks $6.81 |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap | 0 PASS1 PARTIAL4 FAIL | SMALL0/1 usefulMEDIUM0/2 usefulLARGE1/2 useful | AVERAGE SCORE2.20 / 51 is weak · 5 is excellent | CORRECT1.75 / 5STYLE2.75 / 5DECISIONS2.50 / 5REASONING2.00 / 5SCOPE2.00 / 5 | 1 / 520.0% partial or pass | 1.7M total tokensThinking used0 · 0.0%Answer generated0Graded4/5Provider-error tasks0.0% | +8 lessons 9 docs / 9 chunks $1.47 |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap | 0 PASS0 PARTIAL5 FAIL | SMALL0/3 usefulMEDIUM0/2 usefulLARGE0/0 useful | AVERAGE SCORE2.20 / 51 is weak · 5 is excellent | CORRECT1.50 / 5STYLE3.00 / 5DECISIONS2.50 / 5REASONING2.00 / 5SCOPE2.00 / 5 | 0 / 50.0% partial or pass | 1.3M total tokensThinking used0 · 0.0%Answer generated0Graded2/5Provider-error tasks0.0% | +5 lessons 9 docs / 9 chunks $0.67 |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap | 0 PASS1 PARTIAL4 FAIL | SMALL1/1 usefulMEDIUM0/2 usefulLARGE0/2 useful | AVERAGE SCORE1.84 / 51 is weak · 5 is excellent | CORRECT1.40 / 5STYLE2.80 / 5DECISIONS2.00 / 5REASONING1.60 / 5SCOPE1.40 / 5 | 1 / 520.0% partial or pass | 1.8M total tokensThinking used0 · 0.0%Answer generated0Graded5/5Provider-error tasks0.0% | +9 lessons 10 docs / 10 chunks $1.93 |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap | 0 PASS1 PARTIAL4 FAIL | SMALL1/2 usefulMEDIUM0/2 usefulLARGE0/1 useful | AVERAGE SCORE2.24 / 51 is weak · 5 is excellent | CORRECT1.80 / 5STYLE2.60 / 5DECISIONS2.40 / 5REASONING2.00 / 5SCOPE2.40 / 5 | 1 / 520.0% partial or pass | 1.1M total tokensThinking used0 · 0.0%Answer generated0Graded5/5Provider-error tasks0.0% | +10 lessons 5 docs / 5 chunks $1.49 |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap | 0 PASS0 PARTIAL5 FAIL | SMALL0/1 usefulMEDIUM0/1 usefulLARGE0/3 useful | AVERAGE SCORE1.72 / 51 is weak · 5 is excellent | CORRECT1.00 / 5STYLE1.60 / 5DECISIONS2.00 / 5REASONING2.00 / 5SCOPE2.00 / 5 | 0 / 50.0% partial or pass | 1.8M total tokensThinking used0 · 0.0%Answer generated0Graded5/5Provider-error tasks0.0% | +11 lessons 11 docs / 11 chunks $2.22 |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap | 0 PASS0 PARTIAL5 FAIL | SMALL0/4 usefulMEDIUM0/1 usefulLARGE0/0 useful | AVERAGE SCORE2.16 / 51 is weak · 5 is excellent | CORRECT1.20 / 5STYLE3.00 / 5DECISIONS2.40 / 5REASONING2.00 / 5SCOPE2.20 / 5 | 0 / 50.0% partial or pass | 1M total tokensThinking used0 · 0.0%Answer generated0Graded5/5Provider-error tasks0.0% | +4 lessons 6 docs / 6 chunks $1.26 |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap | 0 PASS0 PARTIAL5 FAIL | SMALL0/0 usefulMEDIUM0/1 usefulLARGE0/4 useful | AVERAGE SCORE1.52 / 51 is weak · 5 is excellent | CORRECT1.00 / 5STYLE1.40 / 5DECISIONS2.00 / 5REASONING1.60 / 5SCOPE1.60 / 5 | 0 / 50.0% partial or pass | 1.9M total tokensThinking used0 · 0.0%Answer generated0Graded5/5Provider-error tasks40.0% | +7 lessons 7 docs / 7 chunks $1.84 |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | ReasoningProvider defaultPrecisionHost-dependentReply cap12,000 / callTime limit300 s / callToken budgetno explicit cap | 0 PASS4 PARTIAL14 FAIL | SMALL3/5 usefulMEDIUM0/3 usefulLARGE1/10 useful | AVERAGE SCORE2.48 / 51 is weak · 5 is excellent | CORRECT1.88 / 5STYLE2.63 / 5DECISIONS2.94 / 5REASONING2.56 / 5SCOPE2.38 / 5 | 4 / 1822.2% partial or pass | 5.8M total tokensThinking used0 · 0.0%Answer generated0Graded16/18Provider-error tasks16.7% | +8 lessons 20 docs / 20 chunks $8.26 |
HOW THE STUDENT MODEL WAS CHOSEN
One-off trials between candidate models, each sitting the same saved exams and graded the same way. They chose the engine and taught the Brain nothing. Open one to compare what worked, code quality, time and cost.
TEST 10CONFIGURATION TEST · 25 JUL 2026GLM-5.2 native Z.AI integration smoke · one examVIEW RESULTS
CURRENT TEST WORDING VERIFIED
| MODEL | SETTINGS | OUTCOMES | TASK SIZES | TEACHER AVERAGE | 5 QUALITY SCORES | GRADED | ENGINEER USAGE | GRADING + TOTAL COST |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2glm-5.2 | LEADING RESULT Reasoning levelThinking on · Max effort (provider default)PrecisionFP8Output cap33K tokens / replyTime limit900s | 0 PASS1 PARTIAL0 FAIL | SMALL1/1 usefulMEDIUM0/0 usefulLARGE0/0 useful | AVERAGE SCORE3.20 / 51 is weak · 5 is excellent | CORRECT3.00 / 5STYLE3.00 / 5DECISIONS4.00 / 5REASONING3.00 / 5SCOPE3.00 / 5 | 1/1complete | 212K total tokensWork used212KThinking used27K · 92.1%Answer generated2KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time7.6 min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops1Thinking-cap stops0Tool-flow errors1Unknown calls0 | $0.33 grading $0.33 total $0.33 per task 0 unknown teacher calls |
TEST 09CONFIGURATION TEST · 24 JUL 2026GLM-5.2 fixed-harness final validation · Z.AI FP8VIEW RESULTS
CURRENT TEST WORDING VERIFIED
| MODEL | SETTINGS | OUTCOMES | TASK SIZES | TEACHER AVERAGE | 5 QUALITY SCORES | GRADED | ENGINEER USAGE | GRADING + TOTAL COST |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2z-ai/glm-5.2 | OUTCOME LEADER Reasoning levelhigh reasoningPrecisionFP8Output cap33K tokens / replyTime limit900s | 1 PASS0 PARTIAL3 FAIL | SMALL1/4 usefulMEDIUM0/0 usefulLARGE0/0 useful | AVERAGE SCORE3.10 / 51 is weak · 5 is excellent | CORRECT2.50 / 5STYLE4.25 / 5DECISIONS3.00 / 5REASONING2.75 / 5SCOPE3.00 / 5 | 4/4complete | 927K total tokensWork used872KThinking used55K · 88.8%Answer generated7KWork allowance200K / examThinking allowance150K / examAverage time5.0 min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops4Thinking-cap stops0Tool-flow errors10Unknown calls0 | $1.11 grading $1.81 total $0.45 per task 0 unknown teacher calls |
TEST 08CONFIGURATION TEST · 24 JUL 2026GLM-5.2 reasoning and token-cap testVIEW RESULTS
CURRENT TEST WORDING VERIFIED
| MODEL | SETTINGS | OUTCOMES | TASK SIZES | TEACHER AVERAGE | 5 QUALITY SCORES | GRADED | ENGINEER USAGE | GRADING + TOTAL COST |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2z-ai/glm-5.2 | HIGHEST SCORE Reasoning levelhigh reasoningPrecisionFP8Output cap33K tokens / replyTime limit900s | 0 PASS0 PARTIAL2 FAIL | SMALL0/0 usefulMEDIUM0/0 usefulLARGE0/2 useful | AVERAGE SCORE2.40 / 51 is weak · 5 is excellent | CORRECT2.00 / 5STYLE2.50 / 5DECISIONS3.00 / 5REASONING2.50 / 5SCOPE2.00 / 5 | 2/2complete | 920K total tokensWork used831KThinking used88K · 85.3%Answer generated15KWork allowance400K / examThinking allowance150K / examAverage time11.1 min/taskReply-limit stops1Timeout truncations0Provider-error tasks0.0%Work-cap stops2Thinking-cap stops0Tool-flow errors8Unknown calls0 | $1.17 grading $1.99 total $0.99 per task 0 unknown teacher calls |
TEST 07CONFIGURATION TEST · 24 JUL 2026GLM-5.2 quality probe v3 · Z.AI FP8VIEW RESULTS
CURRENT TEST WORDING VERIFIED
| MODEL | SETTINGS | OUTCOMES | TASK SIZES | TEACHER AVERAGE | 5 QUALITY SCORES | GRADED | ENGINEER USAGE | GRADING + TOTAL COST |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2z-ai/glm-5.2 | LEADING RESULT Reasoning levelhigh reasoningPrecisionFP8Output cap33K tokens / replyTime limit900s | 0 PASS1 PARTIAL2 FAIL | SMALL1/1 usefulMEDIUM0/0 usefulLARGE0/2 useful | AVERAGE SCORE2.13 / 51 is weak · 5 is excellent | CORRECT2.00 / 5STYLE2.33 / 5DECISIONS2.33 / 5REASONING2.00 / 5SCOPE2.00 / 5 | 3/3complete | 1.1M total tokensWork used1MThinking used111K · 85.4%Answer generated19KWork allowance200K / 400K / examThinking allowance150K / examAverage time9.5 min/taskReply-limit stops1Timeout truncations0Provider-error tasks0.0%Work-cap stops2Thinking-cap stops0Tool-flow errors11Unknown calls0 | $1.17 grading $2.18 total $0.73 per task 0 unknown teacher calls |
TEST 06CONFIGURATION TEST · 24 JUL 2026GLM-5.2 quality probe · Z.AI FP8VIEW RESULTS
CURRENT TEST WORDING VERIFIED
| MODEL | SETTINGS | OUTCOMES | TASK SIZES | TEACHER AVERAGE | 5 QUALITY SCORES | GRADED | ENGINEER USAGE | GRADING + TOTAL COST |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2z-ai/glm-5.2 | HIGHEST SCORE Reasoning levelhigh reasoningPrecisionFP8Output cap33K tokens / replyTime limit900s | 0 PASS0 PARTIAL3 FAIL | SMALL0/1 usefulMEDIUM0/0 usefulLARGE0/2 useful | AVERAGE SCORE2.00 / 51 is weak · 5 is excellent | CORRECT1.00 / 5STYLE1.67 / 5DECISIONS2.67 / 5REASONING2.00 / 5SCOPE2.67 / 5 | 3/3complete | 916K total tokensWork used817KThinking used98K · 86.2%Answer generated16KWork allowance200K / 400K / examThinking allowance150K / examAverage time8.5 min/taskReply-limit stops1Timeout truncations0Provider-error tasks0.0%Work-cap stops0Thinking-cap stops0Tool-flow errors17Unknown calls0 | $0.98 grading $1.86 total $0.62 per task 0 unknown teacher calls |
OLDER COMPARISONS · 5 MORE ON RECORD
Earlier model and setting trials, same exams and same grading. Kept for the record, folded away because the choice they informed is already made.
TEST 05FAIR-BUDGET HEAD-TO-HEAD · 24 JUL 2026GLM and Kimi · equal work and thinking limitsVIEW RESULTS
| MODEL | SETTINGS | OUTCOMES | TASK SIZES | TEACHER AVERAGE | 5 QUALITY SCORES | GRADED | ENGINEER USAGE | GRADING + TOTAL COST |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2z-ai/glm-5.2 | PARTIAL DATA Reasoning levelThinking on · Max effort (provider default)PrecisionHost-dependentOutput cap66K tokens / replyTime limit900s | 0 PASS3 PARTIAL9 FAIL | SMALL2/7 usefulMEDIUM0/2 usefulLARGE1/3 useful | AVERAGE SCORE2.28 / 51 is weak · 5 is excellent | CORRECT1.90 / 5STYLE2.90 / 5DECISIONS2.40 / 5REASONING2.20 / 5SCOPE2.00 / 5 | 10/122 ungraded | 3.3M total tokensWork used2.7MThinking used507K · 96.3%Answer generated20KWork allowance200K / 400K / examThinking allowance150K / examAverage time17.0 min/taskReply-limit stops5Timeout truncations0Provider-error tasks41.7%Work-cap stops5Thinking-cap stops2Tool-flow errors36Unknown calls7 | $2.80 grading $5.01 total $0.42 per task 0 unknown teacher calls |
GLM-5.2 GLM-5.2z-ai/glm-5.2 | PARTIAL DATA Reasoning levelxhigh reasoningPrecisionHost-dependentOutput cap66K tokens / replyTime limit900s | 0 PASS0 PARTIAL4 FAIL | SMALL0/2 usefulMEDIUM0/2 usefulLARGE0/0 useful | AVERAGE SCORE1.85 / 51 is weak · 5 is excellent | CORRECT1.25 / 5STYLE2.25 / 5DECISIONS2.25 / 5REASONING1.75 / 5SCOPE1.75 / 5 | 4/4complete | 876K total tokensWork used784KThinking used92K · 91.6%Answer generated8KWork allowance200K / examThinking allowance150K / examAverage time13.1 min/taskReply-limit stops0Timeout truncations0Provider-error tasks50.0%Work-cap stops2Thinking-cap stops0Tool-flow errors13Unknown calls2 | $1.07 grading $1.59 total $0.40 per task 0 unknown teacher calls |
TEST 04DEFAULT REASONING BASELINE · 24 JUL 2026Provider-default baseline · deterministically regradedVIEW RESULTS
CURRENT TEST WORDING VERIFIED
| MODEL | SETTINGS | OUTCOMES | TASK SIZES | TEACHER AVERAGE | 5 QUALITY SCORES | GRADED | ENGINEER USAGE | GRADING + TOTAL COST |
|---|---|---|---|---|---|---|---|---|
GLM-5.2 GLM-5.2z-ai/glm-5.2 | LEADING RESULT Reasoning levelprovider-default reasoningPrecisionHost-dependentOutput cap32K tokens / replyTime limit300s | 0 PASS2 PARTIAL7 FAIL | SMALL1/4 usefulMEDIUM0/2 usefulLARGE1/3 useful | AVERAGE SCORE2.56 / 51 is weak · 5 is excellent | CORRECT1.89 / 5STYLE2.33 / 5DECISIONS3.11 / 5REASONING2.56 / 5SCOPE2.89 / 5 | 9/9complete | 2M total tokensWork used2MThinking used101K · 82.0%Answer generated22KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time5.4 min/taskReply-limit stops0Timeout truncations0Provider-error tasks22.2%Work-cap stops0Thinking-cap stops0Tool-flow errors19Unknown calls2 | $2.78 grading $3.63 total $0.40 per task 0 unknown teacher calls |
Kimi K2.7 Code Kimi K2.7 Codemoonshotai/kimi-k2.7-code | RANK 2 Reasoning levelprovider-default reasoningPrecisionHost-dependentOutput cap32K tokens / replyTime limit300s | 0 PASS2 PARTIAL7 FAIL | SMALL2/4 usefulMEDIUM0/2 usefulLARGE0/3 useful | AVERAGE SCORE2.13 / 51 is weak · 5 is excellent | CORRECT1.89 / 5STYLE2.44 / 5DECISIONS2.33 / 5REASONING2.11 / 5SCOPE1.89 / 5 | 9/9complete | 2.3M total tokensWork used2.3MThinking used32K · 64.6%Answer generated18KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time1.7 min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops0Thinking-cap stops0Tool-flow errors16Unknown calls0 | $2.60 grading $3.37 total $0.37 per task 0 unknown teacher calls |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | RANK 3 Reasoning levelprovider-default reasoningPrecisionHost-dependentOutput cap12K tokens / replyTime limit300s | 0 PASS1 PARTIAL8 FAIL | SMALL1/4 usefulMEDIUM0/2 usefulLARGE0/3 useful | AVERAGE SCORE2.05 / 51 is weak · 5 is excellent | CORRECT1.50 / 5STYLE2.38 / 5DECISIONS2.50 / 5REASONING1.88 / 5SCOPE2.00 / 5 | 8/91 ungraded | 2.4M total tokensWork used2.4MThinking used27K · 45.4%Answer generated32KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time0.6 min/taskReply-limit stops0Timeout truncations0Provider-error tasks22.2%Work-cap stops1Thinking-cap stops0Tool-flow errors4Unknown calls2 | $2.30 grading $2.90 total $0.32 per task 0 unknown teacher calls |
TEST 03MAXIMUM REASONING TEST · 24 JUL 2026Maximum-reasoning model comparisonVIEW RESULTS
CURRENT TEST WORDING VERIFIED
TWO DIFFERENT LEADERS — Kimi is the outcome leader because this ranking counts full passes first, and Kimi earned the only pass. GLM is the average-score leader (2.26 vs 2.00). One pass in nine tasks does not prove Kimi is generally better.
| MODEL | SETTINGS | OUTCOMES | TASK SIZES | TEACHER AVERAGE | 5 QUALITY SCORES | GRADED | ENGINEER USAGE | GRADING + TOTAL COST |
|---|---|---|---|---|---|---|---|---|
Kimi K2.7 Code Kimi K2.7 Codemoonshotai/kimi-k2.7-code | OUTCOME LEADER Reasoning levelxhigh reasoningPrecisionHost-dependentOutput cap66K tokens / replyTime limit900s | 1 PASS1 PARTIAL7 FAIL | SMALL2/4 usefulMEDIUM0/2 usefulLARGE0/3 useful | AVERAGE SCORE2.00 / 51 is weak · 5 is excellent | CORRECT1.78 / 5STYLE2.11 / 5DECISIONS2.22 / 5REASONING2.00 / 5SCOPE1.89 / 5 | 9/9complete | 2.4M total tokensWork used2.4MThinking used32K · 66.3%Answer generated17KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time1.6 min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops0Thinking-cap stops0Tool-flow errors18Unknown calls0 | $2.31 grading $3.16 total $0.35 per task 0 unknown teacher calls |
GLM-5.2 GLM-5.2z-ai/glm-5.2 | RANK 2 Reasoning levelxhigh reasoningPrecisionHost-dependentOutput cap32K / 66K tokens / replyTime limit300 / 900s | 0 PASS2 PARTIAL7 FAIL | SMALL1/4 usefulMEDIUM0/2 usefulLARGE1/3 useful | AVERAGE SCORE2.26 / 51 is weak · 5 is excellent | CORRECT1.89 / 5STYLE2.22 / 5DECISIONS2.67 / 5REASONING2.22 / 5SCOPE2.33 / 5 | 9/9complete | 2.6M total tokensWork used2.4MThinking used328K · 90.7%Answer generated34KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time11.8 min/taskReply-limit stops1Timeout truncations2Provider-error tasks11.1%Work-cap stops0Thinking-cap stops0Tool-flow errors39Unknown calls1 | $2.79 grading $4.50 total $0.50 per task 0 unknown teacher calls $0.08 aborted/no verdict · 140K tokens |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | RANK 3 Reasoning levelxhigh reasoningPrecisionHost-dependentOutput cap32K tokens / replyTime limit300s | 0 PASS2 PARTIAL7 FAIL | SMALL1/4 usefulMEDIUM0/2 usefulLARGE1/3 useful | AVERAGE SCORE2.06 / 51 is weak · 5 is excellent | CORRECT1.83 / 5STYLE2.67 / 5DECISIONS2.00 / 5REASONING2.00 / 5SCOPE1.83 / 5 | 6/93 ungraded | 2.5M total tokensWork used2.5MThinking used42K · 59.6%Answer generated29KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time2.6 min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops3Thinking-cap stops0Tool-flow errors24Unknown calls0 | $1.80 grading $2.45 total $0.27 per task 0 unknown teacher calls |
TEST 02MODEL COMPARISON · 23 JUL 2026Coding model capability comparisonVIEW RESULTS
AUDIT HOLD — These saved results used a grading rule that treated equivalent work differently. They remain visible for history but are not ranked until deterministic regrading.
| MODEL | SETTINGS | OUTCOMES | TASK SIZES | TEACHER AVERAGE | 5 QUALITY SCORES | GRADED | ENGINEER USAGE | GRADING + TOTAL COST |
|---|---|---|---|---|---|---|---|---|
Kimi K2.7 Code Kimi K2.7 Codemoonshotai/kimi-k2.7-code | HISTORICAL LABEL · PENDING REGRADE Reasoning levelProvider defaultPrecisionHost-dependentOutput cap32K tokens / replyTime limit300s | 1 PASS1 PARTIAL7 FAIL | SMALL2/4 usefulMEDIUM0/2 usefulLARGE0/3 useful | AVERAGE SCORE2.17 / 51 is weak · 5 is excellent | CORRECT1.89 / 5STYLE2.22 / 5DECISIONS2.33 / 5REASONING2.11 / 5SCOPE2.33 / 5 | 9/9complete | 2.3M total tokensWork used2.3MThinking used32K · 64.6%Answer generated18KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time1.7 min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops0Thinking-cap stops0Tool-flow errors16Unknown calls0 | $2.81 grading $3.57 total $0.40 per task 0 unknown teacher calls |
GLM-5.2 GLM-5.2z-ai/glm-5.2 | HISTORICAL LABEL · PENDING REGRADE Reasoning levelThinking on · Max effort (provider default)PrecisionHost-dependentOutput cap32K tokens / replyTime limit300s | 0 PASS2 PARTIAL7 FAIL | SMALL2/4 usefulMEDIUM0/2 usefulLARGE0/3 useful | AVERAGE SCORE2.71 / 51 is weak · 5 is excellent | CORRECT1.78 / 5STYLE2.44 / 5DECISIONS3.33 / 5REASONING2.67 / 5SCOPE3.33 / 5 | 9/9complete | 2M total tokensWork used2MThinking used101K · 82.0%Answer generated22KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time5.4 min/taskReply-limit stops0Timeout truncations0Provider-error tasks22.2%Work-cap stops0Thinking-cap stops0Tool-flow errors19Unknown calls2 | $2.87 grading $3.71 total $0.41 per task 0 unknown teacher calls |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | HISTORICAL LABEL · PENDING REGRADE Reasoning levelProvider defaultPrecisionHost-dependentOutput cap12K tokens / replyTime limit300s | 0 PASS2 PARTIAL7 FAIL | SMALL1/4 usefulMEDIUM0/2 usefulLARGE1/3 useful | AVERAGE SCORE2.08 / 51 is weak · 5 is excellent | CORRECT1.63 / 5STYLE2.38 / 5DECISIONS2.38 / 5REASONING2.13 / 5SCOPE1.88 / 5 | 8/91 ungraded | 2.4M total tokensWork used2.4MThinking used27K · 45.4%Answer generated32KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time0.6 min/taskReply-limit stops0Timeout truncations0Provider-error tasks22.2%Work-cap stops1Thinking-cap stops0Tool-flow errors22Unknown calls2 | $2.20 grading $2.80 total $0.31 per task 0 unknown teacher calls |
GLM-5.2 GLM-5.2z-ai/glm-5.2 | CONTEXT ONLY · NOT CAPABILITY RANKED Reasoning levelThinking on · Max effort (provider default)PrecisionHost-dependentOutput cap12K tokens / replyTime limit300s | 0 PASS0 PARTIAL5 FAIL | SMALL0/2 usefulMEDIUM0/2 usefulLARGE0/1 useful | AVERAGE SCORE1.40 / 51 is weak · 5 is excellent | CORRECT1.50 / 5STYLE2.00 / 5DECISIONS1.00 / 5REASONING1.50 / 5SCOPE1.00 / 5 | 2/53 ungraded | 1.3M total tokensWork used1.3MThinking used132K · 99.6%Answer generated2KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time— min/taskReply-limit stops11Timeout truncations0Provider-error tasks0.0%Work-cap stops3Thinking-cap stops0Tool-flow errors15Unknown calls0 | $0.46 grading $1.15 total $0.23 per task 0 unknown teacher calls $0.90 aborted/no verdict · 362K tokens |
TEST 01CONFIGURATION TEST · 23 JUL 2026DeepSeek V4 Pro reasoning and token-cap testVIEW RESULTS
OLDER EXPERIMENT — The exact task wording was not saved with these results. Use them to understand this older test, not to rank new models against it.
| MODEL | SETTINGS | OUTCOMES | TASK SIZES | TEACHER AVERAGE | 5 QUALITY SCORES | GRADED | ENGINEER USAGE | GRADING + TOTAL COST |
|---|---|---|---|---|---|---|---|---|
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | HIGHEST SCORE Reasoning levelProvider defaultPrecisionHost-dependentOutput cap12K tokens / replyTime limit300s | 0 PASS0 PARTIAL5 FAIL | SMALL0/2 usefulMEDIUM0/2 usefulLARGE0/1 useful | AVERAGE SCORE2.30 / 51 is weak · 5 is excellent | CORRECT1.25 / 5STYLE2.25 / 5DECISIONS2.75 / 5REASONING2.00 / 5SCOPE3.25 / 5 | 4/51 ungraded | 2.2M total tokensWork used2.2MThinking used10K · 38.3%Answer generated17KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time— min/taskReply-limit stops0Timeout truncations0Provider-error tasks20.0%Work-cap stops1Thinking-cap stops0Tool-flow errors0Unknown calls0 | $1.17 grading $1.47 total $0.29 per task 1 unknown teacher calls |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | RANK 2 Reasoning levelhigh reasoningPrecisionHost-dependentOutput cap12K tokens / replyTime limit300s | 0 PASS0 PARTIAL5 FAIL | SMALL0/2 usefulMEDIUM0/2 usefulLARGE0/1 useful | AVERAGE SCORE2.08 / 51 is weak · 5 is excellent | CORRECT1.20 / 5STYLE2.00 / 5DECISIONS2.60 / 5REASONING2.20 / 5SCOPE2.40 / 5 | 5/5complete | 2.1M total tokensWork used2.1MThinking used12K · 33.4%Answer generated24KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time— min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops0Thinking-cap stops0Tool-flow errors0Unknown calls0 | $1.47 grading $1.77 total $0.35 per task 0 unknown teacher calls |
DeepSeek V4 Pro DeepSeek V4 Prodeepseek/deepseek-v4-pro | RANK 3 Reasoning levelhigh reasoningPrecisionHost-dependentOutput cap12K tokens / replyTime limit300s | 0 PASS0 PARTIAL5 FAIL | SMALL0/2 usefulMEDIUM0/2 usefulLARGE0/1 useful | AVERAGE SCORE1.40 / 51 is weak · 5 is excellent | CORRECT1.00 / 5STYLE2.00 / 5DECISIONS1.50 / 5REASONING1.50 / 5SCOPE1.00 / 5 | 4/51 ungraded | 1.3M total tokensWork used1.3MThinking used7K · 26.9%Answer generated19KWork allowanceby task size · 200K–400K / examThinking allowanceshared with work budget / examAverage time— min/taskReply-limit stops0Timeout truncations0Provider-error tasks0.0%Work-cap stops1Thinking-cap stops0Tool-flow errors0Unknown calls0 | $1.09 grading $1.32 total $0.26 per task 0 unknown teacher calls |
HOW IS THE STUDENT DOING?
How the student is doing right now. Only results whose exact task wording and reference answer were saved are counted, so two versions of an exam are never compared as if they were the same.
TARGET NOT MET
TARGET MET
TARGET NOT MET
NO COMPARABLE CURRENT RESULT
TARGET MET
CAN THE BRAIN FIND ITS OWN ANSWERS? A fixed list of questions whose right answers are already known. Retrieval runs each one against the live index and is marked on whether the right documents came back, and how high. It costs nothing to run, so a ranking change can be judged immediately.
Everything above depends on search working. If the right guide never reaches the engineer, no amount of grading will fix the answer. This section measures search on its own, against known-correct answers, so a ranking change can be judged the moment it lands instead of waiting for the next paid cycle.
A fixed list of 34 questions with known right answers, every one taken from something that actually happened in this project: a shipped decision, a recorded bug, or a retrieval fault found in review. It runs against the live index without spending anything, so a ranking change can be judged the moment it lands instead of waiting for the next graded cycle. This is the isolated companion to BRAIN HELP FOUND above, which measures the same thing from inside real graded work.
HOW SEARCH HAS IMPROVED
Each row is one setting of the search itself, scored on the same 34 questions. Reading down the rows is the changelog of retrieval getting better.
- SEARCH SETTING
- the retrieval configuration in that row
- ALL DOCS FOUND
- questions where every required document came back
- DOCS FOUND
- the same, with partial credit for finding some of them
- RANK QUALITY
- how near the top the right document landed, 1.000 is always first
- EXACT-TOKEN
- queries for a bare identifier such as STATUS_RANK
- WRONG ON TOP
- a known-wrong document that outranked the right one
| RUN | SEARCH SETTING | ALL DOCS FOUND | DOCS FOUND | RANK QUALITY | EXACT-TOKEN | WRONG ON TOP |
|---|---|---|---|---|---|---|
| 2026-07-26 16:13 | baseline-vector-only | 70.6% | 73.2% | 0.569 | 40.0% | 0 |
| 2026-07-26 16:17 | lexical+rrf (gold set g010 not yet corrected) | 73.5% | 75.6% | 0.573 | 60.0% | 0 |
| 2026-07-26 16:19 | lexical+rrf | 76.5% | 78.0% | 0.602 | 80.0% | 0 |
| 2026-07-26 16:20 | lexical+rrf | 76.5% | 78.0% | 0.602 | 80.0% | 0 |
| 2026-07-26 16:22 | lexical+rrf | 76.5% | 78.0% | 0.602 | 80.0% | 0 |
| 2026-07-27 03:20 | standing pins for global.* | 79.4% | 80.5% | 0.606 | 80.0% | 0 |
Only shipped settings appear here. A configuration that was measured and rejected never reads as history.
WHAT EACH RUN COST AND DID
Spend, waiting times, and every individual exam. Folded away because it answers what happened rather than how is it going.
WORKFLOW HISTORY · 33 CYCLES
One workflow cycle is one complete five-exam batch: engineer attempts → grading → learning review → Brain update when allowed → saved records. Completed means the workflow finished; the answers may still have failed. Times are Japan Standard Time. Unknown calls are engineer / teacher / harvest.
| CYCLE | TIME | MODE / START | STATUS | TOKENS | SPEND | WAITS | UNKNOWN E/T/H | BRAIN GROWTH | DETAIL |
|---|---|---|---|---|---|---|---|---|---|
| #33 | 2026-07-28 22:28 JST | nightly / manual | completed | 1,106,668 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher/harvest | none | 0/0/0 | +0 units 0 docs · 0 chunks | 5 exams · 2 passed · 0 partial · 3 failed · 0 improvements applied |
| #32 | 2026-07-28 21:56 JST | nightly / manual | completed | 1,675,407 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher/harvest | none | 0/0/0 | +0 units 0 docs · 0 chunks | 5 exams · 1 passed · 1 partial · 3 failed · 0 improvements applied |
| #31 | 2026-07-28 16:51 JST | weekly / manual | completed | 4,315,651 5H START: 75 REQUESTS · 1,929,006 TOKENS | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher/harvest | 11634ms, 58680ms, 59257ms | 0/4/0 | +0 units 0 docs · 0 chunks | 14 exams · 0 passed · 4 partial · 10 failed · 0 improvements applied |
| #30 | 2026-07-27 18:47 JST | nightly / manual | completed | 1,471,516 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher/harvest | none | 0/0/0 | +0 units 0 docs · 0 chunks | 5 exams · 0 passed · 0 partial · 5 failed · 0 improvements applied |
| #29 | 2026-07-25 13:39 JST | nightly / manual | completed | 1,969,521 5H START: 248 REQUESTS · 6,808,731 TOKENS | FLAT-RATE ENGINEER TOKENS ONLY · $2.5492 teacher/harvest | none | 0/0/0 | +8 units 8 docs · 8 chunks | 5 exams · 1 passed · 0 partial · 4 failed · 8 improvements applied |
| #28 | 2026-07-25 12:45 JST | nightly / manual | completed | 1,489,201 5H START: 195 REQUESTS · 5,528,418 TOKENS | FLAT-RATE ENGINEER TOKENS ONLY · $2.1260 teacher/harvest | none | 0/0/0 | +8 units 8 docs · 8 chunks | 5 exams · 0 passed · 2 partial · 3 failed · 8 improvements applied |
| #27 | 2026-07-25 11:55 JST | nightly / manual | completed | 2,378,990 5H START: 135 REQUESTS · 3,457,884 TOKENS | FLAT-RATE ENGINEER TOKENS ONLY · $2.7434 teacher/harvest | none | 0/0/0 | +6 units 9 docs · 9 chunks | 5 exams · 0 passed · 1 partial · 4 failed · 6 improvements applied |
| #26 | 2026-07-25 10:45 JST | nightly / manual | completed | 1,917,048 5H START: 72 REQUESTS · 1,840,168 TOKENS | FLAT-RATE ENGINEER TOKENS ONLY · $2.9183 teacher/harvest | none | 0/0/0 | +9 units 8 docs · 8 chunks | 5 exams · 1 passed · 1 partial · 3 failed · 9 improvements applied |
| #25 | 2026-07-25 09:10 JST | nightly / manual | completed | 2,138,388 | FLAT-RATE ENGINEER TOKENS ONLY · $3.0286 teacher/harvest | none | 1/0/0 | +10 units 11 docs · 11 chunks | 5 exams · 0 passed · 0 partial · 5 failed · 10 improvements applied |
| #24 | 2026-07-25 02:36 JST | nightly / manual | failed | 1,202,600 | $0.0000 | none | — | — | fetch failed |
| #23 | 2026-07-25 02:02 JST | nightly / manual | failed | 314,695 | $0.0000 | none | — | — | Gym shutdown requested: SIGINT |
| #22 | 2026-07-23 01:36 JST | weekly / manual | completed | 6,597,005 | $9.0411 | 41709ms, 42948ms | 1/2/0 | +8 units 20 docs · 20 chunks | 18 exams · 0 passed · 4 partial · 14 failed · 8 improvements applied |
EXAM HISTORY · LATEST 30
One row is one individual evaluation exam inside a workflow cycle. Times are Japan Standard Time. Engineer and teacher spend are persisted separately.
| EXAM | TIME | PROJECT AREA | VERDICT | TOKENS | SPEND | BRAIN CALLS |
|---|---|---|---|---|---|---|
| exam-0057 | 2026-07-28 22:28 JST | web | fail | 209,782 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0027 | 2026-07-28 22:20 JST | web | pass | 216,218 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0060 | 2026-07-28 22:17 JST | web | fail | 235,100 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0056 | 2026-07-28 22:11 JST | web | fail | 217,909 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 1 |
| exam-0053 | 2026-07-28 22:04 JST | web | pass | 227,659 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 1 |
| exam-0052 | 2026-07-28 21:56 JST | web | pass | 232,105 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0051 | 2026-07-28 21:44 JST | web | partial | 420,372 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0050 | 2026-07-28 21:33 JST | web | fail | 408,751 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 1 |
| exam-0049 | 2026-07-28 19:33 JST | web | fail | 211,662 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0048 | 2026-07-28 19:03 JST | web | fail | 402,517 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0063 | 2026-07-28 16:36 JST | web | fail | 422,018 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0044 | 2026-07-28 16:32 JST | web | fail | 425,791 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0033 | 2026-07-28 16:17 JST | web | partial | 188,595 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0021 | 2026-07-28 16:13 JST | web | fail | 234,277 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0020 | 2026-07-28 16:09 JST | web | fail | 440,316 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 2 |
| exam-0019 | 2026-07-28 15:53 JST | web | fail | 421,830 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0018 | 2026-07-28 15:47 JST | ios | fail | 417,188 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0006 | 2026-07-28 15:31 JST | ios | fail | 236,536 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0001 | 2026-07-28 15:16 JST | ios | fail | 424,512 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0047 | 2026-07-28 15:06 JST | web | fail | 239,860 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0016 | 2026-07-28 14:52 JST | ios | fail | 214,196 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0046 | 2026-07-28 14:44 JST | web | partial | 221,021 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0010 | 2026-07-28 14:41 JST | ios | partial | 212,860 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0041 | 2026-07-28 14:36 JST | web | partial | 216,651 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 1 |
| exam-0045 | 2026-07-28 14:26 JST | web | fail | 407,033 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0005 | 2026-07-28 14:21 JST | ios | partial | 369,525 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0043 | 2026-07-28 14:13 JST | web | fail | 400,950 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0004 | 2026-07-28 14:02 JST | ios | fail | 440,986 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0042 | 2026-07-28 13:13 JST | web | fail | 310,512 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 0 |
| exam-0040 | 2026-07-27 18:47 JST | web | fail | 226,079 | FLAT-RATE ENGINEER TOKENS ONLY · $0.0000 teacher | 1 |