# Round 1 of Layatown's self-improvement study: self-consistency distillation (results) Pre-registered in `layatown/docs/SELF-IMPROVE.md` (commit 4f795f2, amendments A1-A3 written before training and evaluation). `self-1` = Layatown v1 fine-tuned on v1's own answers averaged over 4 equivalent views of each state. No Jev label was used for training and no Jev request was made. **Verdict: trade-off.** S1 pass, S2 FAIL, S3 pass. ## Data and cost - Labelled questions: 134,371 from 104,516 states (100,000 decisions; the rest character-sheet and juror anchors), K = 4 views, 537,484 forward passes of v1. - train/character_sheet: 34,120 questions, mean TVD between views 0.0156, top choice differs across views 0.044, teacher vs logged view TVD 0.0100 - train/juror_vote: 251 questions, mean TVD between views 0.0452, top choice differs across views 0.414, teacher vs logged view TVD 0.0471 - train/villager_decision: 50,000 questions, mean TVD between views 0.0389, top choice differs across views 0.085, teacher vs logged view TVD 0.0306 - v1-run/villager_decision: 50,000 questions, mean TVD between views 0.0444, top choice differs across views 0.105, teacher vs logged view TVD 0.0349 - GPU time: labelling 57.76 min of forward passes (58.83 min wall); training 75.1 min wall (6225 steps; about 29.5 min of it stopped by a guard bug, so about 45.6 GPU-min), temperature refit 0.8 min; peak training memory 4.05 GiB (capped share of the shared 12 GB card). - Frozen: embeddings + encoder layers 0-19 of 28; lr encoder 1e-05, head 3e-05, 6144 tokens per step, one epoch. Temperatures refitted on val: [1.0218, 1.0, 1.0171] (v1: 1.0172, 1.0, 1.0014). - Inference cost: unchanged (one view per question, same architecture). ## Pre-registered criteria (self-1 against v1) | criterion | measure | v1 | self-1 | self-1 - v1 [95% paired CI] | margin | result | |---|---|---|---|---|---|---| | S1 | D2 option orders: mean TVD across 4 orders (100 items) | 0.0155 [0.0142, 0.0168] | 0.0147 [0.0135, 0.0159] | -0.0008 [-0.0013, -0.0004] | CI must exclude 0, below | pass | | S1 | D2 option orders: share whose top choice changes | 0.010 | 0.000 | -0.010 [-0.030, 0.000] | not higher | pass | | S2 | held-out decisions: top-1 agreement with Jev, % (n = 238,779) | 86.92 | 86.91 | -0.01 [-0.06, 0.04] | >= -1.0 pt | pass | | S2 | held-out decisions: TVD to Jev | 0.1091 | 0.1093 | 0.0002 [0.0001, 0.0003] | <= +0.01 | pass | | S2 | held-out character sheets: MAE (n = 41,024) | 0.0331 | 0.0332 | 0.0001 [0.0000, 0.0002] | <= +0.01 | pass | | S2 | held-out juror votes: MAE (n = 168) | 0.0615 | 0.0737 | 0.0122 [0.0094, 0.0148] | <= +0.01 | FAIL | | S3 | D1 written-rule compliance, 400 pairs | 0.968 [0.949, 0.983] | 0.968 [0.949, 0.983] | 0.000 [0.000, 0.000] | >= -0.02 | pass | ## Reported, not required (S4 and context) | measure | v1 | self-1 | self-1 - v1 [95% paired CI] | |---|---|---|---| | D2 state-key orders: mean TVD | 0.0528 [0.0474, 0.0586] | 0.0471 [0.0423, 0.0523] | -0.0057 [-0.0069, -0.0046] | | D2 state-key orders: top choice changes | 0.040 [0.010, 0.080] | 0.020 [0.000, 0.050] | -0.020 [-0.050, 0.000] | | D1 mean signed effect | 0.361 [0.335, 0.387] | 0.359 [0.334, 0.384] | -0.002 [-0.003, -0.001] | | D1 compliance: eat_starving (50 pairs) | 1.000 [1.000, 1.000] | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] | | D1 compliance: sleep_exhausted (50 pairs) | 1.000 [1.000, 1.000] | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] | | D1 compliance: danger_here (50 pairs) | 1.000 [1.000, 1.000] | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] | | D1 compliance: cowardly_vs_brave (50 pairs) | 0.980 [0.940, 1.000] | 0.980 [0.940, 1.000] | 0.000 [0.000, 0.000] | | D1 compliance: guard_fight (50 pairs) | 0.840 [0.740, 0.940] | 0.840 [0.740, 0.940] | 0.000 [0.000, 0.000] | | D1 compliance: child_flee (50 pairs) | 1.000 [1.000, 1.000] | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] | | D1 compliance: wounded_flee (50 pairs) | 1.000 [1.000, 1.000] | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] | | D1 compliance: work_healthy (50 pairs) | 0.920 [0.840, 0.980] | 0.920 [0.840, 0.980] | 0.000 [0.000, 0.000] | | D3 chat_moderation (38 rows) | 0.737 [0.605, 0.868] | 0.737 [0.605, 0.868] | 0.000 [0.000, 0.000] | | D3 chat_neighbour_effect (20 rows) | 0.000 [0.000, 0.000] | 0.000 [0.000, 0.000] | 0.000 [0.000, 0.000] | | D3 username_screening (36 rows) | 0.778 [0.639, 0.917] | 0.778 [0.639, 0.917] | 0.000 [0.000, 0.000] | | D3 event_reading (148 rows) | 0.426 [0.345, 0.507] | 0.439 [0.358, 0.520] | 0.013 [-0.013, 0.041] | | D3 juror_factorial (368 rows) | 0.677 [0.628, 0.723] | 0.679 [0.630, 0.726] | 0.003 [-0.014, 0.019] | | held-out decisions: Brier to Jev | 0.0223 | 0.0223 | -0.0000 [-0.0001, 0.0000] | | held-out sheets: side agreement | 0.9412 | 0.9410 | -0.0001 [-0.0005, 0.0001] | How far self-1 moved from v1 on the held-out day: mean TVD between them 0.0084, same top choice 99.00%. On the level D probes: {'answers': 1500, 'meanTvdToV1': 0.0094, 'topSameAsV1': 0.9913}. Check: v1 re-served today gives D2 option-order TVD 0.0155 [0.0142, 0.0168] (committed answers: 0.0155 [0.0142, 0.0168]). ### Character effects on the held-out day (study criterion A5 contrasts) | contrast | Jev | v1 | self-1 | v1 pass | self-1 pass | |---|---|---|---|---|---| | cowardly flee | +0.431 | +0.420 | +0.421 | True | True | | brave fight | +0.180 | +0.268 | +0.267 | False | False | | guards fight | +0.434 | +0.398 | +0.388 | True | True | | devout pray | +0.199 | +0.201 | +0.200 | True | True | | thieves steal | +0.246 | +0.245 | +0.244 | True | True | ### Village runs (descriptive, same engine working tree for both models) lab-siege-v2 baselines (no cards), seeds 1-3, engine RNG keyed, race draw: | model | seed | deaths | thefts | caught | friendships | collapsed | share changing action per tick | |---|---|---|---|---|---|---|---| | layatown-v1 | s1 | 0 | 23 | 11 | 7 | 18 | 0.322 | | layatown-v1 | s2 | 0 | 22 | 12 | 6 | 8 | 0.330 | | layatown-v1 | s3 | 0 | 28 | 12 | 4 | 6 | 0.337 | | layatown-self-1 | s1 | 0 | 20 | 12 | 5 | 10 | 0.329 | | layatown-self-1 | s2 | 0 | 22 | 12 | 6 | 7 | 0.329 | | layatown-self-1 | s3 | 0 | 23 | 13 | 3 | 6 | 0.342 | Survival v2 (`lab-survival-v2`, seed 2, empty draft): | model | outcome | days survived | deaths | thefts | threats beaten | saved | friendships | score | |---|---|---|---|---|---|---|---|---| | layatown-self-1 | fell | 9 | 26 | 65 | 8 | 101 | 11 | 1020 | | layatown-v1 | fell | 9 | 25 | 70 | 8 | 104 | 12 | 1029 | One run per model and three seeds are far too few for inference: these rows are descriptive only. ## Reading (after the criteria were applied; not part of the pre-registered measures) - S1 passed, but the gain is small: option-order TVD 0.0155 -> 0.0147 (about 5%), state-key TVD 0.053 -> 0.047 (about 11%). v1 was already nearly invariant to option order, so there was little left to remove. - S2 failed on one margin only: juror-vote MAE rose by 0.012 (margin 0.01; n = 168 votes in 14 trials, interval above 0). Undoing the new noul temperature leaves it at +0.012 (0.0732 vs 0.0615), so it is the weights. The juror anchors were the blurriest targets (4 key orders of a juror state disagreed on the side in 41% of the 251 states), repeated 20x as in v1's recipe: the student learned to hedge juror votes toward 0.5 (mean p(guilty) 0.500 vs v1 0.494, Jev 0.479). This is the 'averaging blurs sharp answers' failure the plan warned about. - Everything else is within noise of v1: decision top-1 -0.01 pt, TVD +0.0002, sheet MAE +0.0001, D1 identical, D3 identical or slightly up without significance, character contrasts the same 4 of 5 passes, village runs alike.