The study Can a free local model replace Jev in Jevtown?

Can a free local model replace Jev in Jevtown?

Rendered from layatown/docs/EVALUATION.md in the project repository, 21 September 2026, word for word. Laya is an open model by Convai Innovations (Apache-2.0); Jev is TypeSafe's model. Not affiliated with either. Cited result files are listed in the study.

Pre-registered evaluation plan for Layatown, written 21 September 2026, before any fine-tuned checkpoint is scored on the held-out production day. It fixes the systems, data, metrics, statistics and pass/fail thresholds so that results cannot be chosen after the fact. It follows the project's rule for docs/EVALUATION.md: the plan is committed before the runs, never edited to fit results, and changed only by dated amendments at the end of this file.

1. Summary

Jevtown's citizens are driven by TypeSafe's Jev (jev-1.13.0), a paid System One model. Layatown is Laya (convaiinnovations/laya, ModernBERT-large, 421M parameters, Apache-2.0, revision 1c5edc17), fine-tuned by distillation on Jev's own logged probabilities from Jevtown (training/, see its data card training/README.md), and served on one local RTX 3080 Ti with the same POST /v1/systemone wire shape (layatown/serve.py).

The question is whether Layatown can take Jev's place in the live city. We ask it at three levels, from narrow to broad, because agreement on single answers does not guarantee that a city run on those answers behaves the same:

No part of this plan requires new Jev spend.

2. Research questions

3. Systems

idsystemrolenotes
jevJev jev-1.13.0, the logged answersreference (the teacher)Never called again. Its answers are the targets in training/data and the logs in runs/.
laya-baseLaya at revision 1c5edc17, no fine-tuning, raw temperatureslower referenceTells us how much fine-tuning contributes. Already partly seen on the test day (section 9).
layatown-v1Laya fine-tuned on 150,000 sampled train decisions + all train character-sheet answers + train juror votes repeated x20, one epoch; one temperature per question type fitted on valsystem under testPrimary configuration uses the fitted temperatures; raw temperatures reported alongside.
floorthe train-mean answer (section 5.4)floor any model must beatFrom the data card.
rules, randomthe project's existing free baselinesreferences at level C onlyLogged runs already exist; no new runs needed except seeds 6 to 10 (section 7.3).
later variants (layatown-v2, ...)e.g. all train decisions, more epochs, different juror weightingoptional, exploratorySee section 10.5: chosen on val only, all test results reported.

Before layatown-v1 is scored on the test day, the SHA-256 of its model.safetensors and its rl_agent_config.json (temperatures, training arguments) are recorded as an amendment. That checkpoint is frozen: no change to weights, temperatures, slicing or prompt format after the test day is opened.

4. Data

5. Level A: judgment fidelity on the production day

5.1 Unit and estimand

The unit is one question (one citizen's decision, one yes/no line of a sheet, one juror's vote). Every metric is the item-weighted mean over all questions of the task in test_production, with no subsampling (--every 1). Items that data.py cannot build (an option marker falls off the sequence) are counted and reported; if more than 0.1% of a task's items are dropped, the result for that task is reported as incomplete.

5.2 Metrics (exact definitions)

Let p be the model's distribution and q Jev's target distribution over the same options, in criteria order.

villager_decision (choice, 8 to 11 options)

character_sheet and juror_vote (noul)

5.3 Character effects on the test day (RQ2)

Using the states of the test-day decision items, compute for each pre-registered contrast the difference effect = mean p(action | in-group) - mean p(action | others), once with Jev's targets and once with each model's predictions, on exactly the same items. Group and item selection follow the definitions in bun report (src/cli/report.ts):

contrastin-groupitemsaction
cowardly fleetrait cowardlydanger here or nearflee
brave fighttrait bravedanger here or nearfight
guards fightjob guarddanger here or nearfight
devout praytrait devoutallpray
thieves stealjob thiefallsteal

Jobs or traits absent from the city world are reported as not applicable, not as passes.

5.4 Floor

For each choice question, the floor predicts the mean train target over train questions with the same ordered option set; for nouls, the train mean for that task (0.39 for sheets, 0.49 for jurors). The data card reports TVD 0.517, top-1 0.278, sheet MAE 0.142 and juror MAE 0.112 on the test day. Because no committed script currently produces those numbers, the evaluator recomputes the floor on the identical item set; both the card's and the recomputed values are reported, and thresholds that reference the floor use the recomputed value.

5.5 Uncertainty: bootstrap over ticks

Questions within one tick share the village state, the moment's events and often the same citizen seen a tick earlier, so they are not independent. Confidence intervals are therefore cluster bootstrap intervals with the tick as the cluster:

5.6 Pre-registered thresholds (level A)

layatown-v1 (fitted temperatures) passes level A if all of the following hold:

idcriterion
A1Decision TVD <= 0.15, and the upper 95% bound <= 0.18
A2Decision top-1 agreement >= 0.80, and the lower 95% bound >= 0.77
A3Character-sheet MAE <= 0.07, and the upper 95% bound <= 0.09
A4Juror MAE: upper 95% bound of (model MAE - floor MAE) < 0 (it beats the floor). Target, reported but not required: MAE <= 0.08
A5For every applicable contrast in 5.3: Layatown's effect has the same sign as Jev's, and lies within max(25% of Jev's effect, 0.05) of Jev's effect

A1 to A5 form one conjunction (intersection-union), so no correction for multiple comparisons is applied. For context, the val partials seen during training were TVD 0.114, top-1 0.865, sheet MAE 0.039; the thresholds are set looser than val on purpose, because the test day is a different, larger, live run.

6. Level B: human-labelled evaluation sets

These sets are small and some involve text typed by or about real people. They are evaluated locally only; rows are never copied into results files, only aggregate counts. Layatown was not trained on chat moderation, username screening or event reading: any skill it shows there comes from the base model.

The rows record the case and Jev's answer, not the full request. Each set is scored by rebuilding the request with the exact question wording from the benchmark scripts named in the data card (bench/chat.ts, bench/handles.ts, bench/shapes.ts, bench/trial.ts) and sending it to serve.py locally. Jev is represented by its logged answers.

setrowshuman labelwhat is comparedpre-registered gate
chat_moderation38show / blockaccuracy vs human; lines a person would block that the model shows (misses); harmless lines blockedB1: Layatown misses no more abusive or private lines than Jev
chat_neighbour_effect20none (fairness probe)change in a harmless line's abuse score alone vs beside abusive linesreported; flagged if Layatown's mean shift exceeds Jev's
username_screening36ok / not okmisses (hateful name passed) and false alarms (Scunthorpe cases)B2: Layatown misses no more hateful names than Jev and false-alarms on no more innocent ones than Jev plus 2
event_reading148event type or refusalaccuracy vs human; refusals of fine proposals; accepted proposals that should be refusedB3: accuracy no more than 5 points below Jev's
juror_factorial192none (Jev's p only)MAE to Jev; direction of each factor's main effect (evidence, tie, memory, profile, wording) vs Jev'sreported; a reversed main effect is listed as a failure mode

Consequence of the gates: in --policy laya the city sends every request kind to Layatown, including chat, handle and proposal screening. A failed B1, B2 or B3 means Layatown must not serve that request kind in the live city (it goes to code-only screening or stays with Jev); the finding is reported and does not by itself fail the replacement question for citizen decisions.

7. Level C: city-level behaviour

7.1 Design

The simulator is deterministic given the world, seed and the decision-maker's answers (docs/ARCHITECTURE.md). All conditions share the engine, world data, perception, sampling (jev.floor, jev.stickiness), commitment rule and logging; only the decision-maker changes, exactly as in docs/EVALUATION.md.

7.2 Metrics

The existing bun compare metrics, unchanged, mean over seeds with 95% Student t intervals (n = 5): population at end, deaths, beatable threats overcome, injuries, median ticks until half the village knows, exhaustion collapses, thefts, friendships formed, quarrels, mean food stock, share asked per tick, share changing action per tick, action entropy, individuality (mean JSD to village), story beats, latency. Plus the bun report character contrasts per run (permutation p-values).

7.3 Pre-registered thresholds (level C)

Per scenario, layatown-v1 passes level C if all of the following hold:

idcriterion
C1Viability: population at end >= 75 of 80 in every seed, and mean share of beatable threats overcome >= 0.80 (Jev: 0.96, 1.00, 1.00)
C2Character (the study's H3 rule): each of the five contrasts in 5.3 has p < 0.01 in at least 4 of 5 seeds
C3Closeness: of these 8 metrics (thefts, friendships formed, exhaustion collapses, mean food stock, injuries, action entropy, individuality, share changing action per tick), Layatown's seed mean is closer to Jev's seed mean than rules' is on at least 6
C4Individuality survives: mean individuality >= 0.8 x Jev's mean, and its 95% interval lies above rules' interval

7.4 What level C can and cannot show

All 15 eval6-/eval7- Jev runs are in the train split, so Layatown has been trained on Jev's decisions from exactly these scenarios and seeds. The matched comparison therefore tests whether imitation survives the closed loop (small per-decision errors compounding over 240 ticks into a different village), not generalisation. Trajectories diverge from Jev's within a few ticks because sampling and the world react to different answers, so the model soon sees perceptions that were never in training; but the early ticks are in-sample. The seeds 6 to 10 sensitivity set removes that overlap for the trajectory. If C1 to C4 pass on seeds 1 to 5 but fail on 6 to 10, the level C result is reported as failed.

7.5 The city (descriptive only)

The live city (jevcity) cannot be replayed under matched conditions: city-1 was run on world hash e05c5e46962d5688 and the current world hashes 07806946f4cc1ea5, and the production day depended on live players. One Laya city-week run on the current world is reported descriptively (population, deaths, per-action shares, latency), with no threshold.

7.6 Serving cost and latency (RQ5)

Measured on the RTX 3080 Ti, reported regardless, no pass/fail except one feasibility gate:

8. Verdict rules

9. What was seen before this plan (disclosed exposures)

10. Threats to validity

  1. Fidelity is not correctness. Level A measures agreement with Jev, not with the truth. Distillation copies Jev's biases and errors along with its judgment: Jev's known failure modes (freezing when shown its own last answer, joining the crowd nearby, the neighbour effect in chat, the two-step inference it fails in famine, docs/JEV.md) are expected to transfer. A high score means "as good and as bad as Jev". Only level B touches correctness, on 445 rows.
  2. The test day is not a new domain. It is a new moment of the same city world, with the same prompts, option wording, perception format and names style as train. It tests generalisation to new moments and people, not to new worlds or questions.
  3. Context mismatch. Jev answered each citizen with up to 99 others in the same request; the dataset and serve.py give Layatown only the citizen's own slice (data card, "One item per example"). Level A inherits the dataset's slicing; level C inherits serve.py's. If the two slicings differ (for example in which top-level collections are kept, or option order), A and C measure slightly different inputs.
  4. Temperatures fitted on val. Val is four live city runs, close in distribution to the test day. Fitted temperatures could be slightly optimistic for the city and less right for the village used in level C. Raw temperatures are reported for that reason.
  5. List-position runs excluded. 407,653 decisions were dropped as wrong labels. The remaining data is later in the project, so it over-represents the post-fix world and the live city; the village runs are in train only.
  6. In-sample closed loop. Level C matched seeds are in train (7.4). Mitigated by seeds 6 to 10.
  7. Rounded, near-deterministic, merged labels. Jev's probabilities were logged to two or three decimals and duplicate requests averaged; a perfect student cannot score below the rounding noise, and Jev's own rerun variation (about 0.016 on juries) is a practical floor on MAE.
  8. Small sets. 168 juror votes and 445 human-labelled rows give wide intervals; conclusions from them are directional.
  9. Implementation drift. GPU inference in bf16 with batch-dependent padding is not bit-identical across batch compositions; runs remain replayable because decisions acted on are logged, but a rerun of a Laya run may not reproduce it exactly.
  10. Licensing and terms. Laya is Apache-2.0 (inherited by the fine-tuned weights; attribution and licence notice required). The owner confirmed on 21 September 2026 that TypeSafe's terms permit training on these outputs (data card, "Provenance and permission"). Terms can change; check them before redistributing the weights or the data outside this project.
  11. Same authors. The people who built Layatown wrote this plan and the thresholds. The thresholds are fixed here, before the test, to limit that.

10.5 Later variants

A variant trained after this plan (more data, more epochs, different weights per task) is chosen among others on val only. Each frozen variant is scored on the test day at most once, and every test score of every variant is reported, including the ones that lose. A variant designed after seeing layatown-v1's test results is labelled exploratory, and its test numbers are not used for the verdict of section 8 unless a new, untouched test day (a later production run, built by training/build.ts into a new split) is used for it.

11. Results

To be filled after the frozen run. Do not edit sections 1 to 10 to fit these.

11.1 Frozen checkpoint

fieldvalue
model.safetensors SHA-256TBD
temperatures (choice, score, noul)TBD
training data / steps / minutesTBD
date scored on test dayTBD

11.2 Level A (test_production, 95% tick-bootstrap intervals)

taskmetricfloorlaya-baselayatown-v1 (raw T)layatown-v1 (fitted T)thresholdpass
decisionTVDTBDTBDTBDTBDA1TBD
decisiontop-1 agreementTBDTBDTBDTBDA2TBD
decisionKL / ECETBDTBDTBDTBDreported
sheetMAETBDTBDTBDTBDA3TBD
sheetBrier / side agreementTBDTBDTBDTBDreported
jurorMAETBDTBDTBDTBDA4TBD
jurorBrier / side agreementTBDTBDTBDTBDreported
alldropped itemsTBDTBDTBD<= 0.1%TBD

Character effects (A5):

contrastJev effectlayatown-v1 effectlaya-base effectpass
cowardly fleeTBDTBDTBDTBD
brave fightTBDTBDTBDTBD
guards fightTBDTBDTBDTBD
devout prayTBDTBDTBDTBD
thieves stealTBDTBDTBDTBD

11.3 Level B

setJevlayatown-v1laya-basegatepass
chat_moderationTBDTBDTBDB1TBD
chat_neighbour_effect (mean shift)TBDTBDTBDreported
username_screeningTBDTBDTBDB2TBD
event_readingTBDTBDTBDB3TBD
juror_factorial (MAE to Jev, reversed effects)n/aTBDTBDreported

11.4 Level C (mean over seeds, 95% t intervals)

One table per scenario (hard-week, siege, strange-week), columns jev (eval6/eval7, s1-5), laya (s1-5), laya (s6-10), rules (s1-5), rules (s6-10), random; rows the metrics of 7.2; then C1 to C4 pass/fail.

TBD

Ledger total before / after level C: TBD / TBD.

11.5 Serving (RTX 3080 Ti)

measureJev (logged)layatown-v1
village tick request latency p50 / p95~340 ms / ~660 msTBD
1,000-decision tick, cache off, p50 / p950.98 sTBD
1,000-decision tick, cache on, p50 / p95n/aTBD
energy per 1,000 decisionsn/aTBD Wh
cost per 1,000 decisions$0.015$0 marginal API cost (electricity reported as energy)
D1 gateTBD

11.6 Verdict

TBD (one of the three outcomes in section 8, with the failing criteria named).

Amendments

The plan above is left as first written. Changes are listed here with their date and reason.