# Layatown serving on one RTX 3080 Ti (the production city: jevcity, 1,000 people, 8 s ticks) Model server: layatown/serve.py, checkpoint v1 (sha256 27f210cb…), bf16, length-sorted token-budget batching, exact cache. City: `bun sim --world jevcity --scenario city-week --policy laya` at real speed (`--interval 8`), which uses the same tick loop as the live server. Logs in this folder. | test | result | |---|---| | First scale test (`--fast`, 60 ticks, 4 s request limit inherited from Jev) | 138 shard requests timed out on 29 of 60 ticks (117 of them on shards 4-9): 3 ticks fully skipped, 26 partial (the people in a timed-out shard keep their current action); request p50 2.07 s, p95 3.83 s. The GPU works through a tick's shards one after another, so the later shards met Jev's 4 s limit | | Where the time goes (100 questions) | CPU slicing + tokenizing 39 ms; GPU about 700 ms (after batching fix; 810 ms before). The GPU is the limit: about 140 decisions per second | | Fix | Laya's request limit is the tick minus 1 s (7 s live); batching sorted by length (13% faster) | | Real speed, 75 ticks (10 min) | **0 of 75 skipped**; 40,238 decisions (536 per tick); request p50 1.96 s, p95 5.67 s; exact-cache hits 25%. **Correction:** this run went through `bun sim`, which then gave Laya a 15 s limit, not the live server's 7 s. 4 of 435 requests took 7.05-7.25 s (shard 9 of ticks 6, 12, 14, 18): under the live limit those four ticks would have been partial, not skipped. `bun sim` now uses the live limit at real speed; the run is repeated below | | Energy | GPU mean 207 W over the 10 minutes (idle time between ticks included), max 76 °C: about 34.5 Wh for 40,238 decisions, 0.86 Wh per 1,000 decisions. Jev: $0.015 per 1,000 | | Failure: model server killed at tick 15, back about 96 s later | ticks 15-25 skipped cleanly (each failed in under 1 ms), no crash; ticks resumed by themselves at 26, request p50 back to ~750 ms | | Soak: 900 ticks (2 h, `lt1-soak-jevcity-laya`) | **0 of 900 skipped, 0 failed requests**, 478,971 decisions. Request p50 0.96 s, p95 4.34 s, p99 6.70 s, max 9.13 s. It ran with the sim's then 15 s limit: **36 of 5,203 requests (0.7%) took over 7 s**, so under the live limit about 0.7% of shards would be partial (those ~100 people keep their current action for one tick). Server memory flat at 1.70-1.79 GB, GPU memory at most 4.95 GB, GPU at most 77 °C (`soak-resources.csv`, one row a minute) | | Real speed with the live 7 s limit (75 ticks, `lt1-realtime7-jevcity-laya`) | 0 of 75 skipped; 1 of 432 requests timed out (one partial tick); request p50 0.62 s, p95 1.82 s. Caveat: the model server's cache was warm from the soak of the same scenario, so this is faster than a cold start; the soak's 0.7% is the conservative figure |