BC-250 LLM inference benchmarks: harnesses and raw results for the J.UCS revision (akandr/bc250, jucs-revision-5)
Abstract
Harnesses and raw results for the experiments added in the J.UCS revision of the paper Deploying LLM Inference on a Repurposed UMA APU: Transferable Lessons from a Vulkan-Only, 16 GB Edge Platform. Contents: benchmarks/bench-a*.py, benchmarks/diag-*.py, benchmarks/heads.py: measurement harnesses (KV-cache width versus task accuracy and perplexity, decode-window sensitivity, flash-attention correctness, GGUF per-token byte accounting, competence probe on the canonical stack, key/value-head analysis). benchmarks/results-jucs-revision/: raw per-run JSON and logs for every measurement reported in the paper. jucs-supplementary/: the supplementary document (source and PDF) with the per-cell tables. article/data/_provenance_ledger.yaml: provenance ledger with a revision note. All measurements were taken on one AMD BC-250 (Fedora 43, kernel 6.18.9, Mesa 25.3.4, Ollama 0.20.0 / llama.cpp b9265) on the verified-stock 24-CU configuration and repeated under the 40-CU unlock. See Readme.md and jucs-supplementary/README.md. Licences: code under AGPL-3.0-or-later; the repository Readme under CC BY-SA 4.0 (see LICENSE). This record is a supplement to the preprint of the paper (10.5281/zenodo.19476016). jucs-revision-4 adds, relative to jucs-revision-1: (A14) the Q4_0 KV-cache collapse tested outside the Qwen2 architecture on six models with one to five key/value heads and on a Qwen2-architecture model with different post-training, plus keys and values quantised separately on two collapsing builds (bench-a14-heads-control.py, step-a14-heads-control.json); (A15) the arithmetic item of the competence probe re-run at both KV widths as the first request of fresh sessions and under three request orders (bench-a15-arith-sequence.py, step-a15-arith-sequence.json); and the updated supplementary document (sections S9 and S10, tables S1 to S11). The fourth snapshot adds the A4 GGUF tensor-table byte accounting run over every model file on the board (bench-a4-scan-board.py, step-a4-gguf-bytes-all.json, 52 files, metadata only) and the analysis that turns per-token weight bytes times the measured decode rate into effective bandwidth per quantisation class (analyze-a4-attainment.py, a4-attainment.json), which backs the supplementary effective-bandwidth table; the supplementary document is synchronised with the revised submission (15 pages). jucs-revision-3 was tagged on the wrong commit and is identical in content to jucs-revision-2. jucs-revision-5 (this record) changes the supplementary document only; harnesses and raw results are identical to jucs-revision-4. Two Table S11 cells whose medians are exact halves (granite-4.0-h-tiny at 32K, gpt-oss-20b at 64K) are rounded half-up as in Table 13 of the paper (115.5, 57.7 tok/s), and Figure S2 is labelled with the per-model speed-ups printed in Table 12 of the paper (1.34x, 1.57x for qwen3.5:9b and mistral-small-3.2), computed from the unrounded per-run medians. This is the snapshot cited in the revised manuscript's data-availability statement.