Skip to content

Deploying LLM Inference on a Repurposed UMA APU: Transferable Lessons from a Vulkan-Only, 16 GB Edge Platform

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)
Parallel Computing and Optimization Techniques

Abstract

Cost-driven interest in running large language models (LLMs) on non-mainstream silicon outpaces the maturity of the surrounding software stacks. This paper uses one such platform, the AMD BC-250 (a repurposed cryptocurrency-mining board with a GFX1013 "Cyan Skillfish" accelerated processing unit (APU), 16 GB of unified GDDR6 memory and a Vulkan-only GPU compute path), as a single-board testbed. The primary contribution is a reproducible characterisation of that platform, released with the artefacts needed to repeat it. Three observations are not specific to the board. Two are benchmark-validity cautions: (i) a silent prompt-truncation failure mode in the Ollama runtime, which affects filled-context benchmarking unless the API's prompt_eval_count field is verified per request, and (ii) the distinction between context allocation and context utilisation, which can produce headline "128K context" numbers backed by a much shorter prompt. The third (iii) is confirmatory: sparse mixture-of-experts (MoE) generation throughput tracks the active-parameter count when matrix accelerators are absent, observed here on hardware where, to the author's knowledge, it had not been measured before. A per-token byte accounting from the model tensor tables predicts the observed separation from a same-size dense comparator at matched quantisation to within 4 %, and the deployable edge over 14 B dense models is a modest ≈14–25 %. The evidence rests on a 31-model cohort (3–35 B parameters) and covers generation speed (run-to-run within-cell coefficient of variation ≤0.5 % across the canonical n=3 cells under a pinned governor), filled-context scaling with real-token payloads, a five-task competence-preservation probe (a check that a configuration is not broken, not a capability score), cold-start latency, and clock and thermal telemetry. Three runtime settings the results depend on were measured directly. The 100-token decode window and flash attention hold. Quantising the KV cache to 4 bits costs up to 1.6 % perplexity on most models, more on several others, and collapses the small Qwen2-architecture builds through their key cache, whereas an 8-bit cache stays within ±0.5 % of FP16. The deployment recipe, a kernel TTM (Translation Table Manager) pages_limit adjustment and an FP-validated reuse of a community 40-CU core unlock, is documented as the enabling infrastructure. All measurements come from a single board, driver, and firmware; the scope limits this implies are stated in the threats analysis. Version 3 is the accepted manuscript of Journal of Universal Computer Science submission #204239 (accepted 3 October 2026, scheduled for J.UCS 33(6), June 2027), in the form submitted for copy editing. Version 2 was the manuscript as submitted for review on 16 June 2026. The measurement artefacts (harnesses, raw results, supplementary material) are archived separately at 10.5281/zenodo.22668983.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.