Cost-driven interest in running large language models (LLMs) on non-mainstream silicon outpaces the maturity of the surrounding software stacks. This paper uses one such platform, the AMD BC-250 (a repurposed cryptocurrency-mining board with a GFX1013 "Cyan Skillfish" accelerated processing unit (APU), 16 GB of unified...
Artur Andrzejczak· Zenodo (CERN European Organi...· 0 citations
Cost-driven interest in running large language models (LLMs) on non-mainstream silicon outpaces the maturity of the surrounding software stacks. This paper uses one such platform, the AMD BC-250 (a repurposed cryptocurrency-mining board with a GFX1013 "Cyan Skillfish" accelerated processing unit (APU), 16 GB of unified...
Artur Andrzejczak· Zenodo (CERN European Organi...· 0 citations
Harnesses and raw results for the experiments added in the J.UCS revision of the paper Deploying LLM Inference on a Repurposed UMA APU: Transferable Lessons from a Vulkan-Only, 16 GB Edge Platform. Contents: benchmarks/bench-a*.py, benchmarks/diag-*.py, benchmarks/heads.py: measurement harnesses (KV-cache width versus...
Artur Andrzejczak· Zenodo (CERN European Organi...· 0 citations
Harnesses and raw results for the experiments added in the J.UCS revision of the paper Deploying LLM Inference on a Repurposed UMA APU: Transferable Lessons from a Vulkan-Only, 16 GB Edge Platform. Contents: benchmarks/bench-a*.py, benchmarks/diag-*.py, benchmarks/heads.py: measurement harnesses (KV-cache width versus...
Artur Andrzejczak· Zenodo (CERN European Organi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.