Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference
A conventional all-attention model of the same size on the same data and a conventional all-attention hybrid that beats GPT-2 124M, Pythia-160M, OPT-125M and GPT-neo-125M, and exceeds MobileLLM-125M's published score despite that model seeing a trillion tokens.
Christos Koutsiaris
· 0 citations