Device-Aware Distillation of Small Language Models for Efficient Edge Intelligence
Vinamra Sharma, D. Pau, José Cano
· 0 citations
We have 2 of 9 papers
We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.
Not the right person? Other researchers publish under this name.
Hydra is presented, a common-schema, phase-aware workload characterization framework for LLM inference on edge SoCs that enables reproducible, phase-aware characterization of edge LLM inference.
We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.