Vision Transformers under TinyML Constraints in AIoT: A Systematic Review of Deployment Evidence
Abstract
Background. Vision Transformers (ViTs) lead many visual-recognition benchmarks, but their memory and compute demands exceed the budgets of Artificial Intelligence of Things (AIoT) devices by orders of magnitude. Work claiming to close that gap reports latency, energy and memory on incomparable hardware and under unstated conditions. Objectives. To map which ViT architectures and efficiency techniques are deployed on which classes of constrained hardware (RQ1–RQ2), which domains and modalities they serve (RQ3), and how completely deployment evidence is reported (RQ4). Methods. Eight publisher platforms were searched for peer-reviewed work from January 2020 to September 2026, supplemented by two OpenAlex queries that repaired recall defects found during piloting. Of 4,182 records, 3,883 were screened at title–abstract level by a large language model under a fixed, published guide; an independent blinded second pass measured agreement and bounded the false-exclusion rate. Evidence was mapped against five hardware classes defined by the processor that executes inference. Results. Screening excluded 2,867 records and forwarded 1,016 to full text, 180 of which met every criterion at abstract level. Agreement between passes was κ = 0.93 (n = 382); after adjudication, the 95% upper bound on records wrongly excluded was 8.3% of those retained (11.5% before adjudication). Microcontroller deployment is nearly absent: 3 of the 180 records (1.7%) target an MCU, against 68 (38%) on FPGA/ASIC accelerators, 40 (22%) on embedded GPUs and 36 (20%) on application-processor CPUs; 26 (14%) name no target at all. Only 16 abstracts (9%) report memory and 61 (34%) report energy, whereas 124 (69%) report latency. Medical (174 of 1,016) and agricultural (115) applications dominate the applied work; thermal, depth and event-camera sensing together account for 27 records. Conclusions. The literature labelled "ViT at the edge" is overwhelmingly accelerator- and GPU-class work; the TinyML class that motivates it is almost unoccupied, memory — the binding constraint on a microcontroller — is the least reported metric, and energy appears in only a third of abstracts. A full-text audit of deployment evidence with DERS, a 12-criterion instrument introduced here, is the next stage.