5
Adaptive Model Compression (AMC): Saliency-Driven Resource Allocation for Ultra-Low-Power Transformer Inference
The proposed Adaptive Model Compression framework dynamically allocates hardware resources during transformer inference based on token importance to reduce energy and memory usage on edge devices.
Impact
46/100
Current rank score
5.03
Source tier
Tier 1
Category
Research
Firefly links to the original publisher. The summary above is AI-generated for orientation and may differ from the source. The “current rank score” decays over time so newer significant stories surface first.