Attention offloading distributes LLM inference operations between high-end accelerators and consumer-grade GPUs to reduce costs.
Cibo e viaggi / Food and travel notes by Livio Acerbo

Attention offloading distributes LLM inference operations between high-end accelerators and consumer-grade GPUs to reduce costs.