Kafila: Serving Large Language Models on Heterogeneous Consumer Devices
Kafila serves LLMs on a small, trusted group of heterogeneous consumer devices by planning block placement and NAT-traversing ring order from each device’s measured bandwidth, capacity, and reachability. This cuts the pipeline’s slowest stage by 4.2x versus uniform division and enables serving models uniform division can’t fit at all.