Kafila: Serving Large Language Models on Heterogeneous Consumer Devices

Kafila serves LLMs on a small, trusted group of heterogeneous consumer devices by planning block placement and NAT-traversing ring order from each device’s measured bandwidth, capacity, and reachability. This cuts the pipeline’s slowest stage by 4.2x versus uniform division and enables serving models uniform division can’t fit at all.

September 2026 · Murtaza Rangwala, Richard Sinnott, Rajkumar Buyya