FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

AK's Picks LLM MQIA MoE AI Systems and Hardware EDIO
前沿开源权重模型正日益普及,但其部署服务目前仍主要依赖数据中心级别的基础设施。本文提出 FreeToken——一种面向边缘设备原生设计的混合专家(MoE)推理服务系统,它不再将个人计算机简单视为一块小型 GPU,而是将其建模为一个统一、弹性可伸缩的推理平台。FreeToken 针对本地人工智能应用的两大现实特征,对整个服务栈进行了协同设计:一是智能体(agent)类工作负载的执行模式持续动态变化;二是边缘硬件资源具有高度异构性,且不同设备间的资源配比差异显著。这两大特征具体体现在模型布局与加载、专家模块驻留策略、CPU–GPU 协同执行、智能体状态复用机制以及运行时内存管理等关键环节。FreeToken 并不预先固化某种卸载策略,而是持续地、动态地将计算任务与模型状态映射到设备当前实际可用的资源之上。FreeToken 支持超过 20 种 MoE 模型,并已在涵盖从仅配备 8GB 显存的笔记本 GPU 到单块工作站级 GPU 的多种硬件平台上,成功运行真实的代码生成与工具调用类智能体。更重要的是,它显著拓展了各类终端设备的实际服务能力:在普通笔记本上即可部署 35B 参数规模的模型,在高端游戏台式机上可运行高达 284B 参数的模型,甚至可在单块工作站级 GPU 上直接部署参数量达 753B 的 GLM-5.2 模型。FreeToken 将开源权重真正转化为可即装即用的本地化软件,使用户手中已有的终端设备成为支撑前沿级人工智能能力的切实可行平台。本系统已在 flashml.ai 开源发布。
Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure. We present FreeToken, an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform. FreeToken co-designs the full serving stack, including model layout and loading, expert residency, CPU--GPU execution, agentic state reuse, and runtime memory management, around two realities of local AI: agent workloads continuously change their execution pattern, and edge hardware exposes heterogeneous resources whose balance differs from machine to machine. Rather than committing to a fixed offloading strategy, FreeToken continuously maps computation and model state onto the resources actually available. FreeToken supports more than 20 MoE models and real coding and tool-using agents across hardware ranging from an 8GB laptop GPU to a single workstation GPU. More importantly, it changes what these machines can practically serve, from a 35B model on a laptop to a 284B model on a gaming desktop and the 753B GLM-5.2 on a single workstation GPU. FreeToken turns open weights into deployable local software, making the machines users already own a practical platform for frontier-scale intelligence. We release the system at flashml.ai.