GEOZ
让优质内容被 AI 引用
首页
GEO 分析工具
GEO
AI大模型
DeepSeek
llms.txt
schema
RSS
菜单
Entity
Disaggregated Serving
将推理过程拆分为prefill(预填充)和decode(解码)两个阶段,分别由独立服务器处理,以降低首token延迟。
相关文章
如何在Kubernetes上实现LLM分布式推理SOTA性能?llm-d v0.5实测50k tok/s