Entity
KV cache
Key-Value cache that stores intermediate computations during transformer model inference, significantly reducing redundant computations for repeated tokens and improving inference efficiency.
术语属性
- category:Performance Optimization
- purpose:Computation Reuse
- impact:Reduces Latency, Improves Throughput