Entity
Tensor-parallelism
A distributed computing technique that spreads individual neural network layers across multiple GPUs or servers to handle models that exceed single-device memory and compute capabilities.
术语属性
- category:Distributed Computing
- application:Large Language Models
- challenges:Orchestration Complexity,KV Cache Coordination