KV cache
Runtime attention keys and values retained from prior tokens so decoding can reuse them instead of recomputing the full prefix.
Runtime attention keys and values retained from prior tokens so decoding can reuse them instead of recomputing the full prefix.