Skip to content
CodeWiki
Practice
Paths
Tracks
Cheatsheets
Playground
Glossary
AI era
Search
⌘K
English
English
Chinese
Practice
Paths
Tracks
Cheatsheets
Playground
Glossary
AI era
English
English
Chinese
Practice
/
Quiz
/
AI and LLM engineering
/
AI and LLM engineering
A quantized local model file fits in GPU memory. What must capacity planning still includ…
from AI and LLM engineering
advanced
1 min
A quantized local model file fits in GPU memory. What must capacity planning still include?
KV cache, runtime buffers, context length, concurrent requests, and runner-specific overhead
Nothing else; model file size is an exact upper bound on runtime memory
Only network bandwidth, because quantization eliminates activation and cache memory
Check
Ask AI about this kata
Report an error
previous kata
Which observability record is most useful and safest for debugging a slow RAG answer?
Quiz
next kata
Name the observability unit that represents one complete application request as related s…
Fill in
A new version is available
Reload