GPU memory12 min read
Memory requirements of language model inference on a consumer GPU
What occupies GPU memory while a language model runs, calculated for Qwen3-14B on an 8 GB laptop GPU: system use, number formats, model parts, and the KV cache.
Running large language models on small GPUs · Part 1
Field note / 01Model memory map
Qwen3-14B · BF16 weights
29.5 GB
Measured against an 8 GB laptop GPUOney ErgeApplied AI + systems