Running large language models on small GPUs · Part 1
Memory requirements of language model inference on a consumer GPU
I run language models on a laptop GPU with 8 GB of memory, a size I think is common for consumer hardware. This series is about what happens when the model I want to run is much larger than that memory, and what can be done about it without changing the released model weights. This first part calculates what has to fit in that memory, and why this 14B model does not.
A language model is, in large part, a collection of numbers called weights that are learned during training. Text is converted into lists of numbers and passed through a series of layers. Each layer transforms the list using its own weights, then hands the result to the next layer.
A GPU is fast at this work because it runs thousands of multiplications at the same time. That speed only helps if the weights arrive as fast as the GPU can use them. The GPU's own memory sits next to its processor on the graphics card and, on my GPU, can be read at about 448 GB/s [5]. Data in the computer's main memory or on disk must first travel over the PCIe connection, which carries at most about 32 GB/s [5], more than 10 times less. Storage is slower still. For this reason, the weights are normally loaded into GPU memory before inference, the process of using a trained model to generate text.
The practical problem is size.
Throughout this series, I use Qwen3-14B, the 14B model of the Qwen3 family that Alibaba's Qwen team released on April 29, 2025 [1]. Its released BF16 weights require about 29.5 GB. My GPU, an NVIDIA RTX 3080 Laptop GPU, has 8 GB of advertised memory.
There are two practical ways to deal with this mismatch. One is to make the weights smaller, normally through quantization. The other is to keep most of the weights outside GPU memory and load them only when they are needed.
GPU memory during inference
Everything using the GPU competes for the same memory. For this discussion, I divide that memory into five parts:
- Msys
- System and display. Memory used by the operating system, the desktop, connected monitors, and other open programs before the model starts. On my laptop, with Windows, a browser, and a code editor open,
nvidia-smireported 2.2 GB in use. - Mrt
- GPU software. Memory held by the NVIDIA driver and by libraries such as PyTorch, outside the model itself. This depends on the software stack; measure it on your own machine with
nvidia-smi, the same way as Msys, before and after loading the library. - MW
- Weights. Fixed by the model and its number format. Calculated below.
- MKV
- Stored keys and values. Memory kept for every token of the input prompt and of the generated answer. It grows as the conversation grows. Explained below.
- Mact
- Activations. Temporary results inside the layer that is being computed, released when the layer finishes.
MGPU is the memory on the graphics card. GPU makers label memory in binary units, so my "8 GB" card holds 8.6 GB in the decimal units this post uses throughout. I still call it an "8 GB GPU" in prose, but every chart and calculation below uses 8.6 GB.
On my laptop, that leaves about 6.4 GB of physical capacity before any model weights, KV cache, runtime allocations, or activations are added.
Only MW and MKV can be calculated directly from the model. The other three terms depend on the machine and software stack and must be measured separately. The weights therefore cannot use the whole GPU by themselves.
Number formats for weights
The weights are by far the largest fixed term in Eq. (1), so the first question is how much memory one weight requires. The memory needed for the weights follows directly from how many there are and how many bytes each one takes:
where N is the number of weights and b is the number of bytes each weight uses.
A model is normally trained and released using one fixed number format for every weight. One way to fit a model into less memory is to store each weight with fewer bits than it was released with. This is called quantization, and it directly lowers b in Eq. (2). Fewer bits give fewer possible values per weight, so quantization discards some information and can reduce accuracy.
Table 1. Common number formats for language model weights. The size under each format is the storage for one weight; 1 byte is 8 bits.
| Format | Storage per weight | Typical use |
|---|---|---|
| FP32 | 4 bytes | Traditional training and reference calculations |
| FP16 | 2 bytes | Inference on GPUs, mixed-precision training |
| BF16 | 2 bytes | Training and release of recent open models |
| 8-bit integer | About 1 byte | Quantized inference |
| 4-bit integer | About 0.5 byte | Local inference on consumer GPUs |
BF16 was designed to keep the range of FP32 in half the memory [2]. It is the released format for Qwen3-14B. With N = 14.8 billion weights and b = 2 bytes, Qwen3-14B needs 29.5 GB.
Quantization reduces this requirement. For Qwen3-14B, 8-bit weights need about 14.8 GB and 4-bit weights need about 7.4 GB, before the additional metadata used by practical quantization formats.
This series keeps BF16. In this series, exact weights means that inference ultimately uses the same BF16 values released with the model, not an 8-bit or 4-bit approximation. Afterimage [6], the engine this series builds toward, is designed to run these exact weights: it does not introduce weight error through lossy quantization.
Where this post reports the same output tokens across methods, such as in Table 6, that is a separate measured result and not a guarantee that follows automatically from using exact weights.
Parts of a language model
To understand how a model larger than the GPU can still run, we need to know which parts of the model are used in sequence.
Model files and configuration
Downloading Qwen3-14B from Hugging Face gives three kinds of files. The weights are stored in several .safetensors files. A tokenizer splits text into tokens and gives each token an ID number. The file config.json describes the shape of the model.
Table 2. Entries of Qwen3-14B's config.json [1] used in this post.
| Entry | Meaning | Symbol | Value |
|---|---|---|---|
vocab_size | Number of different tokens | V | 151,936 |
hidden_size | Length of the numerical representation of each token | d | 5,120 |
num_hidden_layers | Number of decoder layers | L | 40 |
num_attention_heads | Attention heads per layer | 40 | |
num_key_value_heads | Sets of keys and values per layer | nkv | 8 |
head_dim | Length of each key and value | dh | 128 |
torch_dtype | Released number format | bfloat16 |
The number of layers is not a fixed property of language models in general. It is an architecture choice made when the model is designed and trained. Qwen3-14B has 40 decoder layers, and that value is recorded in config.json.
A token enters the model through the embedding table, passes through the decoder layers in order, and leaves through the output head.
Embedding table
The embedding table is the first stop for the input prompt. It takes each token of the prompt and looks up its row in a table, converting that piece of text into a list of numbers. Those numbers are what the decoder layers work with.
Table 3. A simple embedding table. The numbers are illustrative.
| Token ID | Token | Row of numbers |
|---|---|---|
| 0 | the | [0.2, −0.1, 0.7] |
| 1 | dog | [0.9, 0.3, −0.4] |
| 2 | ran | [−0.5, 0.8, 0.1] |
| 3 | in | [0.0, −0.6, 0.2] |
The text “the dog” splits into the tokens “the” and “dog,” and the embedding table replaces each one with its row from Table 3. No arithmetic happens at this step; it is a lookup.
Decoder layers
Coming out of the embedding table, every token has a list of numbers. Each decoder layer updates these representations using its own weights and information from earlier tokens.
Qwen3-14B repeats this 40 times, one decoder layer after another, allowing each token representation to incorporate information from the tokens that came before it.
Inside each layer, attention lets a token use information from earlier tokens, while a feed-forward step transforms the token representation further. Both operations use the layer's own weights. Each Qwen3-14B decoder layer contains about 330M weights.
Output head
The output head turns the decoder layers' final numerical representation into scores for possible next tokens. It compares the final representation with rows corresponding to the model's vocabulary and assigns each candidate token a score.
With the greedy decoding used in this series, the highest-scoring token is selected as the next token of the answer.
In Qwen3-14B, the output head contains about 778M weights.
Weight memory of each part
Table 4 applies Eq. (2) to each part.
Table 4. Weights of Qwen3-14B by part, in BF16. Calculated from Table 2. Rows may not sum exactly because of rounding.
| Part | What it does | Weights | Memory |
|---|---|---|---|
| Embedding table | Looks up the numerical representation for each token | 778M | 1.6 GB |
| One decoder layer | Mixes information between tokens and transforms each token | 330M | 0.66 GB |
| All 40 decoder layers | 13.2B | 26.4 GB | |
| Output head | Scores every token as the possible next token | 778M | 1.6 GB |
| Total | 14.8B | 29.5 GB |
Each individual model component fits on the GPU. The problem is keeping all 40 decoder layers resident together: they require about 26.4 GB.
Figure 1 draws the weights to scale against the GPU.
- Weights
- 29.5 GB
- GPU memory
- 8.6 GB
- Layers that fit
- 10 of 40
This sequential structure is important. The model contains 29.5 GB of weights, but all 29.5 GB do not have to be on the GPU at the same instant.
Memory for the prompt and the answer: the KV cache
LLMs generate text one token at a time. After reading the prompt, they produce the first token of the answer, add it to the sequence, and repeat. Each new token depends on the tokens before it.
The operation that retrieves information from earlier tokens is called attention.
The next subsection shows how queries, keys, and values are used in attention. If you already know how attention works, skip ahead to Storing keys and values instead of recomputing them.
A closer look at attention
Queries, keys, values, and attention are fundamental pieces of the transformer architecture [3].
In each layer, attention computes three short lists of numbers for every token: a query, a key, and a value. The query of the newest token is compared with the keys of the tokens available to it. A stronger match gives a higher score, which causes more of that token's value to contribute to the output.
Table 5 works through this for the text “the dog ran.” It scores each available token against the newest one, “ran,” converts those scores into weights, and blends the values into one output for “ran.”
For this simplified example, I omit the usual score scaling and the learned projections that produce queries, keys, and values, and show only the part needed to explain the KV cache.
Table 5. A simple attention step for the newest token, “ran,” whose query is [0, 2]. The numbers are illustrative.
| Token | Key | Value | Score, query × key | Weight | Contribution, weight × value |
|---|---|---|---|---|---|
| the | [1, 0] | [0.2, 0.1] | 0 | 0.06 | [0.01, 0.01] |
| dog | [0, 1] | [0.9, 0.4] | 2 | 0.47 | [0.42, 0.19] |
| ran | [1, 1] | [0.5, 0.7] | 2 | 0.47 | [0.24, 0.33] |
| Blended output for “ran” | [0.67, 0.52] |
Each score comes from multiplying the query by that token's key, number by number, and adding the results. Softmax then converts the scores into weights that add up to 1. The output is the sum of every token's value multiplied by its weight:
where si is the score of token i, wi is its weight, and vi is its value.
Storing keys and values instead of recomputing them
When the model adds the next token, “in,” the keys and values of the earlier tokens do not change.
Without storing them, the model would have to recompute those values every time another token is generated. Instead, inference software computes them once and keeps them in GPU memory.
This stored information is the KV cache.
During token-by-token decoding, the new token produces a new key and value in every decoder layer. The keys and values already stored for earlier tokens are reused.
Size of the KV cache
Qwen3-14B has 40 query heads per layer but only 8 distinct sets of keys and values. Five query heads therefore share each key/value set.
This design is called grouped-query attention [4]. Compared with storing independent keys and values for all 40 heads, it reduces the KV-storage requirement by a factor of 5.
The memory of the KV cache is
where the factor 2 counts keys and values, L, nkv, and dh come from Table 2, b is the bytes per number, and T is the number of tokens in the input prompt and answer together.
For Qwen3-14B in BF16,
or about 0.16 MB per token.
An input prompt and answer of 8K tokens together need about 1.3 GB.
At the model's standard 32K-token context, the KV cache alone needs about 5.4 GB.
At 128K tokens, it needs about 21 GB, more than the whole GPU.
Total memory for Qwen3-14B on an 8 GB laptop GPU
Eq. (1) can now be filled in.
Qwen3-14B needs about 29.5 GB for its BF16 weights. A short 3K-token input prompt and answer add about 0.5 GB for the KV cache. My laptop's system already occupies about 2.2 GB.
That is already more than 32 GB, before GPU software and temporary activations are even counted, far more than the GPU's 8 GB.
Even the simple 4-bit weight estimate leaves essentially no room for the rest of the workload. About 7.4 GB of weights plus the measured 2.2 GB already used by the system exceed the physical capacity considered here before the KV cache, runtime allocations, and quantization metadata are added.
That does not mean a 4-bit Qwen3-14B cannot run on an 8 GB system using memory-management techniques. It means the model weights cannot be treated as the only thing that has to fit.
The calculator below tests other combinations.
Memory check for Qwen3-14B
8.6 GB laptop GPU- System
- Weights
- KV cache (input prompt + answer)
- Total
Weights come from Eq. (2), and the KV cache from Eq. (4). GPU-software memory and activations are not included and must be measured separately. Quantized files are also slightly larger than N × b because they store scale factors and other metadata.
Loading one part of the model at a time
The model still does not fit, but the calculation above suggests another option: the GPU does not need every weight at the same time.
The weights can stay outside GPU memory, and only the model part currently needed is loaded onto the GPU: load it, use it, release its weight memory, then continue to the next part.
Inference happens in two stages.
First, the prefill stage processes the input prompt and builds the initial KV cache.
After that, decoding begins. For one generated token, the simplified sequence is:
- Step 1Load the embedding information for the newest token and construct its initial representation.
- Step 2Load decoder layer 1 and process the new token using the keys and values already cached for earlier tokens. Store the new token's key and value for that layer, then release the layer's weights.
- Step 3Repeat the same operation through decoder layers 2 to 40, loading and releasing each layer's weights as execution progresses.
- Step 4Load the output head, score the possible next tokens, and select one. With greedy decoding, the highest-scoring token is selected.
That sequence produces one new token. The newly selected token is then fed back into the same process.
The important change is in weight residency. Instead of keeping all 29.5 GB of model weights on the GPU, streamed execution keeps only the model part currently being used, together with the KV cache, runtime allocations, and temporary working memory.
For Qwen3-14B, one decoder layer requires about 0.66 GB, while the output head requires about 1.6 GB.
The weight-residency requirement is therefore determined mainly by the model part currently loaded rather than by the sum of all model weights.
Measured streamed execution
I measured two systems that load weights incrementally: AirLLM 3.2.0 and Afterimage exact minimum-memory [6].
Both used the same Qwen3-14B BF16 checkpoint on my RTX 3080 Laptop GPU. The test used 4 prompts and 4 generated tokens per prompt with greedy decoding. Both methods produced the same output tokens.
Table 6. Minimum-memory measurements for Qwen3-14B BF16 on the RTX 3080 Laptop GPU. Source: Afterimage results [6].
| Method | Measured PyTorch peak | Request time per generated token |
|---|---|---|
| AirLLM 3.2.0 | 1.6 GB | 27 s |
| Afterimage exact minimum-memory | 1.7 GB | 33 s |
Peak GPU memory is PyTorch's measured whole-run peak. It can include activations and other PyTorch working allocations as well as weights, and it does not include every source of GPU memory use.
Request time per generated token is the total request wall time divided by the four generated tokens. It includes prompt prefill and is therefore not an isolated steady-state decode measurement.
These results solve the capacity problem, but not the performance problem.
Both systems run the 29.5 GB BF16 model with measured PyTorch peaks below 2 GB in this configuration. The measured requests average 27 to 33 seconds of wall time per generated token.
The model is no longer prevented from running because all of its weights cannot fit on the GPU. Instead, execution depends on repeatedly retrieving, preparing, transferring, and executing the weights required by each stage.
AirLLM is faster at this operating point. Afterimage exact minimum-memory also reconstructs its losslessly stored weights before execution, adding work that is examined later in this series.
Loading every layer from storage minimizes GPU-memory use, but it treats every weight as equally disposable. In practice, some weights may be worth keeping on the GPU, some may be worth keeping in system memory, and others may be cheap enough to reload.
This starts to look like an optimization problem.
Given fixed GPU and system-memory budgets, the question becomes which weights should remain in each memory tier and which should be retrieved again during execution.
The best choice may not be determined by weight size alone. It depends on when the weight is needed, how expensive it is to prepare and move, and whether that work delays generation or can overlap with something already running.
Later parts of this series will examine that problem.
Conclusions
The GPU memory needed to run Qwen3-14B was calculated from its configuration and compared with my 8 GB laptop GPU. The following were concluded:
- GPU memory holds five main items: system use, GPU software, weights, the KV cache, and activations. On my laptop, the system alone used 2.2 GB before any model was loaded, leaving about 6.4 GB of physical capacity.
- The released BF16 weights need 29.5 GB, more than 3 times the GPU memory. Most of this is in 40 decoder layers of about 0.66 GB each.
- The KV cache stores keys and values for the input prompt and generated answer. It needs about 0.16 MB per token and about 5.4 GB at 32K tokens.
- The model uses its major weight groups sequentially. Loading each group only when it is needed lowers the measured GPU peak to under 2 GB in the tested AirLLM and Afterimage configurations.
Therefore, running the released BF16 model on this GPU requires most weights to remain outside GPU memory and be delivered as execution progresses.
The memory problem has become an optimization problem: which weights should remain on the GPU, which should remain in system memory, and which should be retrieved again?
The next part examines the execution path of streamed inference to determine which operations lie on the critical path to the next token and which work can be overlapped or avoided.
References
- Qwen Team. 2025. “Qwen3: Think Deeper, Act Faster.” April 29, 2025. Qwen3 Technical Report. Model card and configuration: Qwen/Qwen3-14B.
- Number formats and quantization: D. Kalamkar et al. 2019, “A Study of BFLOAT16 for Deep Learning Training”; T. Dettmers et al. 2022, “LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale,” NeurIPS 2022; T. Dettmers and L. Zettlemoyer. 2023, “The Case for 4-bit Precision: k-bit Inference Scaling Laws,” ICML 2023.
- A. Vaswani et al. 2017. “Attention Is All You Need.” NeurIPS 2017.
- J. Ainslie et al. 2023. “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.” EMNLP 2023.
- Hardware specifications: GeForce RTX 3080 Mobile, 8 GB GDDR6, 448 GB/s memory bandwidth; PCIe 4.0 x16, about 32 GB/s per direction.
- AirLLM 3.2.0. O. Erge, Afterimage; measured tools and benchmark details in
docs/HOW_IT_WORKS.md.