Skip to project details
Oney Erge
Inference systemsResearch system

Afterimage

Run large, full-precision models on smaller GPUs.

A research toolbox for testing lossless compression, streaming, and speculative methods beyond VRAM. A measured 29.5 GB Qwen3-14B BF16 model ran on an 8 GB GPU, with some tested modes outperforming AirLLM and Hugging Face Accelerate.

Model layers streaming from storage through memory to a smaller GPU
29.5 GB modelon8 GB GPU
Lossless weight streaming

What Afterimage does.

Aim

Run a model larger than GPU memory without reducing the precision of its stored weights by loading only the layers needed at each moment.

What it does

Lossless weight storage and layerwise CUDA streaming executed a 29.536 GB BF16 model on an 8 GB GPU. The tradeoff is disk-bound latency measured in seconds per token.

Good at

  • 29.5 GB model on 8 GB GPU
  • Lossless weights

The Afterimage flow.

  1. 01

    Inspect

    Estimate download, store, host memory, and VRAM requirements.

  2. 02

    Compress

    Build a lossless on-disk store from the original model weights.

  3. 03

    Plan

    Choose exact residency and optional speculative decoding settings.

  4. 04

    Stream

    Move each required layer through the available GPU memory.

  5. 05

    Measure

    Record latency, peak memory, exactness, and named baselines.

Run the estimator for your model and hardware, then compare a supported exact mode with a named baseline.

Setup and examples

Explore another project.

View all projects