Skip to content

Will KOLDOS run on your computer?

Enter your graphics card, memory and context length. The checker estimates how much memory each model needs and where it would run: GPU inference, GPU + RAM offload, or CPU inference. It estimates memory, not speed. How the estimate works

Your system

Fill in from this browser

Reads the operating system, CPU threads and GPU name from this browser. Browsers can't read VRAM, so it comes from the GPU's published spec when the model is recognized. Nothing is sent anywhere.

Operating system

Dedicated VRAM

From the published spec of the RTX 4060.

System RAM

Context length

How much text the model can read at once. More context needs more memory.

Free storage

CPU threads

Only used to label shared results. The estimate is based on memory.

Result

Compatibility: Good. 7B quantized, up to 6-bit.

RTX 4060 · 8 GB VRAM · 16 GB RAM · 4K context

KOLDOS compatibility
Good
Expected model class
7B quantized, up to 6-bit
Expected experience
Local inference suitable for smaller quantized models. KOLDOS 7B fits entirely in GPU memory. KOLDOS 22B would need GPU + RAM offload, which is much slower.

Why this result

  • Your GPU has 8 GB of VRAM. After keeping 0.5 GB free for the desktop and other programs, 7.5 GB is usable.
  • KOLDOS 7B at 6-bit with a 4K context needs about 6.1 GB: 5.3 GB of weights, 0.1 GB of context cache and 0.6 GB of runtime overhead.
  • That fits within the available GPU memory with 1.4 GB to spare, so the whole model can run on the GPU.

Performance estimate unavailable

Benchmark database

The checker estimates memory, not speed. Tokens per second depend on memory bandwidth, the inference backend, drivers, power limits and settings, so no figure is shown until it has been measured on real hardware.

  • Windows support: In development.

Fit by model and quantization

Estimated fit for each model and quantization. Select a cell to see its memory breakdown.
Model3-bit4-bit5-bit6-bit8-bit16-bit
KOLDOS 7BAvailable
KOLDOS 22BFuture
GPU
The whole model and its context fit in GPU memory with room to spare.
Tight
It fits in GPU memory, with little room left for other programs or a longer context.
Offload
Part of the model is offloaded to system RAM. It works, but noticeably slower.
CPU
The GPU can't hold this model, so it runs on the CPU from system RAM, which is slow.
No
Not enough free VRAM and RAM combined for this model at this quantization.

Memory breakdown

KOLDOS 7B at 6-bit (Q6_K-style), 4K context

GPU inference

GPU memory

  • Weights
  • Context (KV cache)
  • Runtime overhead
  • Beyond VRAM
Weights
5.3 GB
Context
0.1 GB
Overhead
0.6 GB
Total
6.1 GB
Usable VRAM
7.5 GB
Usable RAM
12.8 GB
Share of the model on the GPU
100%
Model file size
5.3 GB

Suggested starting point

Model
KOLDOS 7B
Quantization
6-bit
Context that stays in VRAM
Up to 32K
GPU offload
All layers

Share or compare

Copy a link with these inputs, or compare common GPUs side by side in the benchmark database.

Compare GPUs