Will KOLDOS run on your computer?
Your system
Fill in from this browser
Reads the operating system, CPU threads and GPU name from this browser. Browsers can't read VRAM, so it comes from the GPU's published spec when the model is recognized. Nothing is sent anywhere.
Operating system
Dedicated VRAM
From the published spec of the RTX 4060.
System RAM
Context length
How much text the model can read at once. More context needs more memory.
Free storage
CPU threads
Only used to label shared results. The estimate is based on memory.
Result
RTX 4060 · 8 GB VRAM · 16 GB RAM · 4K context
- KOLDOS compatibility
- Good
- Expected model class
- 7B quantized, up to 6-bit
- Expected experience
- Local inference suitable for smaller quantized models. KOLDOS 7B fits entirely in GPU memory. KOLDOS 22B would need GPU + RAM offload, which is much slower.
Why this result
- Your GPU has 8 GB of VRAM. After keeping 0.5 GB free for the desktop and other programs, 7.5 GB is usable.
- KOLDOS 7B at 6-bit with a 4K context needs about 6.1 GB: 5.3 GB of weights, 0.1 GB of context cache and 0.6 GB of runtime overhead.
- That fits within the available GPU memory with 1.4 GB to spare, so the whole model can run on the GPU.
Performance estimate unavailable
Benchmark databaseThe checker estimates memory, not speed. Tokens per second depend on memory bandwidth, the inference backend, drivers, power limits and settings, so no figure is shown until it has been measured on real hardware.
- Windows support: In development.
Fit by model and quantization
| Model | 3-bit | 4-bit | 5-bit | 6-bit | 8-bit | 16-bit |
|---|---|---|---|---|---|---|
| KOLDOS 7BAvailable | ||||||
| KOLDOS 22BFuture |
- GPU
- The whole model and its context fit in GPU memory with room to spare.
- Tight
- It fits in GPU memory, with little room left for other programs or a longer context.
- Offload
- Part of the model is offloaded to system RAM. It works, but noticeably slower.
- CPU
- The GPU can't hold this model, so it runs on the CPU from system RAM, which is slow.
- No
- Not enough free VRAM and RAM combined for this model at this quantization.
Memory breakdown
KOLDOS 7B at 6-bit (Q6_K-style), 4K context
GPU inference
GPU memory
- Weights
- Context (KV cache)
- Runtime overhead
- Beyond VRAM
- Weights
- 5.3 GB
- Context
- 0.1 GB
- Overhead
- 0.6 GB
- Total
- 6.1 GB
- Usable VRAM
- 7.5 GB
- Usable RAM
- 12.8 GB
- Share of the model on the GPU
- 100%
- Model file size
- 5.3 GB
Suggested starting point
- Model
- KOLDOS 7B
- Quantization
- 6-bit
- Context that stays in VRAM
- Up to 32K
- GPU offload
- All layers
Share or compare
Copy a link with these inputs, or compare common GPUs side by side in the benchmark database.