KOLDOS by graphics card
Results measured by the KOLDOS team on real hardware, next to what the memory estimate says for every other card. A configuration that has not been measured shows Benchmark unavailable.
- Data source
- Bundled file
- Measured results
- …
GPU table
Vendor
Form factor
| Compare | GPU | Architecture | VRAM | Tier | 7B · 4-bit estimate | 22B · 4-bit estimate | Benchmark status |
|---|---|---|---|---|---|---|---|
| RTX 5050NVIDIA | Blackwell | 8 GB | Mainstream | Fits | Needs RAM | … | |
| RTX 5060NVIDIA | Blackwell | 8 GB | Mainstream | Fits | Needs RAM | … | |
| RTX 5070NVIDIA | Blackwell | 12 GB | Upper | Fits | Needs RAM | … | |
| RTX 5080NVIDIA | Blackwell | 16 GB | High | Fits | Fits | … | |
| RTX 5090NVIDIA | Blackwell | 32 GB | Enthusiast | Fits | Fits | … | |
| RTX 4060NVIDIA | Ada Lovelace | 8 GB | Mainstream | Fits | Needs RAM | … | |
| RTX 4070NVIDIA | Ada Lovelace | 12 GB | Upper | Fits | Needs RAM | … | |
| RTX 4070 TiNVIDIA | Ada Lovelace | 12 GB | Upper | Fits | Needs RAM | … | |
| RTX 4080NVIDIA | Ada Lovelace | 16 GB | High | Fits | Fits | … | |
| RTX 4090NVIDIA | Ada Lovelace | 24 GB | Enthusiast | Fits | Fits | … | |
| RTX 4050 LaptopNVIDIA | Ada Lovelace | 6 GB | Entry | Fits | Needs RAM | … | |
| RTX 4060 LaptopNVIDIA | Ada Lovelace | 8 GB | Mainstream | Fits | Needs RAM | … | |
| RTX 4070 LaptopNVIDIA | Ada Lovelace | 8 GB | Mainstream | Fits | Needs RAM | … | |
| RTX 3060 12 GBNVIDIA | Ampere | 12 GB | Upper | Fits | Needs RAM | … | |
| RX 7900 XTXAMD | RDNA 3 | 24 GB | Enthusiast | Fits | Fits | … |
Estimates assume a 4K context and count only GPU memory. "Needs RAM" means part of the model would be offloaded to system RAM. Tiers group cards by VRAM: Entry (Up to 6 GB), Mainstream (7 to 8 GB), Upper (9 to 12 GB), High (13 to 16 GB), Enthusiast (More than 16 GB). Select up to 4 GPUs to compare them below.
Side by side
3 of 4 selected
| Configuration | RTX 4050 Laptop | RTX 4060 | RTX 4090 |
|---|---|---|---|
| VRAM | 6 GB | 8 GB | 24 GB |
| Architecture | Ada Lovelace | Ada Lovelace | Ada Lovelace |
| Series | GeForce RTX 40 | GeForce RTX 40 | GeForce RTX 40 |
| Form factor | Laptop | Desktop | Desktop |
| 7B at 4-bit | Fits | Fits | Fits |
| 7B at 8-bit | Needs RAM | Needs RAM | Fits |
| 7B context kept in VRAM | Up to 32K | Up to 64K | Up to 128K |
| 7B measured speed | Benchmark unavailable | Benchmark unavailable | Benchmark unavailable |
| 22B at 4-bit | Needs RAM | Needs RAM | Fits |
| 22B at 8-bit | Needs RAM | Needs RAM | Tight |
| 22B context kept in VRAM | — | — | Up to 32K |
| 22B measured speed | Benchmark unavailable | Benchmark unavailable | Benchmark unavailable |
Measured results
Benchmark runs
Every run records the full configuration, so results can be compared and reproduced.
Loading benchmark data
What every result records
A tokens-per-second figure means little without the configuration behind it.
- GPU catalog IDstring | null
- ID from the hardware database, or null for unlisted hardware.
- GPUstring
- Exact GPU model, including Laptop and memory variant.
- VRAMnumber | null
- Dedicated GPU memory in GB. Null for unified memory.
- RAMnumber | null
- Installed system memory in GB. Null when it was not recorded.
- CPUstring
- CPU model as reported by the operating system.
- Operating system{ id, version }
- windows, linux or macos, plus the exact version when it was recorded.
- Modelstring
- ID of the KOLDOS model that was measured.
- Quantizationstring
- Quantization name exactly as published.
- Contextnumber
- Context length configured for the run, in tokens.
- Backend{ name, version }
- Inference backend and its version.
- GPU layers"all" | number
- How many layers were offloaded to the GPU.
- Tokens/snumber | null
- End-to-end throughput: all tokens divided by total time.
- Prompt tokens/snumber | null
- Prompt processing speed.
- Generation tokens/snumber | null
- Output generation speed.
- KOLDOS versionstring | null
- Version of KOLDOS used for the run. Null for runs made before versions were recorded.
- MeasuredISO 8601 date
- When the run was recorded.
- Sourcekoldos-team | community
- Who ran the benchmark.
- Verifiedboolean
- Whether the KOLDOS team reproduced the result.