Skip to content

KOLDOS by graphics card

Results measured by the KOLDOS team on real hardware, next to what the memory estimate says for every other card. A configuration that has not been measured shows Benchmark unavailable.
Data source
Bundled file
Measured results
…

GPU table

Vendor

Form factor

GPUs with architecture, VRAM tier, estimated fit at 4-bit and benchmark status. 15 rows.
CompareGPUArchitectureVRAMTier7B · 4-bit estimate22B · 4-bit estimateBenchmark status
RTX 5050NVIDIABlackwell8 GBMainstreamFitsNeeds RAM…
RTX 5060NVIDIABlackwell8 GBMainstreamFitsNeeds RAM…
RTX 5070NVIDIABlackwell12 GBUpperFitsNeeds RAM…
RTX 5080NVIDIABlackwell16 GBHighFitsFits…
RTX 5090NVIDIABlackwell32 GBEnthusiastFitsFits…
RTX 4060NVIDIAAda Lovelace8 GBMainstreamFitsNeeds RAM…
RTX 4070NVIDIAAda Lovelace12 GBUpperFitsNeeds RAM…
RTX 4070 TiNVIDIAAda Lovelace12 GBUpperFitsNeeds RAM…
RTX 4080NVIDIAAda Lovelace16 GBHighFitsFits…
RTX 4090NVIDIAAda Lovelace24 GBEnthusiastFitsFits…
RTX 4050 LaptopNVIDIAAda Lovelace6 GBEntryFitsNeeds RAM…
RTX 4060 LaptopNVIDIAAda Lovelace8 GBMainstreamFitsNeeds RAM…
RTX 4070 LaptopNVIDIAAda Lovelace8 GBMainstreamFitsNeeds RAM…
RTX 3060 12 GBNVIDIAAmpere12 GBUpperFitsNeeds RAM…
RX 7900 XTXAMDRDNA 324 GBEnthusiastFitsFits…

Estimates assume a 4K context and count only GPU memory. "Needs RAM" means part of the model would be offloaded to system RAM. Tiers group cards by VRAM: Entry (Up to 6 GB), Mainstream (7 to 8 GB), Upper (9 to 12 GB), High (13 to 16 GB), Enthusiast (More than 16 GB). Select up to 4 GPUs to compare them below.

Side by side

3 of 4 selected

ConfigurationRTX 4050 LaptopRTX 4060RTX 4090
VRAM6 GB8 GB24 GB
ArchitectureAda LovelaceAda LovelaceAda Lovelace
SeriesGeForce RTX 40GeForce RTX 40GeForce RTX 40
Form factorLaptopDesktopDesktop
7B at 4-bitFitsFitsFits
7B at 8-bitNeeds RAMNeeds RAMFits
7B context kept in VRAMUp to 32KUp to 64KUp to 128K
7B measured speedBenchmark unavailableBenchmark unavailableBenchmark unavailable
22B at 4-bitNeeds RAMNeeds RAMFits
22B at 8-bitNeeds RAMNeeds RAMTight
22B context kept in VRAM——Up to 32K
22B measured speedBenchmark unavailableBenchmark unavailableBenchmark unavailable

Measured results

Benchmark runs

Every run records the full configuration, so results can be compared and reproduced.
Loading benchmark data

What every result records

A tokens-per-second figure means little without the configuration behind it.

GPU catalog IDstring | null
ID from the hardware database, or null for unlisted hardware.
GPUstring
Exact GPU model, including Laptop and memory variant.
VRAMnumber | null
Dedicated GPU memory in GB. Null for unified memory.
RAMnumber | null
Installed system memory in GB. Null when it was not recorded.
CPUstring
CPU model as reported by the operating system.
Operating system{ id, version }
windows, linux or macos, plus the exact version when it was recorded.
Modelstring
ID of the KOLDOS model that was measured.
Quantizationstring
Quantization name exactly as published.
Contextnumber
Context length configured for the run, in tokens.
Backend{ name, version }
Inference backend and its version.
GPU layers"all" | number
How many layers were offloaded to the GPU.
Tokens/snumber | null
End-to-end throughput: all tokens divided by total time.
Prompt tokens/snumber | null
Prompt processing speed.
Generation tokens/snumber | null
Output generation speed.
KOLDOS versionstring | null
Version of KOLDOS used for the run. Null for runs made before versions were recorded.
MeasuredISO 8601 date
When the run was recorded.
Sourcekoldos-team | community
Who ran the benchmark.
Verifiedboolean
Whether the KOLDOS team reproduced the result.