Does the model fit in memory?
The same sum our product pages use: weights, KV cache and runtime overhead against usable memory.
Write memoryNeededGb(params, quant, kvMbPerToken, contextTokens), which returns the GB needed:
- weights: params (billions) × BITS_PER_WEIGHT[quant] / 8;
- KV cache: kvMbPerToken × contextTokens / 1000;
- plus RUNTIME_OVERHEAD_GB (1.5 GB).
Also write fitStatus(neededGb, usableGb): return 'fits' when the need is <= 85 % of usable memory, 'tight' when it fits but goes past 85 %, and 'no' when it does not fit. Do not round anything: the results must match the site's.
Challenges 0/4
- Llama 3.1 8B at Q4 with 8192 tokens needs 7.42 GB
- Llama 3.1 70B at Q8 with 32,768 tokens: 87.26 GB
- fits, tight and no verdicts around the 85 % line
- On a DGX Spark (122 GB usable) the 70B does not fit in fp16 but fits in Q8
function memoryNeededGb(params, quant, kvMbPerToken, contextTokens) {
const weights = (params * BITS_PER_WEIGHT[quant]) / 8;
// add the KV cache and the runtime overhead
return weights;
}
function fitStatus(neededGb, usableGb) {
// 'fits' up to 85 % of usableGb, 'tight' up to 100 %, otherwise 'no'
return neededGb <= usableGb ? 'fits' : 'no';
}
// Llama 3.1 8B at Q4 with 8192 tokens of context (0.131 MB of KV per token)
console.log(memoryNeededGb(8.0, 'q4', 0.131, 8192));Console output appears here (console.log).
Go deeper: the related tool →
This in production, with your data? Let's talk for 15 minutes →