NVIDIA (custom build) · GPU PC
RTX 4090 24GB Build
Maximum speed, 30B models at 130 tok/s
- 24 GBGPU memory (VRAM)
- 14Bfits in memory (Q4, 8k)
- 14Bpractical (estimate)
- 600 Wmax draw
- ~€2,500indicative price
What can it run?
Pick a job or a model size. The bar compares what it needs with this machine's usable memory.
- Model weights
- Context cache
- Runtime + extras (Whisper, embeddings)
How we calculate this
Weights = parameters × bits per weight / 8 (Q4_K_M ≈ 4.85; Q8_0 ≈ 8.5). Cache = context tokens × the reference model's per-token size (layers × KV heads × head size × 2 × 2 bytes). Plus 1.5 GB of runtime. Usable memory = total minus what the system keeps (3–6 GB on shared memory, 0.5 GB on a GPU). Comfortable = fits in 85%. It is a ±20% estimate: measure it on your machine before you decide. How to pick the GGUF file
When does it beat the cloud?
Set your usage. We compare the purchase plus electricity with what you would pay a cloud API.
| Month | Own hardware | Cloud |
|---|
- Own hardware (purchase + electricity)
- Cloud API
Assumptions: €0.25/kWh (Spain average, editable), USD 1 = €0.92, 3 input tokens per output token, and the machine at maximum draw for every hour it is on (worst case). API prices verified 2026-09-09. Your time and maintenance are not included. Note: the cloud side is a frontier model and here you would run a smaller open one; this compares cost, not quality.
Learn with this machine
- ComfyUI Batch Image Generation: Create 100 Product Images in MinutesTutorial on using ComfyUI's batch generation workflow to create consistent product images, social media graphics, and marketing assets at scale — all.
- CUDA on Fedora: a Local AI Workstation in an AfternoonInstall the NVIDIA driver from RPM Fusion, get Ollama on the GPU in five minutes, then build llama.cpp with the CUDA toolkit using the official Fedora guide.
- Llama 3.3 70B on Your Own Hardware: What It Takes to Run ItThe memory a 70B model really needs, measured and cited speeds on a 64 GB Mac, two 24 GB GPUs and our GB10, and the SME tasks where it beats an 8B model.
- GGUF Quantization: Pick the Right File for Your MachineHow to read a GGUF file name, choose between Q4_K_M, Q5 and Q8 for the memory you have, and measure speed, memory and quality yourself with three commands.
Skip it if…
- it will run all day in an office: it draws up to 600 W and a desktop GPU is not quiet.
- you need 70B models: they do not fit in this much GPU memory.
Specs and where the numbers come from
| CPU | AMD Ryzen 7 / Intel i7 |
|---|---|
| GPU | RTX 4090 24GB |
| NPU | N/A (CUDA cores) |
| Memory | 24 GB (GPU memory (VRAM)); usable by the model ≈ 23.5 GB |
| Speed | 130 tok/s with 30B vendor or community estimate, not measured by us |
| Price | ~€2,500 indicative, checked 2026-09-19; check the live price in the shop |
We only call something "measured" when it ran on our machines. Editorial policy