Updated 4 October 2026. This post is from April 2026 and names Qwen 2.5 as current. The method and the advice still hold; today you would run it with:
- Gemma 4: Use the small Gemma 4 models for the same fine-tuning workflows on Mac or GPU hardware.
- DeepSeek V4: The medium-sized DeepSeek V4 models can be fine-tuned using the same LoRA and QLoRA techniques described here.
Fine-tuning sounds like a task for a GPU cluster and a team of machine learning specialists. For a 7B model, you can use a Mac Mini M4 or a single consumer GPU. LoRA and QLoRA compress the training process so much that a model trained on your company data runs on the same hardware you use for inference.
This guide explains how it works, what it costs, and when it makes sense for your business. It is simpler than it looks.

What Fine-Tuning Actually Does
A pre-trained model like Qwen 2.5 7B or Gemma 3 4B knows a lot about everything. Fine-tuning teaches it to be exceptional at your specific task.
flowchart LR
BASE["Base Model<br/>(General Knowledge)"] --> LORA["LoRA Training<br/>(Your Data)"]
LORA --> CUSTOM["Custom Model<br/>(Your Domain Expert)"]
DATA["Your Training Data<br/>(500-5,000 examples)"] --> LORA
style BASE fill:#1E293B,color:#FAFAFA
style LORA fill:#F5A623,color:#0B1628
style CUSTOM fill:#059669,color:#FAFAFABefore fine-tuning: “Summarize this contract” → generic legal summary After fine-tuning on your firm’s contracts: “Summarize this contract” → summary in your firm’s format, highlighting the clauses your lawyers care about, using your terminology
LoRA vs QLoRA: The Techniques That Changed Everything
Traditional fine-tuning updates every parameter in the model, for a 7B model, that’s 7 billion numbers. With standard mixed-precision training and the Adam optimizer, weights, gradients and optimizer state take roughly 16 bytes per parameter, so well over 100 GB before activations. Impractical on consumer hardware.
LoRA (Low-Rank Adaptation) freezes the original model and trains only small “adapter” matrices. The LoRA paper reports that on GPT-3 it cut trainable parameters by 10,000 times and GPU memory by 3 times, while matching or beating full fine-tuning quality on the models the authors tested. A 7B model’s LoRA adapter is typically tens of megabytes, depending on the rank you choose, instead of 14 GB.
QLoRA goes further by quantizing the frozen base model to 4-bit precision during training, a quarter of its 16-bit size. The QLoRA paper reports fine-tuning a 65B model on a single 48 GB GPU while preserving full 16-bit fine-tuning task performance on its benchmarks.
| Method | What is trained | Base model in memory | Best For |
|---|---|---|---|
| Full fine-tune | Every parameter | 16-bit weights + gradients + optimizer state | Research, large budgets |
| LoRA | Small adapter matrices | 16-bit weights, frozen | Best quality on a large GPU or Mac |
| QLoRA | Small adapter matrices | 4-bit weights, frozen | Consumer hardware sweet spot |
Those papers measured quality on their own benchmarks. Whether an adapter is good enough for your task is something you check on your own held-out examples, with the “Test your fine-tuned model” command in the MLX walkthrough below.
What You Need: Hardware Requirements
| Your Hardware | Practical Model Size | Best Tool |
|---|---|---|
| Mac Mini M4 16GB | 7B (QLoRA) | MLX-LM (LORA.md) |
| Mac with 32GB | 7B (LoRA) or 14B (QLoRA) | MLX-LM |
| RTX 3080 10GB | 7B (QLoRA) | Unsloth |
| RTX 3090 / 4090 24GB | 13B-14B (QLoRA) | Unsloth |
Training time depends on the number of examples, their length and the number of iterations, so we do not quote a generic figure. Run 50 iterations first, read the time per iteration from the log, and multiply.
Step-by-Step: Fine-Tune on Mac with MLX
Apple’s MLX framework makes fine-tuning native on Apple Silicon:
# Install MLX-LM
pip install mlx-lm
# Prepare your training data: --data takes a DIRECTORY holding train.jsonl
# (valid.jsonl is optional and reports validation loss)
mkdir -p data
cat > data/train.jsonl << 'EOF'
{"prompt": "Summarize this contract clause:", "completion": "This clause establishes..."}
{"prompt": "Extract the payment terms:", "completion": "Payment is due within..."}
EOF
# Fine-tune with LoRA (the adapter is saved to --adapter-path)
mlx_lm.lora \
--model mlx-community/Qwen2.5-7B-Instruct-4bit \
--train \
--data ./data \
--batch-size 2 \
--num-layers 16 \
--iters 500 \
--adapter-path ./my-custom-adapter
# Test your fine-tuned model
mlx_lm.generate \
--model mlx-community/Qwen2.5-7B-Instruct-4bit \
--adapter-path ./my-custom-adapter \
--prompt "Summarize this contract clause: ..."
# Optional: fuse the adapter into a standalone model
mlx_lm.fuse \
--model mlx-community/Qwen2.5-7B-Instruct-4bit \
--adapter-path ./my-custom-adapter \
--save-path ./my-fused-modelThe adapter (adapters.safetensors plus adapter_config.json in the adapter folder) is typically tens of megabytes. The base model stays unchanged. You can swap adapters for different tasks without downloading new models.
Step-by-Step: Fine-Tune on NVIDIA GPU with Unsloth
For GPU-accelerated training on Windows or Linux hardware (Unsloth docs):
# Install Unsloth (fastest QLoRA library)
pip install unsloth
# Python training script
python << 'EOF'
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="unsloth/Qwen2.5-7B-Instruct-bnb-4bit",
max_seq_length=2048,
load_in_4bit=True,
)
model = FastLanguageModel.get_peft_model(model, r=16, lora_alpha=16)
# Train on your data
from trl import SFTTrainer
trainer = SFTTrainer(model=model, tokenizer=tokenizer, dataset=your_dataset)
trainer.train()
# Export to GGUF for Ollama
model.save_pretrained_gguf("./output", tokenizer, quantization_method="q4_k_m")
EOFThe final step, exporting to GGUF, means your fine-tuned model runs directly in Ollama. Same deployment, same infrastructure, just better at your specific task.
When Fine-Tuning Makes Sense (And When It Doesn’t)
| Scenario | Fine-Tune? | Why |
|---|---|---|
| ”Answer questions about our product catalog” | No: use RAG | RAG retrieves current data; fine-tuning bakes in stale data |
| ”Write emails in our brand voice” | Yes | Style and tone are learned through examples |
| ”Classify support tickets into 12 categories” | Yes | Domain-specific classification is a good fit for examples |
| ”Extract structured data from our invoice format” | Yes | Consistent extraction patterns are trainable |
| ”Summarize contracts in our template format” | Yes | Output format is a fine-tuning strength |
| ”Answer general questions” | No | Base models already handle this well |
Rule of thumb: Fine-tune when the format or style of the output matters. Use RAG when the data needs to be current.
The Economics
| Cost Item | Local Fine-Tuning | Cloud Fine-Tuning Service |
|---|---|---|
| Hardware | The machine you already run inference on | N/A |
| Training compute | Electricity for the run | Billed per training token (see each provider’s pricing page) |
| Per-inference cost | None beyond electricity | Per-token price of the fine-tuned model |
| Data privacy | 100% local | Data sent to provider |
| Iterations | As many as you have time for | Each run is billed |
The ability to iterate freely is the hidden advantage. With cloud training, every experiment costs money. With local training, you can run many experiments in a day at no marginal cost beyond electricity, finding the right dataset and parameters for your use case.
What We Offer
At VORLUX AI, fine-tuning is available as an add-on to our Edge AI deployment:
- Data preparation: We help structure your training examples (typically 500-5,000 samples)
- Model selection: Choose the right base model for your task and hardware
- Training: LoRA/QLoRA fine-tuning on our hardware or yours
- Evaluation: Test the fine-tuned model against your quality criteria
- Deployment: Export to Ollama and integrate with your existing workflows
The fine-tuned model runs on the same Mac Mini as your base deployment. No additional hardware needed.
Want a model that speaks your business language? Schedule a free 15-minute assessment to discuss whether fine-tuning makes sense for your use case.
Related: Quantization Guide | Best Local LLMs | Hardware Guide | n8n RAG Pipeline
Sources: LoRA paper (Hu et al., 2021) | QLoRA paper (Dettmers et al., 2023) | MLX-LM (LORA.md) | Unsloth docs | LoRA on Apple Silicon (Towards Data Science)
Next steps
- Check your current hardware against our requirements for Mac Mini or NVIDIA GPU setups.
- Evaluate if your specific task requires fine-tuning or if RAG is a better fit.
- Review our hardware guide to see if your existing machines can handle the training load.
- Compare the efficiency of different model sizes like Gemma 3 4B for your specific use case.
Related reading
- Edge AI Hardware Guide 2026: Jetson vs Mac Mini vs NUC, Real Specs, Real Costs
- Quantization Explained: Run 70B AI Models on Consumer Hardware
- NPU vs GPU: Why Neural Processing Units Are the Future of Edge AI
Work with us
We size the model and the machine by measuring, not by guessing. If you want to see your own task running on real hardware, book a 15-minute call or see how we work in consulting.