View all articles
fine-tuningloraedge-aitutorialhardware

Fine-Tune AI Models on Your Own Hardware: The LoRA Guide for SMEs

JG
Jacobo González Jaspe
|

Reviewed:

Abstract illustration: nested glass spheres with an amber light at the centre. Models
Illustration generated with AI on our own machine.

Updated 4 October 2026. This post is from April 2026 and names Qwen 2.5 as current. The method and the advice still hold; today you would run it with:

  • Gemma 4: Use the small Gemma 4 models for the same fine-tuning workflows on Mac or GPU hardware.
  • DeepSeek V4: The medium-sized DeepSeek V4 models can be fine-tuned using the same LoRA and QLoRA techniques described here.

Fine-tuning sounds like a task for a GPU cluster and a team of machine learning specialists. For a 7B model, you can use a Mac Mini M4 or a single consumer GPU. LoRA and QLoRA compress the training process so much that a model trained on your company data runs on the same hardware you use for inference.

This guide explains how it works, what it costs, and when it makes sense for your business. It is simpler than it looks.

Fine-tuning AI models locally

What Fine-Tuning Actually Does

A pre-trained model like Qwen 2.5 7B or Gemma 3 4B knows a lot about everything. Fine-tuning teaches it to be exceptional at your specific task.

flowchart LR
    BASE["Base Model<br/>(General Knowledge)"] --> LORA["LoRA Training<br/>(Your Data)"]
    LORA --> CUSTOM["Custom Model<br/>(Your Domain Expert)"]
    DATA["Your Training Data<br/>(500-5,000 examples)"] --> LORA
    
    style BASE fill:#1E293B,color:#FAFAFA
    style LORA fill:#F5A623,color:#0B1628
    style CUSTOM fill:#059669,color:#FAFAFA
Diagram

Before fine-tuning: “Summarize this contract” → generic legal summary After fine-tuning on your firm’s contracts: “Summarize this contract” → summary in your firm’s format, highlighting the clauses your lawyers care about, using your terminology

LoRA vs QLoRA: The Techniques That Changed Everything

Traditional fine-tuning updates every parameter in the model, for a 7B model, that’s 7 billion numbers. With standard mixed-precision training and the Adam optimizer, weights, gradients and optimizer state take roughly 16 bytes per parameter, so well over 100 GB before activations. Impractical on consumer hardware.

LoRA (Low-Rank Adaptation) freezes the original model and trains only small “adapter” matrices. The LoRA paper reports that on GPT-3 it cut trainable parameters by 10,000 times and GPU memory by 3 times, while matching or beating full fine-tuning quality on the models the authors tested. A 7B model’s LoRA adapter is typically tens of megabytes, depending on the rank you choose, instead of 14 GB.

QLoRA goes further by quantizing the frozen base model to 4-bit precision during training, a quarter of its 16-bit size. The QLoRA paper reports fine-tuning a 65B model on a single 48 GB GPU while preserving full 16-bit fine-tuning task performance on its benchmarks.

MethodWhat is trainedBase model in memoryBest For
Full fine-tuneEvery parameter16-bit weights + gradients + optimizer stateResearch, large budgets
LoRASmall adapter matrices16-bit weights, frozenBest quality on a large GPU or Mac
QLoRASmall adapter matrices4-bit weights, frozenConsumer hardware sweet spot

Those papers measured quality on their own benchmarks. Whether an adapter is good enough for your task is something you check on your own held-out examples, with the “Test your fine-tuned model” command in the MLX walkthrough below.

What You Need: Hardware Requirements

Your HardwarePractical Model SizeBest Tool
Mac Mini M4 16GB7B (QLoRA)MLX-LM (LORA.md)
Mac with 32GB7B (LoRA) or 14B (QLoRA)MLX-LM
RTX 3080 10GB7B (QLoRA)Unsloth
RTX 3090 / 4090 24GB13B-14B (QLoRA)Unsloth

Training time depends on the number of examples, their length and the number of iterations, so we do not quote a generic figure. Run 50 iterations first, read the time per iteration from the log, and multiply.

Step-by-Step: Fine-Tune on Mac with MLX

Apple’s MLX framework makes fine-tuning native on Apple Silicon:

bash
# Install MLX-LM
pip install mlx-lm

# Prepare your training data: --data takes a DIRECTORY holding train.jsonl
# (valid.jsonl is optional and reports validation loss)
mkdir -p data
cat > data/train.jsonl << 'EOF'
{"prompt": "Summarize this contract clause:", "completion": "This clause establishes..."}
{"prompt": "Extract the payment terms:", "completion": "Payment is due within..."}
EOF

# Fine-tune with LoRA (the adapter is saved to --adapter-path)
mlx_lm.lora \
  --model mlx-community/Qwen2.5-7B-Instruct-4bit \
  --train \
  --data ./data \
  --batch-size 2 \
  --num-layers 16 \
  --iters 500 \
  --adapter-path ./my-custom-adapter

# Test your fine-tuned model
mlx_lm.generate \
  --model mlx-community/Qwen2.5-7B-Instruct-4bit \
  --adapter-path ./my-custom-adapter \
  --prompt "Summarize this contract clause: ..."

# Optional: fuse the adapter into a standalone model
mlx_lm.fuse \
  --model mlx-community/Qwen2.5-7B-Instruct-4bit \
  --adapter-path ./my-custom-adapter \
  --save-path ./my-fused-model

The adapter (adapters.safetensors plus adapter_config.json in the adapter folder) is typically tens of megabytes. The base model stays unchanged. You can swap adapters for different tasks without downloading new models.

Step-by-Step: Fine-Tune on NVIDIA GPU with Unsloth

For GPU-accelerated training on Windows or Linux hardware (Unsloth docs):

bash
# Install Unsloth (fastest QLoRA library)
pip install unsloth

# Python training script
python << 'EOF'
from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="unsloth/Qwen2.5-7B-Instruct-bnb-4bit",
    max_seq_length=2048,
    load_in_4bit=True,
)

model = FastLanguageModel.get_peft_model(model, r=16, lora_alpha=16)

# Train on your data
from trl import SFTTrainer
trainer = SFTTrainer(model=model, tokenizer=tokenizer, dataset=your_dataset)
trainer.train()

# Export to GGUF for Ollama
model.save_pretrained_gguf("./output", tokenizer, quantization_method="q4_k_m")
EOF

The final step, exporting to GGUF, means your fine-tuned model runs directly in Ollama. Same deployment, same infrastructure, just better at your specific task.

When Fine-Tuning Makes Sense (And When It Doesn’t)

ScenarioFine-Tune?Why
”Answer questions about our product catalog”No: use RAGRAG retrieves current data; fine-tuning bakes in stale data
”Write emails in our brand voice”YesStyle and tone are learned through examples
”Classify support tickets into 12 categories”YesDomain-specific classification is a good fit for examples
”Extract structured data from our invoice format”YesConsistent extraction patterns are trainable
”Summarize contracts in our template format”YesOutput format is a fine-tuning strength
”Answer general questions”NoBase models already handle this well

Rule of thumb: Fine-tune when the format or style of the output matters. Use RAG when the data needs to be current.

The Economics

Cost ItemLocal Fine-TuningCloud Fine-Tuning Service
HardwareThe machine you already run inference onN/A
Training computeElectricity for the runBilled per training token (see each provider’s pricing page)
Per-inference costNone beyond electricityPer-token price of the fine-tuned model
Data privacy100% localData sent to provider
IterationsAs many as you have time forEach run is billed

The ability to iterate freely is the hidden advantage. With cloud training, every experiment costs money. With local training, you can run many experiments in a day at no marginal cost beyond electricity, finding the right dataset and parameters for your use case.

What We Offer

At VORLUX AI, fine-tuning is available as an add-on to our Edge AI deployment:

  1. Data preparation: We help structure your training examples (typically 500-5,000 samples)
  2. Model selection: Choose the right base model for your task and hardware
  3. Training: LoRA/QLoRA fine-tuning on our hardware or yours
  4. Evaluation: Test the fine-tuned model against your quality criteria
  5. Deployment: Export to Ollama and integrate with your existing workflows

The fine-tuned model runs on the same Mac Mini as your base deployment. No additional hardware needed.


Want a model that speaks your business language? Schedule a free 15-minute assessment to discuss whether fine-tuning makes sense for your use case.

Related: Quantization Guide | Best Local LLMs | Hardware Guide | n8n RAG Pipeline


Sources: LoRA paper (Hu et al., 2021) | QLoRA paper (Dettmers et al., 2023) | MLX-LM (LORA.md) | Unsloth docs | LoRA on Apple Silicon (Towards Data Science)


Next steps

  • Check your current hardware against our requirements for Mac Mini or NVIDIA GPU setups.
  • Evaluate if your specific task requires fine-tuning or if RAG is a better fit.
  • Review our hardware guide to see if your existing machines can handle the training load.
  • Compare the efficiency of different model sizes like Gemma 3 4B for your specific use case.

Work with us

We size the model and the machine by measuring, not by guessing. If you want to see your own task running on real hardware, book a 15-minute call or see how we work in consulting.

Diagram
Share: LinkedIn X
Veredicto semanal

Get new guides before anyone else

Subscribe and we tell you when new guides, templates and workflows go up. One email a week, no spam.

Already published: 69 guides and 25 templates. All free, no signup.

Bonus: the local-AI starter pack PDF when you subscribe
Once a week No spam Unsubscribe anytime

See what you get

The EU AI Act now applies: a checklist you can complete

Tell us what you want to run

Tell us what you want to run and on what budget. We will tell you which hardware you need, which model fits, and what to expect from it, before you spend anything.

First call free, 15 min Local-first: your data stays on your network Open tools and guides

69 free guides · 17 compliance templates