View all articles
newsopinionno-hypeai

No Hype: Beam, Watermarking, and Local Optimisation

JG
Jacobo González Jaspe
|

Reviewed:

No Hype: Beam, Watermarking, and Local Optimisation No Hype
This article is also available in Spanish:Sin humo: Modelos abiertos, regulación y hardware doméstico

A new open-weight rival from Reflection AI, OpenAI watermarking text for the EU, a 125B model squeezed onto a 12 GB gaming GPU, and Google trimming what cheaper Gemini plans get: what changed, and what it means if you run AI yourself.

Reflection AI unveils Beam model

What happened. Reflection AI announced Beam, a text-only mixture-of-experts model, on Monday (techcrunch.com). The 501-billion-parameter model has 23 billion active parameters and a 1 million token context window. The company has raised approximately $4.7 billion and reached a $25 billion valuation. Weights are planned for release this month.

What we think. Another massive model enters the arena with a massive valuation to match. The $7 billion chip deal suggests they are prepared for the compute costs.

The good bit. The release of open weights provides a high-performance alternative for those avoiding closed systems.

What you can do. Use our VRAM calculator to see if your setup can handle large models.

OpenAI implements text watermarking in the EU

Illustration generated with AI on our own machine. Illustration generated with AI on our own machine.

What happened. OpenAI is adding an optional text watermark API for users in the European Union (cryptobriefing.com). This follows the June 2026 signing of the EU Code of Practice on Transparency of AI-Generated Content. The company aims to comply with Article 50 of the AI Act. The deadline for existing systems is December 2, 2026.

What we think. OpenAI is adding features to meet EU AI Act rules. An internal prototype was reported to be 99.9% effective in 2024.

The good bit. It provides a standardised way to identify AI-generated text within the European regulatory framework.

What you can do. Read our EU AI Act guide to understand how these rules affect your projects.

Running 125B parameters on a 12GB GPU

Illustration generated with AI on our own machine. Illustration generated with AI on our own machine.

What happened. A project by Strata demonstrated running the 125 billion parameter Qwen3.8-Flash-Next model on a gaming PC (hwupgrade.it, 2026-10-05). The setup requires 12 GB of VRAM, 32 GB of system RAM, and 80 GB of free disk space. On a GeForce RTX 5070, Q2_0 quantization generates 94 tokens/s. The system loads between 35 and 55 GB into memory during the first startup.

What we think. It is impressive that a 70 GB download can run on consumer hardware. However, the 32 GB system RAM requirement remains a hurdle for many.

The good bit. Speculative decoding provides a gain of 1.6-1.8 times, making large models more usable locally.

What you can do. Explore local AI for developers to learn more about running models on your own hardware.

Google restricts Gemini access for budget users

Illustration generated with AI on our own machine. Illustration generated with AI on our own machine.

What happened. From 9 October, Google limits which models non-paying users and AI Plus subscribers can use (androidauthority.com). Users without a Google AI subscription get Gemini 3.5-Flash-Lite only. AI Plus, which costs $4.99 a month, is restricted to Flash and Flash-Lite models and loses Gemini 3.1-Pro and the newer reasoning models. Gemini AI Pro and Ultra are not affected.

What we think. A cheap tier that quietly gets cheaper to run is still a cheap tier, just with less in it. Fair enough as a business decision; less fair if you built a workflow on Gemini 3.1-Pro last month on the AI Plus plan and find out from a blog post.

The good bit. Pro and Ultra subscribers keep full access and also get a Deep Thinking option, so the paid tiers are at least clearly labelled.

What you can do. If anything you rely on runs on Gemini 3.1-Pro through a free or AI Plus account, check it before 9 October. For work that should not change when a vendor changes its tiers, see what hardware runs which model.

In one line

Big tech is tightening access and compliance while local optimisation makes massive models more accessible.

Work with us

Contact us for more information or to discuss our services. /en/contact /en/consulting

Diagram
Share: LinkedIn X
Veredicto semanal

Get new guides before anyone else

Subscribe and we tell you when new guides, templates and workflows go up. One email a week, no spam.

Already published: 69 guides and 25 templates. All free, no signup.

Bonus: the local-AI starter pack PDF when you subscribe
Once a week No spam Unsubscribe anytime

See what you get

The EU AI Act now applies: a checklist you can complete

Tell us what you want to run

Tell us what you want to run and on what budget. We will tell you which hardware you need, which model fits, and what to expect from it, before you spend anything.

First call free, 15 min Local-first: your data stays on your network Open tools and guides

69 free guides · 17 compliance templates