I am not trying to build a leaderboard in my flat. I want to know whether AI compute can move between different local machines, vendors and power envelopes while the right to make a consequential decision stays bounded, auditable and local. Two Tiiny AI Labs are already on the way. The AORUS RTX 5090 AI Box and Intel Arc Pro B70 / Dual B60 are wishlist instruments: if I can get access to them, these are the questions I would test — and publish, including negative results.
A useful experiment should end with a statement about authority, reliability, portability, cost or capability — not merely a faster token counter. The unit I care about is a verified decision: a result whose evidence, rules and permitted action can be reconstructed later.
Intel Core Ultra 9 275HX, RTX 5090 Mobile 24GB, 96GB RAM, Thunderbolt 5. This is the CUDA reference and the machine that already ran my Nemotron-FOL and Sovereign Decision Plane experiments.
Role: heavy local perception, orchestration, reference execution path.
ARM + NPU/dNPU, 80GB LPDDR5X split across SoC and dNPU, 30W TDP. The important feature is not “120B in a tiny box”; it is two separate physical machines that can fail, disagree and arrive late.
Role: physically separate proposer / reviewer / deferred-evidence nodes.
Desktop RTX 5090, 32GB GDDR7, Thunderbolt 5. Eight extra GB over the internal 5090M only matter if they unlock a class of work that 24GB cannot do cleanly.
Role: escalation / second-read node for genuinely heavier evidence.
Xe2 / XMX, 32GB GDDR6, 608GB/s spec (~500GB/s measured read in BITCOS), up to 367 INT8 TOPS. The point is not “Intel versus NVIDIA”; it is a third execution stack for testing software sovereignty and decision invariance.
Role: XPU execution-diversity check and portability test.
Two B60 GPUs, 24GB each — not one unified 48GB heap. The card needs a host that exposes both x8 links; I would not assume my M1 Lite can do that without a compatibility proof.
Role: two-worker topology or multi-GPU test, only on a compatible host.
FreeToken showed that model placement can violate the usual “VRAM = capability” intuition. Nemotron-FOL showed something more important: a confident model error can be prevented from becoming authority by a deterministic gate.
Next: move those guarantees across physical machines and backends.
Each question has a failure condition. If the result is “no”, that is still a useful public artifact. The devices do not get to define success.
A Tiiny can propose facts at T0, Olares can perform a heavier second read later, and a second Tiiny can deliver deferred evidence — while caps, TTL, revocation and ai_authorised_to_override = false remain invariant.
Why two Tiinys matter: I can physically disconnect one, delay it, replay an old packet, skew its clock or let it return FAIL after another node still holds stale PROVISIONAL state.
Falsifier: any path where delayed, stale or conflicting AI output widens authority, silently extends TTL, or authorises an action after revocation.
Run the same evidence and decision contract through CUDA, Tiiny runtime, ROCm where compatible, and later Intel XPU. I do not require identical prose or logits. I require the control plane either to reach the same permitted action or to detect material divergence before action.
Important distinction: different hardware is an independent execution path, not independent evidence. A true witness must bring separate evidence or verification; silicon diversity alone does not make a claim true.
Falsifier: backend/quantisation differences change a consequential verdict without triggering escalation. If that happens, “AI tokens” are not a fungible commodity for that task class.
Start with the smallest local node. Escalate only when a deterministic gate, missing evidence, multimodality or context requirements demand a stronger read. The metric is not tokens/second; it is the fraction of real tasks completed correctly at each rung.
Tiiny: cheap first read. Olares 5090M: heavy second read. AI Box: only if a named class of evidence cannot be handled cleanly inside 24GB.
Falsifier: the small node causes so many missed facts / false escalations that the “cheap perception” ladder costs more time or errors than simply using the 5090 reference.
I would look for named workloads that are qualitatively awkward in 24GB but clean in 32GB: less aggressive quantisation, no CPU offload, useful long context, concurrent multimodal components, image/video generation, or an escalation model that can keep all required state resident.
Why Olares is interesting: the host already has a 24GB 5090M and TB5. Olares documents NVIDIA eGPU support, plus a current Olares OS Gen1 workaround in its tested setup; Windows can renegotiate to Gen4 after restart. That makes bus behaviour part of the measurement rather than a footnote.
Falsifier: none of my real workloads gains a new reliable mode from 32GB, or the eGPU path introduces enough transfer / power / driver friction that the internal 24GB remains the practical choice.
Port the real pipeline: structured extraction, evidence packets, constrained output where supported, Prolog/Z3/OPA gates, regression cases and audit logs. B70 is useful because it provides one 32GB Intel/XMX surface; the research result is how much of the system survives a vendor/runtime change.
B70: cleanest first Intel target. Dual B60: a later topology experiment — two 24GB workers, not a magic 48GB monolith. It requires x8/x8 host support, so the host must be verified before promising a Minisforum pairing.
Falsifier: critical features or model/runtime support make the pipeline materially less reproducible, less auditable or impractical outside CUDA. That negative result is evidence about real software sovereignty.
Only after Q02 establishes which task classes are safe to move. Then schedule movable work by latency, energy, availability and tariff: Tiiny for always-on low-power work, GPU nodes for expensive escalations, with the same authority contract after them.
The unit: joules and pence per verified decision, not tokens per joule in isolation. A cheap wrong answer is not efficient.
Falsifier: energy-aware routing changes outcomes, creates unacceptable delay, or saves too little to justify the orchestration complexity.
This is deliberately not a shopping recommendation. It is a list of new falsifiable questions unlocked by access to the hardware.
| Instrument | What it adds | Question it unlocks | What would make it unnecessary |
|---|---|---|---|
| 2× Tiiny to try on Olares One, Minisforum or Morefine | Two low-power physical nodes with separate state/failure | Authority under partition; cheap perception; delayed evidence | Nothing to buy — these are already coming |
| AORUS RTX 5090 Infinity to try on Minisforum or Morefine | Nvidia RTX 5090 + 32GB over Oculink | Does +8\16GB GPU unlock a new escalation workload? | instead of my current 24\12GB AMD\Nvidia capacity there [check my re-design vision] |
| AORUS 5090 AI Box to try on Olares One | Desktop RTX 5090 + 32GB over TB5 beside a 24GB 5090M | Does +8GB / desktop GPU unlock a new escalation workload? | If Q03 produces no workload that is qualitatively constrained by 24GB |
| Arc Pro B70 32GB to try on Minisforum or Morefine | Intel Xe2/XMX and a different software/runtime path | Can the decision pipeline leave CUDA without losing semantics or auditability? | If the pipeline cannot be matched closely enough to make the comparison meaningful |
| MAXSUN Dual B60 48G to try on Minisforum or Morefine | Two independent 24GB Intel GPUs on one board | Two-worker topology / multi-GPU scaling after the worker-fungibility question matters | If a compatible x8/x8 host is unavailable, or two-worker independence adds no value |
If a £4k-ish GPU does not unlock a new capability, that is a useful result. If Intel portability is painful, publish the pain. If two Tiinys agree with themselves and still miss the truth, publish that too. A public artifact is more valuable when it can say “this idea failed here”.
The community value is not that I own unusual boxes. It is that somebody publishes the method, raw traces, failure modes and exact boundary conditions on hardware people can actually obtain.
Reproducible two/three-host test: stale state, delayed FAIL, replayed packet, clock skew, node loss, TTL expiry, revocation. One command, pass/fail invariants.
Same evidence packet and matched model/checkpoint where runtimes permit; versions, quantisation, seeds, raw outputs, extracted claims, final formal verdict. Material disagreements highlighted.
Real workload classes routed Tiiny → current 5090M → optional Box / XPU. Report escalations, misses, false escalations, latency, wall power and total cost per verified result.
Not “5090 is fast.” List the exact workloads that fail / degrade / require offload in 24GB and become clean in 32GB — plus the cost of TB5 and OS/driver caveats.
Time-to-first-result, unsupported features, model/quant gaps, structured-output behaviour, reproducibility and the patches needed. Installation pain is data when sovereignty is the claim.
Broken configurations, contradictory outputs, “not worth buying” outcomes and exact logs. No deleting a test because it makes a favourite device look bad.
Device specifications are linked to current vendor documentation; research claims remain hypotheses until the experiment is run.
AORUS RTX 5060 Ti AI BOX delivers portable desktop-class graphics with a GeForce RTX™ 5060 Ti 16GB GDDR7 GPU, Thunderbolt™ 5 connectivity, and AI-ready performance.
Gigabyte AORUS RTX 5090 AI Gaming Box, Thunderbolt 5, WaterForce AIO Cooling, 240mm Radiator, 2x 120mm Fans, FEP Tubes.
Dual GPU 48GB GDDR6 VRAM,PCIe 5.0 x8, Blower Fan+Vapor Chamber Cooling, for AI Inference&Large Model Deployment.
32 Xe2 Cores, 2600MHz Clock, 367 TOPS AI Compute, 256-Bit 608GB/s Bandwidth, 290W TBP for AI Inference and 3D Rendering.
A powerful 20L compact system purpose-built for AI: Intel® Core™ Ultra 9 275HX (24 cores, up to 5.4GHz), Dual PCIe Gen5 GPU support – accelerate your AI workflows, 4x 64GB DDR5 up to 7200MT/s (OC) – massive bandwidth, Thunderbolt 5 + M.2 + SlimSAS +PCIe x4 + 10G LAN + Wi-Fi 7, Independent CPU & GPU thermal chambers.
A testable hypothesis, not a product. BITCOS (Intel Research, arXiv:2609.16338, Sep 2026) cuts how many bits a ternary weight needs — storage cost 2 − z bits/weight, 1.485 on the sparsest (dense) checkpoint, zeros up to 51.5%. FreeToken decides which MoE experts must sit on the GPU, cross PCIe, or be computed on the CPU. The emerging FreeToken-Intel port shows the same offload architecture leaving CUDA for Xe2. Ternary sparse-MoE checkpoints already exist — the paper’s own Table I lists three: CAT-Q Qwen3-30B-A3B, CAT-Q Qwen3-235B-A22B and Maple 20B-A1B. Their zero density is only 33–41%, so their experts cost 1.61–1.80 bit/weight with scales, not 1.485. The two ideas still compound, but mostly through ternarisation itself: 2.4–2.8× fewer bytes per expert than MXFP4/NVFP4. The BITCOS layout adds 1.18–1.25× over 2-bit packing and 0.97–1.02× over 5-trit. The expert bank shrinks, every cache miss moves fewer bytes, the same VRAM cache holds more hot experts — and the move-vs-compute split should stay put, unless CPU unpack stops being bandwidth-bound. Essay: The 1.5-Bit Expert · thread on X. ▲ measured · ◇ still to be built
expert bank shrinks toward (2−z) bits/weight + scales. CAT-Q 235B-A22B: ≈52 GB of ternary weights vs ≈125 GB at MXFP4 — fits Olares’ 96 GB.
every cache miss moves fewer bytes across the bus. the miss gets cheaper.
same cache budget holds more hot experts — + free slot. hit-rate rises.
split = bandwidth ratio. Morefine: 3.3 : 19.9 GB/s → 14.2% predicted, 14.8% measured. a smaller payload speeds both paths alike; the split moves only if CPU unpack turns instruction-bound.
does a 1.6–1.8-bit ternary expert change FreeToken’s offload economics for models that still don’t fit — e.g. CAT-Q 235B-A22B on 24 GB? a ternary 30B-A3B is ≈7 GB of weights and needs no offload at all.
measurable. not yet built. →
· OLARES 24GB · MOREFINE 12GB · AMD 24GB · [INTEL 32GB?]