SOLODKIY.CV // NEXT DRAM.GOLD LAB
WORKING NOTE · 19 SEPTEMBER 2026 · REVISED 20 SEPTEMBER · CONTINUES NEMOTRON-FOL · PROVISIONAL AUTHORITY · FREETOKEN · POWER→INTELLIGENCE

FUNGIBLE COMPUTE.
NON-FUNGIBLE AUTHORITY.

I am not trying to build a leaderboard in my flat. I want to know whether AI compute can move between different local machines, vendors and power envelopes while the right to make a consequential decision stays bounded, auditable and local. Two Tiiny AI Labs are already on the way. The AORUS RTX 5090 AI Box and Intel Arc Pro B70 / Dual B60 are wishlist instruments: if I can get access to them, these are the questions I would test — and publish, including negative results.

QUESTIONS FIRST · HARDWARE SECOND · RAW FAILURES INCLUDED
IN LAB · OLARES ONE · 96GB RAM · RTX 5090 MOBILE 24GB IN LAB · MINISFORUM M1 LITE · 96GB RAM · AMD 24GB IN LAB · MOREFINE M900 · 48GB RAM · NVIDIA 12GB IN TRANSIT · 2× TIINY AI LAB WISHLIST · AORUS RTX 5090 AI BOX 32GB WISHLIST · ARC PRO B70 32GB WISHLIST · MAXSUN DUAL B60 24+24GB HYPOTHESIS · BITCOS × FREETOKEN × B70
01

The hardware is an instrument, not the hypothesis

A useful experiment should end with a statement about authority, reliability, portability, cost or capability — not merely a faster token counter. The unit I care about is a verified decision: a result whose evidence, rules and permitted action can be reconstructed later.

CURRENT REFERENCE

Olares One

Intel Core Ultra 9 275HX, RTX 5090 Mobile 24GB, 96GB RAM, Thunderbolt 5. This is the CUDA reference and the machine that already ran my Nemotron-FOL and Sovereign Decision Plane experiments.

Role: heavy local perception, orchestration, reference execution path.

ARRIVING ×2

Tiiny AI Lab

ARM + NPU/dNPU, 80GB LPDDR5X split across SoC and dNPU, 30W TDP. The important feature is not “120B in a tiny box”; it is two separate physical machines that can fail, disagree and arrive late.

Role: physically separate proposer / reviewer / deferred-evidence nodes.

WISHLIST

AORUS RTX 5090 AI Box

Desktop RTX 5090, 32GB GDDR7, Thunderbolt 5. Eight extra GB over the internal 5090M only matter if they unlock a class of work that 24GB cannot do cleanly.

Role: escalation / second-read node for genuinely heavier evidence.

WISHLIST

Intel Arc Pro B70

Xe2 / XMX, 32GB GDDR6, 608GB/s spec (~500GB/s measured read in BITCOS), up to 367 INT8 TOPS. The point is not “Intel versus NVIDIA”; it is a third execution stack for testing software sovereignty and decision invariance.

Role: XPU execution-diversity check and portability test.

WISHLIST / SPECIAL CASE

MAXSUN Dual B60 48G

Two B60 GPUs, 24GB each — not one unified 48GB heap. The card needs a host that exposes both x8 links; I would not assume my M1 Lite can do that without a compatibility proof.

Role: two-worker topology or multi-GPU test, only on a compatible host.

ALREADY LEARNED

FreeToken / FOL

FreeToken showed that model placement can violate the usual “VRAM = capability” intuition. Nemotron-FOL showed something more important: a confident model error can be prevented from becoming authority by a deterministic gate.

Next: move those guarantees across physical machines and backends.

02

Six questions worth adding silicon for

Each question has a failure condition. If the result is “no”, that is still a useful public artifact. The devices do not get to define success.

Q01 · LIVE CASE 03 · AUTHORITY UNDER PARTITION

Does bounded authority survive when perception, verification and deferred evidence live on different machines?

A Tiiny can propose facts at T0, Olares can perform a heavier second read later, and a second Tiiny can deliver deferred evidence — while caps, TTL, revocation and ai_authorised_to_override = false remain invariant.

Why two Tiinys matter: I can physically disconnect one, delay it, replay an old packet, skew its clock or let it return FAIL after another node still holds stale PROVISIONAL state.

Falsifier: any path where delayed, stale or conflicting AI output widens authority, silently extends TTL, or authorises an action after revocation.

Q02 · FREETOKEN M2 · ARE AI WORKERS FUNGIBLE?

Can compute move while decision semantics stay put?

Run the same evidence and decision contract through CUDA, Tiiny runtime, ROCm where compatible, and later Intel XPU. I do not require identical prose or logits. I require the control plane either to reach the same permitted action or to detect material divergence before action.

Important distinction: different hardware is an independent execution path, not independent evidence. A true witness must bring separate evidence or verification; silicon diversity alone does not make a claim true.

Falsifier: backend/quantisation differences change a consequential verdict without triggering escalation. If that happens, “AI tokens” are not a fungible commodity for that task class.

Q03 · SMALLEST SUFFICIENT PERCEPTION

What proportion of real work deserves an expensive GPU at all?

Start with the smallest local node. Escalate only when a deterministic gate, missing evidence, multimodality or context requirements demand a stronger read. The metric is not tokens/second; it is the fraction of real tasks completed correctly at each rung.

Tiiny: cheap first read. Olares 5090M: heavy second read. AI Box: only if a named class of evidence cannot be handled cleanly inside 24GB.

Falsifier: the small node causes so many missed facts / false escalations that the “cheap perception” ladder costs more time or errors than simply using the 5090 reference.

Q04 · AORUS 5090 AI BOX · CAPABILITY, NOT SPEED

Do 32GB outside the box unlock a new task — or merely make the same task prettier on a chart?

I would look for named workloads that are qualitatively awkward in 24GB but clean in 32GB: less aggressive quantisation, no CPU offload, useful long context, concurrent multimodal components, image/video generation, or an escalation model that can keep all required state resident.

Why Olares is interesting: the host already has a 24GB 5090M and TB5. Olares documents NVIDIA eGPU support, plus a current Olares OS Gen1 workaround in its tested setup; Windows can renegotiate to Gen4 after restart. That makes bus behaviour part of the measurement rather than a footnote.

Falsifier: none of my real workloads gains a new reliable mode from 32GB, or the eGPU path introduces enough transfer / power / driver friction that the internal 24GB remains the practical choice.

Q05 · INTEL · SOFTWARE SOVEREIGNTY

Is my “sovereign” reasoning pipeline actually portable — or secretly a CUDA application?

Port the real pipeline: structured extraction, evidence packets, constrained output where supported, Prolog/Z3/OPA gates, regression cases and audit logs. B70 is useful because it provides one 32GB Intel/XMX surface; the research result is how much of the system survives a vendor/runtime change.

B70: cleanest first Intel target. Dual B60: a later topology experiment — two 24GB workers, not a magic 48GB monolith. It requires x8/x8 host support, so the host must be verified before promising a Minisforum pairing.

Falsifier: critical features or model/runtime support make the pipeline materially less reproducible, less auditable or impractical outside CUDA. That negative result is evidence about real software sovereignty.

Q06 · POWER → INTELLIGENCE

Once compute is replaceable, can intelligence follow the energy signal?

Only after Q02 establishes which task classes are safe to move. Then schedule movable work by latency, energy, availability and tariff: Tiiny for always-on low-power work, GPU nodes for expensive escalations, with the same authority contract after them.

The unit: joules and pence per verified decision, not tokens per joule in isolation. A cheap wrong answer is not efficient.

Falsifier: energy-aware routing changes outcomes, creates unacceptable delay, or saves too little to justify the orchestration complexity.

01 · PERCEIVE
Tiiny / GPU worker
Propose candidate facts. No authority.
02 · WITNESS
Evidence source
Independent evidence or deterministic validation — not merely “another GPU”.
03 · DECIDE
Prolog · Z3 · OPA
Apply caps, TTL, revocation and policy.
04 · ROUTE
FreeToken scheduler
Choose the smallest eligible worker by need, cost and energy.
05 · RECORD
Evidence packet
Hashes, versions, timestamps, decision trace; reconstruct later.
03

What each wishlist device would actually buy me

This is deliberately not a shopping recommendation. It is a list of new falsifiable questions unlocked by access to the hardware.

InstrumentWhat it addsQuestion it unlocksWhat would make it unnecessary
2× Tiiny
to try on Olares One, Minisforum or Morefine
Two low-power physical nodes with separate state/failureAuthority under partition; cheap perception; delayed evidenceNothing to buy — these are already coming
AORUS RTX 5090 Infinity
to try on Minisforum or Morefine
Nvidia RTX 5090 + 32GB over OculinkDoes +8\16GB GPU unlock a new escalation workload?instead of my current 24\12GB AMD\Nvidia capacity there [check my re-design vision]
AORUS 5090 AI Box
to try on Olares One
Desktop RTX 5090 + 32GB over TB5 beside a 24GB 5090MDoes +8GB / desktop GPU unlock a new escalation workload?If Q03 produces no workload that is qualitatively constrained by 24GB
Arc Pro B70 32GB
to try on Minisforum or Morefine
Intel Xe2/XMX and a different software/runtime pathCan the decision pipeline leave CUDA without losing semantics or auditability?If the pipeline cannot be matched closely enough to make the comparison meaningful
MAXSUN Dual B60 48G
to try on Minisforum or Morefine
Two independent 24GB Intel GPUs on one boardTwo-worker topology / multi-GPU scaling after the worker-fungibility question mattersIf a compatible x8/x8 host is unavailable, or two-worker independence adds no value
Host caveat: MAXSUN specifies GPU1 = PCIe 5.0 x8 and GPU2 = PCIe 5.0 x8; public testing reports that both GPUs require motherboard x8/x8 bifurcation because there is no onboard PCIe bridge. I would therefore label the B60 host TBD / compatible workstation required, not promise “M1 Lite + B60” before the motherboard/adapter path is proven.

THE LAB SHOULD BE ABLE TO DISPROVE ME.

If a £4k-ish GPU does not unlock a new capability, that is a useful result. If Intel portability is painful, publish the pain. If two Tiinys agree with themselves and still miss the truth, publish that too. A public artifact is more valuable when it can say “this idea failed here”.

04

What I would publish for developers

The community value is not that I own unusual boxes. It is that somebody publishes the method, raw traces, failure modes and exact boundary conditions on hardware people can actually obtain.

01 · Authority-under-partition harness

Reproducible two/three-host test: stale state, delayed FAIL, replayed packet, clock skew, node loss, TTL expiry, revocation. One command, pass/fail invariants.

02 · Cross-backend decision corpus

Same evidence packet and matched model/checkpoint where runtimes permit; versions, quantisation, seeds, raw outputs, extracted claims, final formal verdict. Material disagreements highlighted.

03 · Smallest-sufficient-compute ladder

Real workload classes routed Tiiny → current 5090M → optional Box / XPU. Report escalations, misses, false escalations, latency, wall power and total cost per verified result.

04 · AI Box capability map

Not “5090 is fast.” List the exact workloads that fail / degrade / require offload in 24GB and become clean in 32GB — plus the cost of TB5 and OS/driver caveats.

05 · CUDA → XPU porting diary

Time-to-first-result, unsupported features, model/quant gaps, structured-output behaviour, reproducibility and the patches needed. Installation pain is data when sovereignty is the claim.

06 · Negative results folder

Broken configurations, contradictory outputs, “not worth buying” outcomes and exact logs. No deleting a test because it makes a favourite device look bad.

Publication rules

Open invitation: if a vendor, lab or community member can lend access to the AI Box, B70 or Dual B60, I would use the same public harness and publish the complete result — including a result that says the hardware adds nothing useful to this research line.
05

Spec notes / source of truth

Device specifications are linked to current vendor documentation; research claims remain hypotheses until the experiment is run.

Olares One specs — olares.com/docs/one/spec
Olares eGPU on Olares OS — olares.com/docs/one/egpu-olares-os
Olares eGPU on Windows — olares.com/docs/one/egpu-windows
Tiiny Pocket specs — tiiny.ai/pages/tiiny-pocket
GIGABYTE AORUS RTX 5090 Infinity — gigabyte.com/.../GV-N5090AORUS IF-32GD
GIGABYTE AORUS RTX 5090 AI BOX — gigabyte.com/.../GV-N5090IXEB-32GD
My re-design: AORUS RTX 5090 Infinity & AI BOX — my vision #PerfectInfinityCraft
Intel Arc Pro B70 — intel.com/.../arc-pro-b70
MAXSUN Arc Pro B60 Dual 48G — maxsun.com/.../intel-arc-pro-b60-dual-48g-turbo
Gigabyte Aorus RTX 5060 Ti AI BOX 16GB

AORUS RTX 5060 Ti AI BOX delivers portable desktop-class graphics with a GeForce RTX™ 5060 Ti 16GB GDDR7 GPU, Thunderbolt™ 5 connectivity, and AI-ready performance.

Gigabyte AORUS RTX 5090 AI Box Water Coooled eGPU Thunderbolt 5 Gaming Box

Gigabyte AORUS RTX 5090 AI Gaming Box, Thunderbolt 5, WaterForce AIO Cooling, 240mm Radiator, 2x 120mm Fans, FEP Tubes.

MAXSUN Intel Arc Pro B60 48G Turbo Workstation Graphics Card

Dual GPU 48GB GDDR6 VRAM,PCIe 5.0 x8, Blower Fan+Vapor Chamber Cooling, for AI Inference&Large Model Deployment.

GUNNIR Intel Arc Pro B70 TF 32GB Workstation Graphics Card

32 Xe2 Cores, 2600MHz Clock, 367 TOPS AI Compute, 256-Bit 608GB/s Bandwidth, 290W TBP for AI Inference and 3D Rendering.

MS-ARL-HX Micro Station

A powerful 20L compact system purpose-built for AI: Intel® Core™ Ultra 9 275HX (24 cores, up to 5.4GHz), Dual PCIe Gen5 GPU support – accelerate your AI workflows, 4x 64GB DDR5 up to 7200MT/s (OC) – massive bandwidth, Thunderbolt 5 + M.2 + SlimSAS +PCIe x4 + 10G LAN + Wi-Fi 7, Independent CPU & GPU thermal chambers.

06

The 1.5-bit expert — BITCOS × FreeToken

A testable hypothesis, not a product. BITCOS (Intel Research, arXiv:2609.16338, Sep 2026) cuts how many bits a ternary weight needs — storage cost 2 − z bits/weight, 1.485 on the sparsest (dense) checkpoint, zeros up to 51.5%. FreeToken decides which MoE experts must sit on the GPU, cross PCIe, or be computed on the CPU. The emerging FreeToken-Intel port shows the same offload architecture leaving CUDA for Xe2. Ternary sparse-MoE checkpoints already exist — the paper’s own Table I lists three: CAT-Q Qwen3-30B-A3B, CAT-Q Qwen3-235B-A22B and Maple 20B-A1B. Their zero density is only 33–41%, so their experts cost 1.61–1.80 bit/weight with scales, not 1.485. The two ideas still compound, but mostly through ternarisation itself: 2.4–2.8× fewer bytes per expert than MXFP4/NVFP4. The BITCOS layout adds 1.18–1.25× over 2-bit packing and 0.97–1.02× over 5-trit. The expert bank shrinks, every cache miss moves fewer bytes, the same VRAM cache holds more hot experts — and the move-vs-compute split should stay put, unless CPU unpack stops being bandwidth-bound. Essay: The 1.5-Bit Expert · thread on X. ▲ measured · ◇ still to be built

FT·BT—01
LOCAL-AI MEMORY WALL
REPRESENT × PLACE × PORT
THE 1.5-BIT EXPERT
HYPOTHESIS ● REC
THREE MACHINES · ONE BENCH
2026-09-20
WHAT BITCOS × FREETOKEN × AN EMERGING INTEL ARC PORT SUGGEST ABOUT THE NEXT MEMORY WALL ▲ MEASURED   ◇ INFERRED / TO-BE-BUILT
0 C 01 · REPRESENT BITCOS — INTEL RESEARCH ▲ 29 TERNARY LLMs {−1, 0, +1} · ZEROS ≤ 51.5% STORAGE COST = 2 − z BITS / WEIGHT ▲ 1.485 BIT/WEIGHT · ▲ 1.27× DECODE ON XE2 + − + TERNARY WEIGHTS — ZEROS COLLAPSE BITMAP + SIGNS — 1.485 DENSE · 1.61–1.80 MoE FEWER BITS PER EXPERT 02 · PLACE FREETOKEN — EXPERT OFFLOAD ▲ GPT-OSS-120B RUNS ON 24 GB VRAM ▲ 35B-A3B ~20.5 T/S ON RTX 3060 12 GB (~7 GB VRAM) GPU LRU CACHE · HOST RAM · HYBRID CPU PATH HOST RAM — EXPERT POOL (LRU) HYBRID: COMPUTE ON CPU PCIe · CACHE MISS → FETCH GPU — LRU EXPERT CACHE + FREE SLOT QWEN 35B-A3B PREFILL ✓ DECODE ✓ OPENAI-COMPAT API BEYOND CUDA 03 · PORT FREETOKEN-INTEL · XE2 ▲ 35B PREFILL + DECODE ON ARC PRO B70 SYCL · LEVEL ZERO · torch.xpu — COMMUNITY PORT ◇ HYBRID + HONEST PERF NUMBERS — WIP ARC PRO B70 · 32 GB FREETOKEN-INTEL LIVE TOKENS · XE2 · ONEAPI SYCL LEVEL ZERO torch.xpu ◇ NOT COMPETITORS — LAYERS. ◇ A TERNARY SPARSE-MoE WOULD COMPOUND ALL THREE.
▲ EFFECT

RAM ▼

expert bank shrinks toward (2−z) bits/weight + scales. CAT-Q 235B-A22B: ≈52 GB of ternary weights vs ≈125 GB at MXFP4 — fits Olares’ 96 GB.

▲ EFFECT

PCIe ▼

every cache miss moves fewer bytes across the bus. the miss gets cheaper.

▲ EFFECT

VRAM CACHE ▲

same cache budget holds more hot experts — + free slot. hit-rate rises.

◇ PREDICTION

SCHEDULER ⇄

split = bandwidth ratio. Morefine: 3.3 : 19.9 GB/s → 14.2% predicted, 14.8% measured. a smaller payload speeds both paths alike; the split moves only if CPU unpack turns instruction-bound.

◇ OPEN

THE QUESTION

does a 1.6–1.8-bit ternary expert change FreeToken’s offload economics for models that still don’t fit — e.g. CAT-Q 235B-A22B on 24 GB? a ternary 30B-A3B is ≈7 GB of weights and needs no offload at all.

measurable. not yet built. →

▲ 29/29 ternary models measured · zeros up to 51.5% — Intel Research, Sep 2026
▲ BITCOS: storage cost = 2 − z · beats 5-trit packing in 26/29 models · 1.485 = dense CAT-Q Qwen3-1.7B
▲ BITCOS kernels: AVX2 / AVX-512 / Xe2 · 1.28× matvec · 1.27× end-to-end decode
▲ Arrow Lake 8P+16E (Olares’ 275HX topology): GEMV 1.13–1.27× vs 2-bit · Arc Pro B70: GEMV 1.01–1.12×, decode 1.02–1.27× · Lunar Lake CPU: slower than 2-bit
▲ FreeToken: GPT-OSS-120B on 24 GB VRAM · 35B-A3B ~20.5 tok/s on RTX 3060 12 GB (~7 GB VRAM)
▲ FreeToken-Intel (community): 35B prefill+decode on Arc Pro B70 · hybrid WIP · no perf claims yet
▲ ternary sparse-MoE checkpoints exist — CAT-Q Qwen3-30B-A3B / 235B-A22B (z 32.9 / 34.1%), Maple 20B-A1B (z 40.7%) · 1.61–1.80 bit/weight with scales · BITCOS Table I
◇ BITCOS-aware FreeToken expert backend + pack/unpack kernels — to be built
◇ per-expert zero density — not measured anywhere yet; if it varies, picking 5-trit or BITCOS per expert beats both
◇ B70-class card on the bench — for FreeToken-Intel hybrid with ternary experts (BITCOS alone on B70 Intel already measured) · loan / review sample requested, published either way
PUT THEM ON THE SAME BENCH AND SEE WHAT BREAKS.
BITCOS × FREETOKEN × XE2 — A TESTABLE HYPOTHESIS, NOT A PRODUCT ·

· OLARES 24GB · MOREFINE 12GB · AMD 24GB · [INTEL 32GB?]