LOCAL AI LAB // RUN COMPLETE
Milestone 02 — 2026.09.06–07 // Olares One

NEMOTRON ×
EVIDENCE GATE

A local multimodal model read the passport. Then the deterministic layer refused to believe it.

NVIDIA Nemotron 3 Nano Omni RTX 5090 24GB 96GB RAM Prolog / FOL Synthetic data only
THE MODEL PERCEIVES. CODE VERIFIES. LOGIC DECIDES. THE MODEL ONLY EXPLAINS.

// Research Question

Can a single local multimodal model become the perception layer of an evidence-gated compliance agent — without becoming the final authority?

Why Nemotron?

One local model for text + vision (and an advertised audio capability in the installed Ollama build) is attractive for privacy-sensitive workflows: ID documents, OCR, case files, interviews, and multimodal evidence.

What we refused to trust

Model confidence. In regulated workflows, a plausible answer is not evidence. The experiment was designed so a deterministic layer could reject a confident but malformed extraction.

// Test Environment

Hardware

  • Olares One
  • NVIDIA GeForce RTX 5090 Laptop GPU
  • 24,463 MiB VRAM
  • 93 GiB visible system RAM (~96 GB installed)

Runtime

  • Ollama 0.30.11 in Olares / Kubernetes
  • Nemotron 33b-q4_K_M
  • Model payload ~27.6 GB
  • OpenAI-style local HTTP workflow via Ollama API

Formal Layer

  • Python deterministic validator
  • ICAO-style MRZ length / charset checks
  • MRZ check digits + cross-field comparisons
  • SWI-Prolog policy engine
82%
GPU share
Ollama processor split
23.8
GiB-ish VRAM used
23,819 MiB
48.8s
Cold text run
11.6s
Warm text run
12–20s
Vision API runs
observed range

Runtime note

Ollama reported an 18% CPU / 82% GPU split for the loaded model. The GPU reached roughly 23.8 / 24.5 GB VRAM used, so this is not a “huge headroom” deployment: it is a tight local fit with CPU/RAM spillover.

// Architecture

01Passport imageSynthetic specimen only. No real customer PII.
02Nemotron perceptionExtract human-readable fields + candidate MRZ strings.
03Deterministic gateLength, charset, MRZ check digits, cross-field consistency, uncertainty sanity.
04Prolog / FOL policyConvert verified / failed / conflict facts into a policy decision.
05Nemotron explanationExplain the decision to a human. Explicitly forbidden from changing it.
PERCEPTION
“What does the image appear to contain?”
Nemotron
VERIFICATION
“Does the extraction survive deterministic checks?”
Python / ICAO logic
DECISION
“Given verified facts, what policy action follows?”
Prolog / FOL
EXPLANATION
“Tell the reviewer what happened — without overriding the decision.”
Nemotron

// T1 — Clean Synthetic Passport

The first specimen was intentionally simple: high-contrast synthetic text, known ground truth, no privacy risk.

What worked

  • 9 / 9 human-readable fields were extracted correctly in the clean run.
  • The model separated the printed passport number from the MRZ check digit.
  • Structured output via Ollama API produced clean JSON.

What failed

  • MRZ was not reproduced exactly.
  • The model added / dropped characters in machine-readable lines.
  • It reported uncertain_fields: [] despite those errors.

API timing snapshot

total_duration ≈ 19.65 s
load_duration ≈ 15.52 s
prompt_eval_duration ≈ 1.23 s
eval_duration ≈ 2.83 s

// T2 — Degraded Evidence

The same synthetic passport was deliberately damaged: blur, slight rotation, JPEG compression, and occluded characters in the passport number and MRZ.

Adversarial finding: confident reconstruction

Instead of returning null and marking fields uncertain, Nemotron attempted to reconstruct unreadable fragments. The output contained malformed values while still declaring uncertain_fields: [].

{
  "passport_number": "L8989U<..<.5",
  "mrz_line_2": "L898902C36UT0900312: :<<<10",
  "uncertain_fields": []
}
12.17s
T2 wall clock
FAIL
Uncertainty calibration
USEFUL
As candidate extractor

// Deterministic Gate

At this point the model stops being the authority. A small deterministic validator checks what can be checked without asking a language model to “reason” about it.

Checks

  • Passport number format
  • MRZ 44-character structure
  • Allowed MRZ alphabet
  • Document number check digit
  • Date-of-birth check digit
  • Expiry check digit
  • Optional-data check digit
  • Composite check digit
  • Visible field ↔ MRZ consistency

Gate behavior

The validator does not try to “fix” the model. It emits PASS, FAIL, or CONFLICT, then returns a deterministic verdict such as NEED_MORE_EVIDENCE.

// T1b — Formally Valid Control Specimen

We corrected the synthetic MRZ so both lines were exactly 44 characters and the embedded check digits were formally valid. This created a clean control specimen.

The interesting part: the model still failed the machine-readable layer

PASS passport_number_format
FAIL mrz_line_1_structure (model returned 42 chars)
PASS mrz_line_2_structure (44 chars)
PASS document_number_check
FAIL date_of_birth_check
FAIL expiry_date_check
FAIL optional_data_check
FAIL composite_check
CONFLICT nationality / sex / DOB / expiry vs MRZ
FAIL uncertainty_calibration

Deterministic verdict: NEED_MORE_EVIDENCE

12.18s
T1b wall clock
9/9
Readable fields
NO
MRZ trusted automatically

// Prolog / FOL Policy Decision

The validator exports facts. The policy engine reasons over those facts — not over raw model prose.

verification(passport_number_format, pass).
verification(mrz_line_1_structure, fail).
verification(mrz_line_2_structure, pass).
verification(date_of_birth_check, fail).
verification(expiry_date_check, fail).
verification(composite_check, fail).
verification(nationality_vs_mrz, conflict).
verification(sex_vs_mrz, conflict).
...

decision(request_better_evidence) :-
    \+ automated_acceptance.

FOL output

decision=request_better_evidence
reasons=[
  cross_field_conflict,
  deterministic_failure,
  identity_conflict,
  mrz_integrity_failed
]

// Final Loop — The Model Explains, But Does Not Decide

The Prolog decision was passed back to Nemotron with an explicit instruction: explain it to a human reviewer, but do not change, soften, or reverse it.

Nemotron human explanation

{
  "decision": "request_better_evidence",
  "explanation": "Human-readable fields were successfully extracted, but formal MRZ validation and cross-field consistency checks failed due to integrity and identity conflicts.",
  "next_action": "Request a new, valid evidence submission from the applicant."
}
Nemotron perceives. Code verifies. Logic decides. Nemotron explains.

// What This Actually Proves

Proven in this pilot

  • Nemotron 3 Nano Omni Q4_K_M runs locally on the Olares One.
  • A 24GB RTX 5090 can host most of the model with CPU/RAM spillover.
  • Vision extraction works via local Ollama API.
  • Human-readable passport fields can be extracted accurately on a simple synthetic specimen.
  • Deterministic MRZ / cross-field checks can catch model errors.
  • A Prolog policy layer can deterministically refuse automated acceptance.
  • The model can explain the policy result without changing it.

Not proven

  • Production-grade OCR accuracy across real passports or ID types.
  • Reliable uncertainty calibration.
  • Legal or regulatory compliance.
  • Document authenticity, fraud detection, liveness, or biometric verification.
  • That Prolog rules faithfully encode any real law or institution policy.
  • That the model's audio / video paths are production-ready.
  • Any claim involving real customer PII — none was used.

// Why It Matters

The useful lesson was not that a local model can OCR a passport. The useful lesson was that it can be confidently wrong in exactly the place where deterministic evidence exists.

01 // Privacy

Local multimodal perception can keep sensitive evidence on-device while still exposing a structured interface to downstream controls.

02 // Auditability

Checksums, structure rules, and explicit conflicts leave a trace that is reproducible and inspectable without trusting a model's confidence.

03 // Authority Separation

The model is useful without being sovereign. Different layers have different jobs — and different failure modes.

// Experiment Log

SEP 06
Nemotron pulled into Olares Ollama storage (~27 GB tag; ~27.6 GB API size).
SEP 06
Cold / warm runtime measured. ~82% GPU / 18% CPU split; ~23.8 GB VRAM used.
SEP 07
Synthetic passport generated. Clean extraction: 9/9 human-readable fields, MRZ errors, no declared uncertainty.
SEP 07
Degraded document test: model reconstructed unreadable fragments and still reported no uncertainty.
SEP 07
Deterministic MRZ / cross-field validator added; malformed evidence routed to NEED_MORE_EVIDENCE.
SEP 07
SWI-Prolog policy layer added; final decision = request_better_evidence.
SEP 07
Nemotron explanation layer completed the loop without overriding the FOL decision.

// Next

Audio / Video

The installed model reports audio + vision capabilities through Ollama. Next experiment: 30–60 second audio / video → structured evidence → deterministic gate.

Harder Documents

Realistic specimen scans, different document layouts, glare, perspective distortion, and structured confidence benchmarks.

Policy Coverage

Move from one passport decision to a small library of explicit, versioned rules: evidence completeness, expiry, jurisdiction, product eligibility.

Local AI OS

Keep perception, verification, logic, explanation, and later orchestration as separate components rather than one trusted black box.

// Notes & Caveats

This is a home-lab engineering experiment, not a benchmark paper and not a compliance certification. Synthetic passport data was used. The deterministic checks validate properties of the extracted data; they do not prove that an identity document is genuine or that a person is who they claim to be.

Model tag: nemotron3:33b-q4_K_M
Ollama version: 0.30.11
API-reported size: 27,638,631,632 bytes
Ollama processor split: 18% CPU / 82% GPU
GPU after load: 23,819 MiB / 24,463 MiB
Cold text run: 48.759 s
Warm text run: 11.562 s
T1 API total_duration: 19.653508780 s
T1 API load_duration: 15.516499925 s
T1 API prompt_eval_duration: 1.229832000 s
T1 API eval_duration: 2.826628000 s
T2 wall clock: 12.17 s
T1b wall clock: 12.18 s
Final FOL decision: request_better_evidence