A local multimodal model read the passport. Then the deterministic layer refused to believe it.
Can a single local multimodal model become the perception layer of an evidence-gated compliance agent — without becoming the final authority?
One local model for text + vision (and an advertised audio capability in the installed Ollama build) is attractive for privacy-sensitive workflows: ID documents, OCR, case files, interviews, and multimodal evidence.
Model confidence. In regulated workflows, a plausible answer is not evidence. The experiment was designed so a deterministic layer could reject a confident but malformed extraction.
33b-q4_K_MRuntime note
Ollama reported an 18% CPU / 82% GPU split for the loaded model. The GPU reached roughly 23.8 / 24.5 GB VRAM used, so this is not a “huge headroom” deployment: it is a tight local fit with CPU/RAM spillover.
The first specimen was intentionally simple: high-contrast synthetic text, known ground truth, no privacy risk.
uncertain_fields: [] despite those errors.The same synthetic passport was deliberately damaged: blur, slight rotation, JPEG compression, and occluded characters in the passport number and MRZ.
Instead of returning null and marking fields uncertain, Nemotron attempted to reconstruct unreadable fragments. The output contained malformed values while still declaring uncertain_fields: [].
{
"passport_number": "L8989U<..<.5",
"mrz_line_2": "L898902C36UT0900312: :<<<10",
"uncertain_fields": []
}
At this point the model stops being the authority. A small deterministic validator checks what can be checked without asking a language model to “reason” about it.
The validator does not try to “fix” the model. It emits PASS, FAIL, or CONFLICT, then returns a deterministic verdict such as NEED_MORE_EVIDENCE.
We corrected the synthetic MRZ so both lines were exactly 44 characters and the embedded check digits were formally valid. This created a clean control specimen.
Deterministic verdict: NEED_MORE_EVIDENCE
The validator exports facts. The policy engine reasons over those facts — not over raw model prose.
verification(passport_number_format, pass).
verification(mrz_line_1_structure, fail).
verification(mrz_line_2_structure, pass).
verification(date_of_birth_check, fail).
verification(expiry_date_check, fail).
verification(composite_check, fail).
verification(nationality_vs_mrz, conflict).
verification(sex_vs_mrz, conflict).
...
decision(request_better_evidence) :-
\+ automated_acceptance.
decision=request_better_evidence reasons=[ cross_field_conflict, deterministic_failure, identity_conflict, mrz_integrity_failed ]
The Prolog decision was passed back to Nemotron with an explicit instruction: explain it to a human reviewer, but do not change, soften, or reverse it.
{
"decision": "request_better_evidence",
"explanation": "Human-readable fields were successfully extracted, but formal MRZ validation and cross-field consistency checks failed due to integrity and identity conflicts.",
"next_action": "Request a new, valid evidence submission from the applicant."
}
The useful lesson was not that a local model can OCR a passport. The useful lesson was that it can be confidently wrong in exactly the place where deterministic evidence exists.
Local multimodal perception can keep sensitive evidence on-device while still exposing a structured interface to downstream controls.
Checksums, structure rules, and explicit conflicts leave a trace that is reproducible and inspectable without trusting a model's confidence.
The model is useful without being sovereign. Different layers have different jobs — and different failure modes.
The installed model reports audio + vision capabilities through Ollama. Next experiment: 30–60 second audio / video → structured evidence → deterministic gate.
Realistic specimen scans, different document layouts, glare, perspective distortion, and structured confidence benchmarks.
Move from one passport decision to a small library of explicit, versioned rules: evidence completeness, expiry, jurisdiction, product eligibility.
Keep perception, verification, logic, explanation, and later orchestration as separate components rather than one trusted black box.
This is a home-lab engineering experiment, not a benchmark paper and not a compliance certification. Synthetic passport data was used. The deterministic checks validate properties of the extracted data; they do not prove that an identity document is genuine or that a person is who they claim to be.
Model tag: nemotron3:33b-q4_K_M Ollama version: 0.30.11 API-reported size: 27,638,631,632 bytes Ollama processor split: 18% CPU / 82% GPU GPU after load: 23,819 MiB / 24,463 MiB Cold text run: 48.759 s Warm text run: 11.562 s T1 API total_duration: 19.653508780 s T1 API load_duration: 15.516499925 s T1 API prompt_eval_duration: 1.229832000 s T1 API eval_duration: 2.826628000 s T2 wall clock: 12.17 s T1b wall clock: 12.18 s Final FOL decision: request_better_evidence