Σ‑Mem raises peer‑selection accuracy from 46.22 % to 71.10 % under extreme counterfactual reliability shifts, showing that an online reliability memory can turn a weak central model into a discerning arbitrator of peers[1]. The paper attributes this jump to the ability of Σ‑Mem to record and query competence evidence, allowing the central agent to route queries toward trustworthy collaborators as feedback accumulates.
Proposition 1 guarantees provenance non‑amplification for any repair operator in LedgerMind, ensuring every new ledger entry originates from a tool output rather than being fabricated during post‑hoc correction[2]. This formal safeguard forces each reasoning step to cite an active ledger entry, eliminating the classic “hallucination amplification” problem that plagues unconstrained multimodal agents.
The weakest‑link mechanism quarantines 4,057 faulty claims—29 % of the evaluation split—by tracing each rejection to a low‑graded narrator, demonstrating that graded transmitter reliability can automatically block erroneous knowledge propagation[3].
The jarḥ–taʿdīl loop recovers three of four narrator grades from audit evidence alone, proving that source‑grade inference can be refreshed without human intervention and that the system self‑corrects as new evidence arrives[3].
These advances still leave open key failure modes: Σ‑Mem depends on timely correctness feedback, which may be sparse or noisy in real deployments; LedgerMind’s guarantee assumes tool outputs are themselves accurate, offering no protection against systematic tool bias; and the Isnad framework’s grade‑recovery loop missed the highest‑fault narrator, hinting that extreme adversarial chains could slip through. Moreover, all three systems have been evaluated on benchmark suites rather than live user interactions, so their robustness under open‑world queries remains unproven.
Deploying a structured evidence ledger together with an online reliability memory should become the default architecture for multimodal assistants, allowing developers to trace every claim back to a graded source and to suppress hallucinations before they reach the user.













