Track 1 · Retrieval on LoCoMo
- A measured retrieval
What a single agent’s memory must get right before coordination matters: identity, chronology and evidence attribution, and the one result we consider defensible.8
Research programme
Whether a collective of agents that never share raw memory can answer, verify, repair and keep discovering knowledge as well as, or better than, systems that centralise or copy everything, and at what cost in storage, messages and exposure.12
The manuscript’s contributions, in its own framing: a bounded QuestionArtifact, five epistemic functions for continual discovery, a proposed (not validated) sparse question-priority mechanism, and four distinct evidence classes that test retrieval, repair, acquisition and distribution under failure.34 The programme does not claim graph memory, persistent memory, reconstruction or hierarchy as standalone novelties; the contribution is the combination: first-class travelling questions over a lineage-typed fabric, with explicit separation of acquisition, availability and independent support.5
The programme studies three things together: bounded questions that travel between memory holders as first-class objects, a fabric in which every claim carries its lineage, and an evaluation discipline that separates what was acquired, what stayed available and what remained independently supported.61 Graph memory, persistent memory, reconstruction and hierarchy are prior art and are not claimed as inventions.
Each class answers a different question with a different workload and a different grader. The chip in the last column is the one that marks every number elsewhere on this site.
| Class | What it is | How it is scored | Reproducibility | Chip |
|---|---|---|---|---|
| A | A real retrieval and answering system on LoCoMo with two evidence holders, one per speaker7 | LLM-judge accuracy against the gold answer; recall@k as a gold-fragment presence proxy8 | Per-question artifacts committed in the public repository at e054178 (see Artifacts below); leakage-free harness748 | A |
| B | A nine-capability deterministic repair fixture on 5 nodes, 30 seeds9 | Mean capability survival per strategy and intervention | Bitwise reproducible; SHA256-pinned; 277 tests pass10 | B |
| T | Deterministic tests of the cross-node coordinator (mock 4-node fixture and 2-node real-SQLite fixture)11 | Pass/fail invariants; 109 tests12 | Deterministic clocks and seeded RNG13 | T |
| C | A synthetic correlated-evidence acquisition simulation, reported in the manuscript; artifacts not in the current checkout1415 | Accuracy and confident-error rates in a symbolic world | Not reproducible from this repository | C |
| D | A population-scale symbolic distribution simulation of 100 to 100,000 agents with no LLM calls, reported in the manuscript1416 | Knowledge survival, task accuracy, independent-support survival and contradiction F1, kept as four separate outcomes51 | Not reproducible from this repository15 | D |
| Org benchmark | Its own simulator with a five-layer hierarchy; in progress17 | Discovery, privacy, cost, robustness and survival families (definitions only) | Manifests under results/raw; no results published18 | IP |
None is a deployment of the complete architecture. Numerical results from different workloads and graders cannot be pooled into one architectural accuracy.No class is a deployment of the whole architecture, and numbers from different workloads and graders are never pooled into one accuracy figure.14 Every number elsewhere on this site carries one of these chips. |
||||
Each track is mapped onto the evidence classes it draws on. The numbers live on the benchmarks page; this page holds the method.
What a single agent’s memory must get right before coordination matters: identity, chronology and evidence attribution, and the one result we consider defensible.8
Classes T and B, plus the manuscript’s Class C and D simulations: what survives when nodes, roots, routes, authorizations or whole domains fail, and whether lineage-aware placement and repair beat replication at equal cost.1920
Whether a worker → team → department → region → executive hierarchy of local agents can discover effects no single unit can see while exposing less raw data than centralised alternatives.17
Three columns, never merged.
Never shown as shipped; hatched in diagrams.
tools/build.py, which the deploy workflow runs on every push to main, strips it and fails the build if any of it, or a citation to it, survives. It returns only when the build is run with --manuscript-cleared, after the authors confirm that publication does not breach the conditions of review.3550 Where an ungated sentence rests on the manuscript, the site says it in its own words and keeps the citation. A source note that cites the manuscript by section and margin line is a reference, not content: it stays in the deployed page whenever an ungated sentence points to it, and is held out only when every marker that points to it is held out with its copy.One manuscript describing the lineage-first fabric, the QuestionArtifact, continual discovery and the four evidence classes is under review.35 The venue is not named and the manuscript is not distributed from this site while review is in progress; the site describes its content in the lab’s own words and marks every simulation figure as “reported in the manuscript”. After the review decision this section will link the paper or its preprint and the corrections log will record the date.
Note first: source paths on this site cite the NeuralGraph repository, which is public to read at github.com/anovruzov/NeuralGraph but has no licence, so all rights are reserved and nothing in it is offered for reuse.3648 “An auditor can check every sentence” therefore holds for any reader: the rows below link to the commit or branch each artifact is recorded at, and a request to the lab is needed only for what is not in the repository.
| Class | What exists | Where | Access |
|---|---|---|---|
| A LoCoMo retrieval | Per-question answer and retrieval artifacts for every experiment; the 18-experiment record37 | docs/research/RESULTS_ALL.md and demo/results/ at commit e054178, on main4948 | Results document → Per-question JSON → |
| B Repair fixture | Four JSON artifacts, SHA256-pinned, regenerate byte-identically (re-verified 2026-09-06, macOS arm64, Python 3.11); 277 tests pass; at commit 13a1729 on a research branch38 | NeuralGraph/coordination/artifacts/ at commit 13a1729 on branch claude/session-analysis-continuation-sm1a1t, not merged into main3848 | Artifacts and SHA256SUMS → |
| T Coordination tests | 109 deterministic tests including a real-storage restart test12 | NeuralGraph/tests/test_coordination.py, test_coordination_lineage.py, test_coordination_storage.py, test_coordination_experiment_restart.py, on main12 | Test modules → |
| C Acquisition probe | Reported in the manuscript; artifacts not in the current checkout15 | Not in the repository | After review; ask the lab → |
| D Population simulation | Reported in the manuscript, whose current revision is based on the campaign’s written reports rather than on re-running its code15 | Not in the repository | After review; ask the lab → |
| IP Organisational benchmark | results/processed and manifests, in progress; every manifest states the model-backend limitation39 | mycelic-org-benchmark/results/ on branch claude/mycelic-org-benchmark-9cnf35, in progress, not merged into main3948 | Design and code → Results: when the campaign finishes, dated in the corrections log |
qwen2.5:7b-instruct while the code’s default is an OpenAI-compatible endpoint at 127.0.0.1:1234 with google/gemma-4-e4b; the site therefore states no default model and treats model names as configuration.4445Every future discrepancy is added here and dated in the corrections log. Corrections log →
Use the subject “Research collaboration / artifact access”.
Each marker above points to one note below. Paths are relative to the root of the NeuralGraph repository, which is public to read at github.com/anovruzov/NeuralGraph with no licence selected (all rights reserved); they are given as text, with line numbers checked against main on 2026-09-17, and files under mycelic-org-benchmark/ are on branch claude/mycelic-org-benchmark-9cnf35, not yet merged into main, and the Class B artifacts are at commit 13a1729 on branch claude/session-analysis-continuation-sm1a1t, not merged into main; the Artifacts table above links those directly. “Manuscript” is the paper under review, cited by section and printed margin line. Notes that support only manuscript-derived copy are held out of the deployed page together with that copy; a note cited by an ungated sentence stays, since a section-and-line reference reproduces nothing from the paper.