Research programme

What the programme asks, and how it keeps its evidence apart

Whether a collective of agents that never share raw memory can answer, verify, repair and keep discovering knowledge as well as, or better than, systems that centralise or copy everything, and at what cost in storage, messages and exposure.12


What the programme asks

The manuscript’s contributions, in its own framing: a bounded QuestionArtifact, five epistemic functions for continual discovery, a proposed (not validated) sparse question-priority mechanism, and four distinct evidence classes that test retrieval, repair, acquisition and distribution under failure.34 The programme does not claim graph memory, persistent memory, reconstruction or hierarchy as standalone novelties; the contribution is the combination: first-class travelling questions over a lineage-typed fabric, with explicit separation of acquisition, availability and independent support.5

The programme studies three things together: bounded questions that travel between memory holders as first-class objects, a fabric in which every claim carries its lineage, and an evaluation discipline that separates what was acquired, what stayed available and what remained independently supported.61 Graph memory, persistent memory, reconstruction and hierarchy are prior art and are not claimed as inventions.


Evidence classes, kept apart

Each class answers a different question with a different workload and a different grader. The chip in the last column is the one that marks every number elsewhere on this site.

The evidence classes used on this site: what each one is, how it is scored, whether it can be reproduced from the repository, and its chip.
Class What it is How it is scored Reproducibility Chip
A A real retrieval and answering system on LoCoMo with two evidence holders, one per speaker7 LLM-judge accuracy against the gold answer; recall@k as a gold-fragment presence proxy8 Per-question artifacts committed in the public repository at e054178 (see Artifacts below); leakage-free harness748 A
B A nine-capability deterministic repair fixture on 5 nodes, 30 seeds9 Mean capability survival per strategy and intervention Bitwise reproducible; SHA256-pinned; 277 tests pass10 B
T Deterministic tests of the cross-node coordinator (mock 4-node fixture and 2-node real-SQLite fixture)11 Pass/fail invariants; 109 tests12 Deterministic clocks and seeded RNG13 T
C A synthetic correlated-evidence acquisition simulation, reported in the manuscript; artifacts not in the current checkout1415 Accuracy and confident-error rates in a symbolic world Not reproducible from this repository C
D A population-scale symbolic distribution simulation of 100 to 100,000 agents with no LLM calls, reported in the manuscript1416 Knowledge survival, task accuracy, independent-support survival and contradiction F1, kept as four separate outcomes51 Not reproducible from this repository15 D
Org benchmark Its own simulator with a five-layer hierarchy; in progress17 Discovery, privacy, cost, robustness and survival families (definitions only) Manifests under results/raw; no results published18 IP
None is a deployment of the complete architecture. Numerical results from different workloads and graders cannot be pooled into one architectural accuracy.No class is a deployment of the whole architecture, and numbers from different workloads and graders are never pooled into one accuracy figure.14 Every number elsewhere on this site carries one of these chips.

The three tracks

Each track is mapped onto the evidence classes it draws on. The numbers live on the benchmarks page; this page holds the method.

Track 1 · Retrieval on LoCoMo

  • A measured retrieval

What a single agent’s memory must get right before coordination matters: identity, chronology and evidence attribution, and the one result we consider defensible.8

Track 1 on the benchmarks page →

Track 2 · Coordination and capability survival

  • T deterministic coordination tests
  • B deterministic fixture
  • C synthetic acquisition simulation, manuscript
  • D population-scale symbolic simulation, manuscript

Classes T and B, plus the manuscript’s Class C and D simulations: what survives when nodes, roots, routes, authorizations or whole domains fail, and whether lineage-aware placement and repair beat replication at equal cost.1920

Track 2 on the benchmarks page →

Track 3 · Organisational aggregation

  • In progress
  • IP in progress, no results

Whether a worker → team → department → region → executive hierarchy of local agents can discover effects no single unit can see while exposing less raw data than centralised alternatives.17

Track 3 on the benchmarks page →


Measured, demonstrated, proposed

Three columns, never merged.

Measured

  • Measured
  • A measured retrieval
  • Per-agent routing and dialogue-pair back-fill gains; judge sensitivity; candidate dilution.2122
  • Negative results: graph-neighbour expansion, listwise reranking, context-size changes, an LLM router, scoring-formula fixes, reply-only pairs, pair-aware reranking, a 35B answer model, vector+BM25+graph fusion, an LLM-built entity graph.23

Demonstrated

  • Fixture, tests or simulation only
  • B Lineage-aware repair, 7 of 9 capabilities.19
  • T Gate-1 distributed reconstruction across two real SQLite-backed holders.24
  • D Lineage-aware placement and continual questioning effects at up to 100,000 symbolic agents.25
  • Population-scale simulation results, under review.

Proposed only

  • Proposed

Never shown as shipped; hatched in diagrams.

  • The sparse question-priority formula and the stopping rule.4
  • Federated public discovery across regions.26
  • A deployed population of reasoning agents; no experiment validates the entire continual-discovery architecture end to end.no experiment yet exercises the whole continual-discovery architecture at once.27
  • The five epistemic functions (verification, contradiction, relation discovery, hypothesis generation, prediction/falsification) are architectural functions, not all measured natural-language capabilities.The five epistemic functions (verification, contradiction, relation discovery, hypothesis generation, prediction/falsification) are roles the architecture defines; not every one of them has been measured as a natural-language capability.28

Reporting rules an auditor can hold us to

  1. Every figure on the site has a source note with the document and section, and a row in the claims register; the register wins if a page disagrees. Claims register →
  2. Two different system ladders exist. The manuscript’s Class-D ladder (B0 to B7, B3Q, B6C) and the organisational benchmark’s ladder (B0 to B9 plus ORACLE) are always prefixed “manuscript” or “bench” so that “B7” is never ambiguous.2930
  3. LoCoMo category labels are shown “as labelled” because the inherited category constants swap single-hop and multi-hop; aggregates are unaffected.31
  4. Superseded numbers appear only in the corrections log.32 Corrections log →
  5. Wording rules: “behaviour is stable across four simulated scales”, never “scales to 100,000 agents”; “designed so that raw data does not leave the device”, never “guaranteed”.3334
  6. Manuscript-derived content (Class C and D tables, verbatim quotations, formal definitions) is not in the deployed pages. In the committed source it sits between gate markers; tools/build.py, which the deploy workflow runs on every push to main, strips it and fails the build if any of it, or a citation to it, survives. It returns only when the build is run with --manuscript-cleared, after the authors confirm that publication does not breach the conditions of review.3550 Where an ungated sentence rests on the manuscript, the site says it in its own words and keeps the citation. A source note that cites the manuscript by section and margin line is a reference, not content: it stays in the deployed page whenever an ungated sentence points to it, and is held out only when every marker that points to it is held out with its copy.

Manuscript under review

One manuscript describing the lineage-first fabric, the QuestionArtifact, continual discovery and the four evidence classes is under review.35 The venue is not named and the manuscript is not distributed from this site while review is in progress; the site describes its content in the lab’s own words and marks every simulation figure as “reported in the manuscript”. After the review decision this section will link the paper or its preprint and the corrections log will record the date.


Artifact availability by evidence class

Note first: source paths on this site cite the NeuralGraph repository, which is public to read at github.com/anovruzov/NeuralGraph but has no licence, so all rights are reserved and nothing in it is offered for reuse.3648 “An auditor can check every sentence” therefore holds for any reader: the rows below link to the commit or branch each artifact is recorded at, and a request to the lab is needed only for what is not in the repository.

What exists for each evidence class, where it is kept, and how a reader obtains access.
Class What exists Where Access
A LoCoMo retrieval Per-question answer and retrieval artifacts for every experiment; the 18-experiment record37 docs/research/RESULTS_ALL.md and demo/results/ at commit e054178, on main4948 Results document →
Per-question JSON →
B Repair fixture Four JSON artifacts, SHA256-pinned, regenerate byte-identically (re-verified 2026-09-06, macOS arm64, Python 3.11); 277 tests pass; at commit 13a1729 on a research branch38 NeuralGraph/coordination/artifacts/ at commit 13a1729 on branch claude/session-analysis-continuation-sm1a1t, not merged into main3848 Artifacts and SHA256SUMS →
T Coordination tests 109 deterministic tests including a real-storage restart test12 NeuralGraph/tests/test_coordination.py, test_coordination_lineage.py, test_coordination_storage.py, test_coordination_experiment_restart.py, on main12 Test modules →
C Acquisition probe Reported in the manuscript; artifacts not in the current checkout15 Not in the repository After review; ask the lab →
D Population simulation Reported in the manuscript, whose current revision is based on the campaign’s written reports rather than on re-running its code15 Not in the repository After review; ask the lab →
IP Organisational benchmark results/processed and manifests, in progress; every manifest states the model-backend limitation39 mycelic-org-benchmark/results/ on branch claude/mycelic-org-benchmark-9cnf35, in progress, not merged into main3948 Design and code →
Results: when the campaign finishes, dated in the corrections log

Known discrepancies and how this site resolves them

  1. Class B figures. The repository’s benchmark record gives random path diversification 0.704 survival and 5,327 bytes.40 The manuscript gives 6/9 and 5,206 bytes for the same strategy.41 The manuscript reports a different figure for the same strategy; the comparison is held back while review is in progress. This site uses the repository record (bitwise reproducible) throughout and says so.
  2. LoCoMo labels. Inherited category constants swap single-hop and multi-hop; all tables say “as labelled”; a corrected mapping is pending in the repository.31
  3. Two ladders. The manuscript’s B0 to B7/B3Q and the benchmark’s B0 to B9 plus ORACLE share ids; the site always prefixes.30
  4. Two Tesseracts. The repository contains an older single-node retrieval module and the cross-node coordinator, both called Tesseract; this site means the coordinator unless it says otherwise.42
  5. Test counts. The coordination suite was 107 tests before a real-storage restart test added two; the site cites 109.43
  6. Default model. The README lists an Ollama default of qwen2.5:7b-instruct while the code’s default is an OpenAI-compatible endpoint at 127.0.0.1:1234 with google/gemma-4-e4b; the site therefore states no default model and treats model names as configuration.4445
  7. Gossip baseline. Table 3 gives manuscript-B4 knowledge survival .749 at 48.437 replicas per claim under the 50% white-box attack;46 the prose (§5.6 ll.206-207) gives B4 .669 at the same storage in the sentence that follows the 90%-churn comparison.47 We read .669 as the 90%-churn value; the manuscript does not label it explicitly, so the site shows neither as a headline and records the reading here.

Every future discrepancy is added here and dated in the corrections log. Corrections log →

Reproduce a result or ask about an artifact

Use the subject “Research collaboration / artifact access”.

Sources

Each marker above points to one note below. Paths are relative to the root of the NeuralGraph repository, which is public to read at github.com/anovruzov/NeuralGraph with no licence selected (all rights reserved); they are given as text, with line numbers checked against main on 2026-09-17, and files under mycelic-org-benchmark/ are on branch claude/mycelic-org-benchmark-9cnf35, not yet merged into main, and the Class B artifacts are at commit 13a1729 on branch claude/session-analysis-continuation-sm1a1t, not merged into main; the Artifacts table above links those directly. “Manuscript” is the paper under review, cited by section and printed margin line. Notes that support only manuscript-derived copy are held out of the deployed page together with that copy; a note cited by an ungated sentence stays, since a section-and-line reference reproduces nothing from the paper.

  1. NeuralGraph/state.md — Mission ll.11-17
  2. NeuralGraph/mycelic-org-benchmark/docs/REPORT.md — Research question ll.37-40
  3. Manuscript — §1 ll.30-33
  4. Manuscript — §4.1 ll.70-74
  5. Manuscript — §6.2 ll.244-257
  6. NeuralGraph/mycelic-org-benchmark/DESIGN.md — §4 ll.160-192
  7. NeuralGraph/docs/BENCHMARKS.md — Track B header ll.70-72
  8. NeuralGraph/docs/BENCHMARKS.md — Track B ll.70-99
  9. NeuralGraph/docs/BENCHMARKS.md — Track A ll.16-24
  10. NeuralGraph/docs/BENCHMARKS.md — Track A ll.21-23
  11. NeuralGraph/state.md — Gate 1 ll.171-200
  12. NeuralGraph/state.md — Test Status ll.145-154
  13. NeuralGraph/state.md — Gate 1 ll.185-187
  14. Manuscript — §5.1 ll.76-80
  15. Manuscript — §6.1 ll.239-241
  16. Manuscript — §5.3 ll.126-137
  17. NeuralGraph/mycelic-org-benchmark/README.md — ll.1-22
  18. NeuralGraph/mycelic-org-benchmark/docs/REPORT.md — ll.171-180
  19. NeuralGraph/docs/BENCHMARKS.md — Track A ll.16-62
  20. NeuralGraph/state.md — Gate Status ll.163-228
  21. NeuralGraph/docs/BENCHMARKS.md — ll.88-131
  22. NeuralGraph/docs/research/REPORT.md — §2 ll.28-31
  23. NeuralGraph/docs/BENCHMARKS.md — All 18 experiments ll.101-119
  24. NeuralGraph/state.md — Gate 1 ll.171-209
  25. Manuscript — §5.3-5.6 ll.126-211
  26. Manuscript — §6.3 ll.258-264
  27. Manuscript — §6.1 ll.231-232
  28. Manuscript — §4 ll.62-66
  29. Manuscript — Table 2 p.5
  30. NeuralGraph/mycelic-org-benchmark/README.md — Systems ll.139-153
  31. NeuralGraph/docs/BENCHMARKS.md — header note ll.7-12
  32. NeuralGraph/docs/BENCHMARKS.md — Track C ll.147-162
  33. Manuscript — §5.6 ll.199-204
  34. Manuscript — §6.1 l.242
  35. Manuscript — p.1 footer (unnumbered)
  36. NeuralGraph/README.md — License ll.145-147
  37. NeuralGraph/docs/BENCHMARKS.md — Track B ll.70-119
  38. NeuralGraph/docs/BENCHMARKS.md — Track A ll.16-23
  39. NeuralGraph/mycelic-org-benchmark/README.md — Limitation statement ll.31-61
  40. NeuralGraph/docs/BENCHMARKS.md — Track A ll.27-49
  41. Manuscript — §5.2 ll.110-117
  42. NeuralGraph/state.md — Current Architecture ll.32-44
  43. NeuralGraph/state.md — Test Status ll.133-154
  44. NeuralGraph/README.md — Local Models ll.90-98
  45. NeuralGraph/NeuralGraph/llm_backend.py — ll.7-24
  46. Manuscript — Table 3 p.6
  47. Manuscript — §5.6 ll.206-207
  48. Repository check, 2026-09-17: GET api.github.com/repos/anovruzov/NeuralGraph returned private: false, visibility: public, default branch main, license: null; git merge-base --is-ancestor on a fresh fetch puts commit e054178 on main and commit 13a1729 and branch claude/mycelic-org-benchmark-9cnf35 off it
  49. NeuralGraph/docs/BENCHMARKS.md — Track B source lines ll.68-69
  50. This site’s own repository, not NeuralGraph: tools/build.py — resolves the gate; its --check fails the build if a gate marker, the visible text of a stripped manuscript block, a citation to a stripped source note or a withdrawn literal survives in any built page; .github/workflows/pages.yml — runs python3 tools/build.py --check on every push to main and publishes dist/; README.md — Deploy (Pages must publish from GitHub Actions, never from the branch root)
  51. Manuscript — §5.4 ll.145-147

Files cited on this page

  1. NeuralGraph/state.md
  2. NeuralGraph/docs/BENCHMARKS.md
  3. NeuralGraph/docs/research/REPORT.md
  4. NeuralGraph/README.md
  5. NeuralGraph/NeuralGraph/llm_backend.py
  6. NeuralGraph/mycelic-org-benchmark/README.md, DESIGN.md, docs/REPORT.md
  7. Manuscript — the PDF under review; its content is gated, its section-and-line references are not
  8. Repository metadata: api.github.com/repos/anovruzov/NeuralGraph and git merge-base, checked 2026-09-17
  9. This site’s repository: tools/build.py, .github/workflows/pages.yml, README.md