Claims register
Every claim on this site, with its source
If a page and this register disagree, the register wins. Each sidenote superscript on the site resolves to a row here. Verified-on is the date the source was last checked against the row (all rows 2026-09-17).1
How numbers are cited
Every figure carries a superscript that opens a sidenote (or, on narrow screens and with JavaScript off, a native details element) listing: evidence class, judge or grader, dataset slice and n, leakage status, seeds, and the source document with section. Sources cite the NeuralGraph repository at a stated document revision; it is public to read at github.com/anovruzov/NeuralGraph but has no licence, so all rights are reserved.27 Any reader can therefore check every sentence; only artifacts that are not in the repository (Classes C and D) need a request to the lab. Readers who find a figure without a sidenote are asked to report it: Contact (subject: Correction).
Register
One row per claim. The Id is the anchor every sidenote links to; Source is the file and lines in the NeuralGraph repository (public to read; see Sources below for the branch each file is on); Class is the evidence class from the key below. Rows marked gated hold figures from the manuscript under review: while the gate is closed the deployed page shows the row with its source and a withheld notice in place of the figure.
- A measured retrieval
- B deterministic fixture
- C synthetic acquisition simulation
- D population-scale symbolic simulation
- T deterministic coordination tests
- IP in progress, no results
- W withdrawn
- — no evidence class: status, licence or configuration
| Id | Claim | Page | Source | Class | Judge / cohort | Caveat | Verified |
|---|---|---|---|---|---|---|---|
| R1 | +8.9 points single-hop accuracy (as labelled), flat 64.9% → local_pairs 73.8% | index, benchmarks | NeuralGraph/ |
A measured retrieval | Gemma lenient; 282 questions; leakage-free | survives four judges (+11.3 / +9.6 / +4.6; p < .001 for the three LLM judges, p = .01 for substring, two-sided sign tests); routing and pair back-fill measured together | |
| R2 | lineage-aware repair 7/9 (0.778) vs full replication 6/9 (0.667) | index, benchmarks | NeuralGraph/ |
B deterministic fixture | 5-node fixture, 30 seeds | mechanism demonstration on a small-N simulator, not a performance claim; 0.00 under partition and worst-single-domain failure where full replication survives partition (1.00) | |
| R3 | same 282 answers score 13.5 / 21.6 / 51.4 / 64.9 by grader | index, benchmarks | NeuralGraph/ |
A measured retrieval | substring / Qwen strict / Qwen lenient / Gemma lenient | κ 0.17 / 0.30 / 0.72; not a gain | |
| R4 | 72.2% overall on 744 of 1,540 questions | benchmarks | NeuralGraph/ |
A measured retrieval | Gemma lenient; conversations 1 to 5; leakage-free | conversations 6 to 10 and adversarial unrun; judge-dependent | |
| R5 | 584-question subset: 73.8 / 65.6 / 42.5 / 26.7 by grader | benchmarks | NeuralGraph/ |
A measured retrieval | conversations 1 to 4 | not a replacement for R4 | |
| R6 | recall@10 single-hop 39.4 → 46.8 with per-agent routing; matches oracle; gold spoken by the other agent 2.9% | benchmarks | NeuralGraph/ |
A measured retrieval | 282 questions | substring recall proxy | |
| R7 | own name in 0.1% of a speaker’s messages, other speaker’s in 33.4%; name swap 26.6 → 34.8 | benchmarks | NeuralGraph/ |
A measured retrieval | LoCoMo two-speaker | corpus statistic | |
| R8 | matched-budget: neighbours +1.4, embedding back-fill +4.1, pairs +3.2; neighbours in top-50 drop recall@50 48.0 → 32.9 | benchmarks | NeuralGraph/ |
A measured retrieval | 30-candidate budget | regex-entity edges, one dataset | |
| R9 | listwise reranker −4 accuracy for −56% rerank latency | benchmarks | NeuralGraph/ |
A measured retrieval | Gemma lenient | speed only | |
| R10 | context 15/30 → 50/50: gold presence 71 → 73%, accuracy 78 → 69%, accuracy-given-gold 88.7 → 78.1% | benchmarks | NeuralGraph/ |
A measured retrieval | 100 matched questions | small subset | |
| R11 | fusion: recall@10 46.8 → 53.7 (1,219 q); single-hop 78 → 68 (3 wins, 13 losses); multi-hop 75 → 76 | benchmarks | NeuralGraph/ |
A measured retrieval | matched subsets | candidate dilution, not a law | |
| R12 | gold in context → right 88.7%; gold absent 29%; 25.5% of gold answers not a substring of any message | benchmarks | NeuralGraph/ |
A measured retrieval | single-hop as labelled | diagnostic | |
| R13 | harness latency 6.4 to 8.2 s per question | index, benchmarks | NeuralGraph/ |
A measured retrieval | one local machine, one local model | not a product specification | |
| R14 | December 2025 66.7% / 66.9% (superseded) | benchmarks | NeuralGraph/ |
W withdrawn | GPT-4o lenient; two leakage channels | comparison only | |
| R15 | Gate 0 and Gate 1 PASS; Gates 2 to 6 NOT STARTED; 109 tests | index, benchmarks, docs | NeuralGraph/ |
T deterministic coordination tests | deterministic fixtures | no baselines, no scale evidence | |
| R16 | Class B invariants (root failure 0.00 vs 1.00; only oracle survives worst domain; one lost node-failure cell) | benchmarks | NeuralGraph/ |
B deterministic fixture | 30 seeds | small-N simulator | |
| R17 | Class B bytes 5,448 (lineage-aware) vs 9,877 (full replication) | benchmarks | NeuralGraph/ |
B deterministic fixture | serialised transfer | not resident memory | |
| R18 | Class B artifacts bitwise reproducible; 277 tests; re-verified 2026-09-06 | research | NeuralGraph/ |
B deterministic fixture | macOS arm64, Python 3.11 | commit 13a1729 on a research branch | |
| R19 | REDACTED policy fails closed to DENIED; no per-source-memory authorization | docs | NeuralGraph/ |
T deterministic coordination tests | — | documented limitation | |
| R20 | roadmap items unchecked: benchmarks publication, quick-start, API, visualisation, providers, namespaces, packaging | index, architecture, docs | NeuralGraph/ |
— | — | planned | |
| R21 | no licence selected; all rights reserved | site-wide | NeuralGraph/ |
— | — | — | |
| R22 | engine backend defaults: OpenAI-compatible at 127.0.0.1:1234, google/gemma-4-e4b; Ollama when port 11434 | architecture, docs | NeuralGraph/ |
— | — | README lists a different default; site states none | |
| R23 | org benchmark: in progress, no results; simulated model profiles | index, benchmarks | NeuralGraph/ |
IP in progress, no results | — | no number may be shown | |
| R24 | org benchmark: minimum discovery layer α = 1e-8, power 0.8, 1.25 × n_min | benchmarks | NeuralGraph/ |
IP in progress, no results | — | design | |
| R25 | org benchmark: support discount 0 / 0 / 0.5 / 0.8 / 1.0 | architecture, benchmarks | NeuralGraph/ |
IP in progress, no results | — | design | |
| R26 | org benchmark tiers and simulator cost (18 to 25 s at 0.35 GB; 310 s + 150 s at 3.4 GB) | benchmarks, docs | NeuralGraph/ |
IP in progress, no results | 4-vCPU, no GPU | simulator cost | |
| R27 | manuscript under review; content gated | research, benchmarks | paper# |
— | — | venue not named | |
| R28 gated | Class D equal-storage differences +.000 / +.010 / +.046 / +.155 / +.003 | benchmarks | paper# |
D population-scale symbolic simulation | N = 10,000, 30 seeds, 50% white-box attack | published after review | |
| R28 gated | Under review Class D equal-storage result, withheld while the manuscript is under review; published here after the review decision, with the date recorded in the corrections log. | benchmarks | paper# |
D population-scale symbolic simulation | — | manuscript gate closed | |
| R29 gated | Class D questioning B7 vs B6 gains and 1.81× messages; B3Q .628 vs B7 .594 | benchmarks | paper# |
D population-scale symbolic simulation | N = 10,000, 30 seeds | not lineage-specific | |
| R29 gated | Under review Class D questioning-cost result, withheld while the manuscript is under review; published here after the review decision, with the date recorded in the corrections log. | benchmarks | paper# |
D population-scale symbolic simulation | — | manuscript gate closed | |
| R30 gated | Class C 0.871 vs 0.834; worse calibration in 15/23 cells | benchmarks | paper# |
C synthetic acquisition simulation | 10 seeds × 2,000 claims | artifacts absent | |
| R30 gated | Under review Class C acquisition result, withheld while the manuscript is under review; published here after the review decision, with the date recorded in the corrections log. | benchmarks | paper# |
C synthetic acquisition simulation | — | manuscript gate closed | |
| R31 gated | with vs without reranking 68.6% vs 61.9%, 8.2 s vs 0.44 s | benchmarks | paper# |
A measured retrieval | rounded values | not an exact matched effect | |
| R31 gated | Under review With-versus-without-reranking accuracy and latency pair as reported in the manuscript, withheld while it is under review; published here after the review decision, with the date recorded in the corrections log. | benchmarks | paper# |
A measured retrieval | — | manuscript gate closed |
Withdrawn claims from the previous site
| Old wording | Where it was | Why withdrawn | Contradicting or missing source |
|---|---|---|---|
| LoCoMo Benchmark 72% Accuracy (bare) | index.html hero and specifications | no judge, cohort or leakage statement | NeuralGraph/ |
| Retrieval latency <50 ms p95; end-to-end ~1.3 s; embedding ~100 to 200 ms; Sub-50 ms Retrieval | index.html ll.66-73, 152-153, 212-222 | unsourced; measured harness latency is 6.4 to 8.2 s per question | NeuralGraph/ |
| Memory architecture with zero external transmission | index.html l.41 | a guarantee; the engine contacts a configured model endpoint and has an optional open-domain fallback | NeuralGraph/ |
| No centralised index / without explicit indexing; wave-based resonance retrieval as the mechanism | index.html | Stage 1 is vector plus keyword search; wavefront is one reranking signal | NeuralGraph/ |
| pip install mycelic; docker pull mycelic/mycelic; Python SDK; REST /v1 API; mycelic.yaml; all-MiniLM-L6-v2 | docs/ |
no package, image, SDK or API exists | NeuralGraph/ |
| github.com/mycelic organisation and issue links | all footers, contact.html | organisation does not exist in any source | website git remote github.com/ |
| Community $10/month, Professional $25/month, Enterprise; Apache 2.0; open source core | pricing.html ll.40-57, 200 | no pricing exists; no licence selected | NeuralGraph/ |
| Plan: Community (Free) | dashboard.html | contradicts pricing page; no accounts exist | website/ |
| Min RAM 8 GB / 32 GB+; storage; CPU cores | index.html ll.234-241 | no hardware measurement exists | NeuralGraph/ |
| Cross-source aggregation across documents, code and structured data | index.html l.142 | not in any source | NeuralGraph/ |
| Export, audit or purge on demand | index.html l.148 | no such feature; retention is automatic consolidation | NeuralGraph/ |
| Graceful degradation: performance scales linearly with available resources | about.html l.98 | unmeasured | — |
| Enterprise Ready | about.html l.72 | no audit, certification or deployment | NeuralGraph/ |
| Competitive accuracy with cloud alternatives | about.html | no cross-system comparison is defensible under a 51-point judge spread | NeuralGraph/ |
| contact@mycelic.ai | contact.html l.88 | domain appears in no source | — |
| Login, dashboard, Express/JWT auth server | login.html, dashboard.html, server/ |
credentials collected for accounts that did not exist | website/ |
Corrections and changelog
Dated entries, oldest first. Every future correction, results publication and gate change is appended here.
-
LoCoMo 66.7% overall / 66.9% on the 744-question cohort under a GPT-4o lenient judge with two leakage channels (gold-answer acceptance gate; gold-category routing). Superseded; comparison only.3
-
to
The previous website’s headline claims listed in the withdrawn table above were unsourced or contradicted by measurement and are withdrawn.4
-
Repository audit compiled the benchmark record with source commits and evidence classes; Class B artifacts re-verified bitwise.5
-
Site relaunched with this register. Manuscript gate: CLOSED (Class C and D tables, verbatim manuscript quotations and formal definitions are omitted from the deployed HTML until the authors clear publication). Organisational benchmark: in progress, no results. Coordination gates: 0 and 1 passed, 2 to 6 not started. Licence: not selected. Repository: public to read at github.com/anovruzov/NeuralGraph, no licence selected.7
-
future
Manuscript gate opened / paper linked: date to be recorded here.
-
future
Organisational benchmark results published from results/processed with manifests: date to be recorded here.6
Found a figure without a source note, or a row that disagrees with a page?
Write to us with the subject “Correction to a number on the site”.
Sources
Each marker above points to one note below. Paths are relative to the root of the NeuralGraph repository, which is public to read at github.com/anovruzov/NeuralGraph with no licence selected (all rights reserved); they are given as text, with line numbers checked against main on 2026-09-17, and files under mycelic-org-benchmark/ are on branch claude/mycelic-org-benchmark-9cnf35, not yet merged into main, and the Class B artifacts are at commit 13a1729 on branch claude/session-analysis-continuation-sm1a1t, not merged into main. “paper” is the manuscript under review, cited by table or printed margin line; “website/” is the previous site.
- NeuralGraph/docs/BENCHMARKS.md — header ll.3-5 (compiled 2026-09-06 during the repository audit)
- NeuralGraph/README.md — License ll.145-147
- NeuralGraph/docs/BENCHMARKS.md — Track C, the December baseline ll.147-162
- website/*.html — the pages of the previous site
- NeuralGraph/docs/BENCHMARKS.md — ll.3-5, 21-23
- NeuralGraph/mycelic-org-benchmark/README.md — l.118
- Repository check, 2026-09-17: GET api.github.com/repos/anovruzov/NeuralGraph returned private: false, visibility: public, default branch main, license: null; git merge-base --is-ancestor on a fresh fetch puts commit e054178 on main and commit 13a1729 and branch claude/mycelic-org-benchmark-9cnf35 off it
Files cited in the tables on this page
- NeuralGraph/docs/BENCHMARKS.md
- NeuralGraph/docs/research/REPORT.md
- NeuralGraph/state.md
- NeuralGraph/README.md
- NeuralGraph/NeuralGraph/llm_backend.py, NeuralGraph/NeuralGraph/external_retriever.py
- NeuralGraph/NeuralGraph_System_Architecture.md
- NeuralGraph/mycelic-org-benchmark/README.md, DESIGN.md, docs/REPORT.md, docs/EXPERIMENT_PLAN.md
- paper — the manuscript under review, cited by table or printed margin line; gated
- website/index.html, website/about.html, website/contact.html, website/pricing.html, website/login.html, website/dashboard.html, website/docs/getting-started.html, website/docs/api-reference.html, website/server/server.js — the previous site, and its git remote github.com/anovruzov/website