Claims register

Every claim on this site, with its source

If a page and this register disagree, the register wins. Each sidenote superscript on the site resolves to a row here. Verified-on is the date the source was last checked against the row (all rows 2026-09-17).1


How numbers are cited

Every figure carries a superscript that opens a sidenote (or, on narrow screens and with JavaScript off, a native details element) listing: evidence class, judge or grader, dataset slice and n, leakage status, seeds, and the source document with section. Sources cite the NeuralGraph repository at a stated document revision; it is public to read at github.com/anovruzov/NeuralGraph but has no licence, so all rights are reserved.27 Any reader can therefore check every sentence; only artifacts that are not in the repository (Classes C and D) need a request to the lab. Readers who find a figure without a sidenote are asked to report it: Contact (subject: Correction).


Register

One row per claim. The Id is the anchor every sidenote links to; Source is the file and lines in the NeuralGraph repository (public to read; see Sources below for the branch each file is on); Class is the evidence class from the key below. Rows marked gated hold figures from the manuscript under review: while the gate is closed the deployed page shows the row with its source and a withheld notice in place of the figure.

Every quantitative and capability claim on the site: the claim as worded, the pages that carry it, its source, evidence class, judge and cohort, the caveat that travels with it, and the date the source was last checked against the row.
Id Claim Page Source Class Judge / cohort Caveat Verified
R1 +8.9 points single-hop accuracy (as labelled), flat 64.9% → local_pairs 73.8% index, benchmarks NeuralGraph/docs/BENCHMARKS.md#ll.88-99; research/N5_judge_sensitivity.md#ll.20-23 A measured retrieval Gemma lenient; 282 questions; leakage-free survives four judges (+11.3 / +9.6 / +4.6; p < .001 for the three LLM judges, p = .01 for substring, two-sided sign tests); routing and pair back-fill measured together
R2 lineage-aware repair 7/9 (0.778) vs full replication 6/9 (0.667) index, benchmarks NeuralGraph/docs/BENCHMARKS.md#ll.16-62 B deterministic fixture 5-node fixture, 30 seeds mechanism demonstration on a small-N simulator, not a performance claim; 0.00 under partition and worst-single-domain failure where full replication survives partition (1.00)
R3 same 282 answers score 13.5 / 21.6 / 51.4 / 64.9 by grader index, benchmarks NeuralGraph/docs/BENCHMARKS.md#ll.121-131 A measured retrieval substring / Qwen strict / Qwen lenient / Gemma lenient κ 0.17 / 0.30 / 0.72; not a gain
R4 72.2% overall on 744 of 1,540 questions benchmarks NeuralGraph/docs/BENCHMARKS.md#ll.74-86 A measured retrieval Gemma lenient; conversations 1 to 5; leakage-free conversations 6 to 10 and adversarial unrun; judge-dependent
R5 584-question subset: 73.8 / 65.6 / 42.5 / 26.7 by grader benchmarks NeuralGraph/docs/research/REPORT.md#ll.9-15 A measured retrieval conversations 1 to 4 not a replacement for R4
R6 recall@10 single-hop 39.4 → 46.8 with per-agent routing; matches oracle; gold spoken by the other agent 2.9% benchmarks NeuralGraph/docs/research/REPORT.md#l.30; BENCHMARKS.md#ll.135-138 A measured retrieval 282 questions substring recall proxy
R7 own name in 0.1% of a speaker’s messages, other speaker’s in 33.4%; name swap 26.6 → 34.8 benchmarks NeuralGraph/docs/research/REPORT.md#l.30 A measured retrieval LoCoMo two-speaker corpus statistic
R8 matched-budget: neighbours +1.4, embedding back-fill +4.1, pairs +3.2; neighbours in top-50 drop recall@50 48.0 → 32.9 benchmarks NeuralGraph/docs/research/REPORT.md#ll.31, 34-48 A measured retrieval 30-candidate budget regex-entity edges, one dataset
R9 listwise reranker −4 accuracy for −56% rerank latency benchmarks NeuralGraph/docs/BENCHMARKS.md#l.110 A measured retrieval Gemma lenient speed only
R10 context 15/30 → 50/50: gold presence 71 → 73%, accuracy 78 → 69%, accuracy-given-gold 88.7 → 78.1% benchmarks NeuralGraph/docs/research/REPORT.md#ll.40-41 A measured retrieval 100 matched questions small subset
R11 fusion: recall@10 46.8 → 53.7 (1,219 q); single-hop 78 → 68 (3 wins, 13 losses); multi-hop 75 → 76 benchmarks NeuralGraph/docs/research/REPORT.md#ll.83-86 A measured retrieval matched subsets candidate dilution, not a law
R12 gold in context → right 88.7%; gold absent 29%; 25.5% of gold answers not a substring of any message benchmarks NeuralGraph/docs/research/REPORT.md#ll.48, 63 A measured retrieval single-hop as labelled diagnostic
R13 harness latency 6.4 to 8.2 s per question index, benchmarks NeuralGraph/docs/BENCHMARKS.md#ll.90-95 A measured retrieval one local machine, one local model not a product specification
R14 December 2025 66.7% / 66.9% (superseded) benchmarks NeuralGraph/docs/BENCHMARKS.md#ll.147-162 W withdrawn GPT-4o lenient; two leakage channels comparison only
R15 Gate 0 and Gate 1 PASS; Gates 2 to 6 NOT STARTED; 109 tests index, benchmarks, docs NeuralGraph/state.md#ll.145-154, 163-228 T deterministic coordination tests deterministic fixtures no baselines, no scale evidence
R16 Class B invariants (root failure 0.00 vs 1.00; only oracle survives worst domain; one lost node-failure cell) benchmarks NeuralGraph/docs/BENCHMARKS.md#ll.57-62 B deterministic fixture 30 seeds small-N simulator
R17 Class B bytes 5,448 (lineage-aware) vs 9,877 (full replication) benchmarks NeuralGraph/docs/BENCHMARKS.md#ll.39-49 B deterministic fixture serialised transfer not resident memory
R18 Class B artifacts bitwise reproducible; 277 tests; re-verified 2026-09-06 research NeuralGraph/docs/BENCHMARKS.md#ll.16-23 B deterministic fixture macOS arm64, Python 3.11 commit 13a1729 on a research branch
R19 REDACTED policy fails closed to DENIED; no per-source-memory authorization docs NeuralGraph/state.md#ll.270-278 T deterministic coordination tests documented limitation
R20 roadmap items unchecked: benchmarks publication, quick-start, API, visualisation, providers, namespaces, packaging index, architecture, docs NeuralGraph/README.md#ll.131-139 planned
R21 no licence selected; all rights reserved site-wide NeuralGraph/README.md#ll.145-147
R22 engine backend defaults: OpenAI-compatible at 127.0.0.1:1234, google/gemma-4-e4b; Ollama when port 11434 architecture, docs NeuralGraph/NeuralGraph/llm_backend.py#ll.7-24 README lists a different default; site states none
R23 org benchmark: in progress, no results; simulated model profiles index, benchmarks NeuralGraph/mycelic-org-benchmark/README.md#ll.31-61; docs/REPORT.md#ll.171-180 IP in progress, no results no number may be shown
R24 org benchmark: minimum discovery layer α = 1e-8, power 0.8, 1.25 × n_min benchmarks NeuralGraph/mycelic-org-benchmark/DESIGN.md#§3.5 ll.115-143 IP in progress, no results design
R25 org benchmark: support discount 0 / 0 / 0.5 / 0.8 / 1.0 architecture, benchmarks NeuralGraph/mycelic-org-benchmark/DESIGN.md#§6 step 3 ll.238-242 IP in progress, no results design
R26 org benchmark tiers and simulator cost (18 to 25 s at 0.35 GB; 310 s + 150 s at 3.4 GB) benchmarks, docs NeuralGraph/mycelic-org-benchmark/DESIGN.md#§14; docs/EXPERIMENT_PLAN.md#ll.3-14 IP in progress, no results 4-vCPU, no GPU simulator cost
R27 manuscript under review; content gated research, benchmarks paper#p.1 footer (unnumbered) venue not named
R28 gated Class D equal-storage differences +.000 / +.010 / +.046 / +.155 / +.003 benchmarks paper#Table 4 D population-scale symbolic simulation N = 10,000, 30 seeds, 50% white-box attack published after review
R28 gated Under review Class D equal-storage result, withheld while the manuscript is under review; published here after the review decision, with the date recorded in the corrections log. benchmarks paper#Table 4 D population-scale symbolic simulation manuscript gate closed
R29 gated Class D questioning B7 vs B6 gains and 1.81× messages; B3Q .628 vs B7 .594 benchmarks paper#Table 5, ll.163-173 D population-scale symbolic simulation N = 10,000, 30 seeds not lineage-specific
R29 gated Under review Class D questioning-cost result, withheld while the manuscript is under review; published here after the review decision, with the date recorded in the corrections log. benchmarks paper#Table 5, ll.163-173 D population-scale symbolic simulation manuscript gate closed
R30 gated Class C 0.871 vs 0.834; worse calibration in 15/23 cells benchmarks paper#§5.2 ll.118-125 C synthetic acquisition simulation 10 seeds × 2,000 claims artifacts absent
R30 gated Under review Class C acquisition result, withheld while the manuscript is under review; published here after the review decision, with the date recorded in the corrections log. benchmarks paper#§5.2 ll.118-125 C synthetic acquisition simulation manuscript gate closed
R31 gated with vs without reranking 68.6% vs 61.9%, 8.2 s vs 0.44 s benchmarks paper#ll.105-108 A measured retrieval rounded values not an exact matched effect
R31 gated Under review With-versus-without-reranking accuracy and latency pair as reported in the manuscript, withheld while it is under review; published here after the review decision, with the date recorded in the corrections log. benchmarks paper#ll.105-108 A measured retrieval manuscript gate closed

Withdrawn claims from the previous site

Each claim of the previous website that is withdrawn, with its old wording, where it appeared, why it is withdrawn, and the source that contradicts it or the absence of any source.
Old wording Where it was Why withdrawn Contradicting or missing source
LoCoMo Benchmark 72% Accuracy (bare) index.html hero and specifications no judge, cohort or leakage statement NeuralGraph/docs/BENCHMARKS.md#ll.74-86, 121-131
Retrieval latency <50 ms p95; end-to-end ~1.3 s; embedding ~100 to 200 ms; Sub-50 ms Retrieval index.html ll.66-73, 152-153, 212-222 unsourced; measured harness latency is 6.4 to 8.2 s per question NeuralGraph/docs/BENCHMARKS.md#ll.90-95
Memory architecture with zero external transmission index.html l.41 a guarantee; the engine contacts a configured model endpoint and has an optional open-domain fallback NeuralGraph/NeuralGraph/llm_backend.py; external_retriever.py#l.249; paper#l.242
No centralised index / without explicit indexing; wave-based resonance retrieval as the mechanism index.html Stage 1 is vector plus keyword search; wavefront is one reranking signal NeuralGraph/NeuralGraph_System_Architecture.md#§3.3.1, §5.2; NeuralGraph/docs/BENCHMARKS.md#ll.88-99
pip install mycelic; docker pull mycelic/mycelic; Python SDK; REST /v1 API; mycelic.yaml; all-MiniLM-L6-v2 docs/getting-started.html, docs/api-reference.html no package, image, SDK or API exists NeuralGraph/README.md#ll.131-139
github.com/mycelic organisation and issue links all footers, contact.html organisation does not exist in any source website git remote github.com/anovruzov/website; the engine lives in the user repository github.com/anovruzov/NeuralGraph, not an organisation7
Community $10/month, Professional $25/month, Enterprise; Apache 2.0; open source core pricing.html ll.40-57, 200 no pricing exists; no licence selected NeuralGraph/README.md#ll.145-147
Plan: Community (Free) dashboard.html contradicts pricing page; no accounts exist website/pricing.html
Min RAM 8 GB / 32 GB+; storage; CPU cores index.html ll.234-241 no hardware measurement exists NeuralGraph/README.md (none)
Cross-source aggregation across documents, code and structured data index.html l.142 not in any source NeuralGraph/README.md#ll.108-116
Export, audit or purge on demand index.html l.148 no such feature; retention is automatic consolidation NeuralGraph/NeuralGraph_System_Architecture.md#§3.4.3
Graceful degradation: performance scales linearly with available resources about.html l.98 unmeasured
Enterprise Ready about.html l.72 no audit, certification or deployment NeuralGraph/state.md#ll.210-228
Competitive accuracy with cloud alternatives about.html no cross-system comparison is defensible under a 51-point judge spread NeuralGraph/docs/BENCHMARKS.md#ll.121-131
contact@mycelic.ai contact.html l.88 domain appears in no source
Login, dashboard, Express/JWT auth server login.html, dashboard.html, server/server.js credentials collected for accounts that did not exist website/server/server.js#l.12

Corrections and changelog

Dated entries, oldest first. Every future correction, results publication and gate change is appended here.

  1. LoCoMo 66.7% overall / 66.9% on the 744-question cohort under a GPT-4o lenient judge with two leakage channels (gold-answer acceptance gate; gold-category routing). Superseded; comparison only.3

  2. to

    The previous website’s headline claims listed in the withdrawn table above were unsourced or contradicted by measurement and are withdrawn.4

  3. Repository audit compiled the benchmark record with source commits and evidence classes; Class B artifacts re-verified bitwise.5

  4. Site relaunched with this register. Manuscript gate: CLOSED (Class C and D tables, verbatim manuscript quotations and formal definitions are omitted from the deployed HTML until the authors clear publication). Organisational benchmark: in progress, no results. Coordination gates: 0 and 1 passed, 2 to 6 not started. Licence: not selected. Repository: public to read at github.com/anovruzov/NeuralGraph, no licence selected.7

  5. future

    Manuscript gate opened / paper linked: date to be recorded here.

  6. future

    Organisational benchmark results published from results/processed with manifests: date to be recorded here.6

Found a figure without a source note, or a row that disagrees with a page?

Write to us with the subject “Correction to a number on the site”.

Sources

Each marker above points to one note below. Paths are relative to the root of the NeuralGraph repository, which is public to read at github.com/anovruzov/NeuralGraph with no licence selected (all rights reserved); they are given as text, with line numbers checked against main on 2026-09-17, and files under mycelic-org-benchmark/ are on branch claude/mycelic-org-benchmark-9cnf35, not yet merged into main, and the Class B artifacts are at commit 13a1729 on branch claude/session-analysis-continuation-sm1a1t, not merged into main. “paper” is the manuscript under review, cited by table or printed margin line; “website/” is the previous site.

  1. NeuralGraph/docs/BENCHMARKS.md — header ll.3-5 (compiled 2026-09-06 during the repository audit)
  2. NeuralGraph/README.md — License ll.145-147
  3. NeuralGraph/docs/BENCHMARKS.md — Track C, the December baseline ll.147-162
  4. website/*.html — the pages of the previous site
  5. NeuralGraph/docs/BENCHMARKS.md — ll.3-5, 21-23
  6. NeuralGraph/mycelic-org-benchmark/README.md — l.118
  7. Repository check, 2026-09-17: GET api.github.com/repos/anovruzov/NeuralGraph returned private: false, visibility: public, default branch main, license: null; git merge-base --is-ancestor on a fresh fetch puts commit e054178 on main and commit 13a1729 and branch claude/mycelic-org-benchmark-9cnf35 off it

Files cited in the tables on this page

  1. NeuralGraph/docs/BENCHMARKS.md
  2. NeuralGraph/docs/research/REPORT.md
  3. NeuralGraph/state.md
  4. NeuralGraph/README.md
  5. NeuralGraph/NeuralGraph/llm_backend.py, NeuralGraph/NeuralGraph/external_retriever.py
  6. NeuralGraph/NeuralGraph_System_Architecture.md
  7. NeuralGraph/mycelic-org-benchmark/README.md, DESIGN.md, docs/REPORT.md, docs/EXPERIMENT_PLAN.md
  8. paper — the manuscript under review, cited by table or printed margin line; gated
  9. website/index.html, website/about.html, website/contact.html, website/pricing.html, website/login.html, website/dashboard.html, website/docs/getting-started.html, website/docs/api-reference.html, website/server/server.js — the previous site, and its git remote github.com/anovruzov/website