Palace Evidence And Benchmarks
Source basis: Palace of Truth public PROJECT_STATUS.md, docs/benchmark-comparison-2026-05-09.md, docs/public-memory-benchmark-artifacts.md, semantic-memory fixture gates, and the v1.5 load report, last reviewed for this landing surface on July 9, 2026.
This page states the public evidence Palace can defensibly claim today. It is a ceiling for landing-page benchmark copy: broader leaderboard, consumer-readiness, or authority-grounding claims need stronger evidence first.
Defensible Claims
- Palace is public and self-hostable, with real local development, Helm/ArgoCD deployment patterns, AGPL source licensing, and major product surfaces beyond the original MVP plan. Source:
PROJECT_STATUS.md. - Palace exposes durable scoped memory through REST and MCP-compatible entry points, plus ingest, search, chat, feeds, export, graph, jobs, memory APIs, retrieval diagnostics, plugin packaging, and Palace control-plane surfaces. Source:
PROJECT_STATUS.md. - The current strongest evaluation claim is narrow: standalone Palace staging handled a realistic 1000-item NIST SP 800 corpus with ingest, memory jobs, Palace freshness, drained queues, exact retrieval, semantic search, and Palace-scoped retrieval passing. Source:
docs/benchmark-comparison-2026-05-09.md. - Report-only MemPalace and GBrain comparisons are directional adapter checks against checked-in sample artifacts and published target metadata. They are not full apples-to-apples benchmark reruns. Sources:
docs/benchmark-comparison-2026-05-09.md,docs/public-memory-benchmark-artifacts.md. - The checked-in semantic-memory fixture pack defines strict scope, provenance, temporal, mission, empty-result, multi-source, and budget gates. It is an executable contract for the current feature surface, not a public competitive benchmark result.
Current Palace Run
| Evidence | Current result | Source |
|---|---|---|
| Benchmark run id | 20260509-n1k-g | docs/benchmark-comparison-2026-05-09.md |
| Target environment | Standalone Palace staging, not a shared platform staging environment | docs/benchmark-comparison-2026-05-09.md |
| Corpus | Public NIST SP 800 excerpts | docs/benchmark-comparison-2026-05-09.md |
| Tagged items | 1000 | docs/benchmark-comparison-2026-05-09.md |
| Ready tagged items | 1000 | docs/benchmark-comparison-2026-05-09.md |
| Memory jobs | 1000/1000 complete | docs/benchmark-comparison-2026-05-09.md |
| Palace generation | 3412/3412, backlog 0 | docs/benchmark-comparison-2026-05-09.md |
| Queue state | Memory, relationship, dirty-marking, Palace build, and media queues drained | docs/benchmark-comparison-2026-05-09.md |
| Exact search / retrieve | 10 / 10 | docs/benchmark-comparison-2026-05-09.md |
| Search expected-hit ratio | 1.0 across 21 fixed NIST probes | PROJECT_STATUS.md, docs/benchmark-comparison-2026-05-09.md |
| Retrieve expected-hit ratio | 1.0 across 21 fixed NIST probes | PROJECT_STATUS.md, docs/benchmark-comparison-2026-05-09.md |
| Advisory authority-support slice | 12/12 cases passed | PROJECT_STATUS.md, docs/benchmark-comparison-2026-05-09.md |
Report-Only Comparisons
These rows are report-only sample checks. They can motivate further work, but they should not be presented as full competitive benchmark wins.
| Profile | Mode | Sample size | Palace metric | Published target metadata | Public claim status | Source |
|---|---|---|---|---|---|---|
| MemPalace LongMemEval-S Raw ChromaDB | Offline report-only sample | 2 cases | R@5 = 1.0 | R@5 = 0.966 | Directional only | docs/benchmark-comparison-2026-05-09.md, docs/public-memory-benchmark-artifacts.md |
| MemPalace MemBench Raw retrieval | Offline report-only sample | 2 cases | R@5 = 1.0 | R@5 = 0.803 | Directional only | docs/benchmark-comparison-2026-05-09.md, docs/public-memory-benchmark-artifacts.md |
| GBrain rich-prose corpus | Offline report-only sample | 2 cases | R@5 = 1.0 | R@5 = 0.979 | Directional only | docs/benchmark-comparison-2026-05-09.md, docs/public-memory-benchmark-artifacts.md |
| GBrain rich-prose corpus | Offline report-only sample | 2 cases | P@5 = 0.4166 | P@5 = 0.491 | Gap, not a win | docs/benchmark-comparison-2026-05-09.md, docs/public-memory-benchmark-artifacts.md |
Caveats
- Palace is not fully consumer-ready yet. Remaining work includes broader integration smoke coverage, manual retrieval-quality review on realistic corpora, sharper operator docs for self-hosters, and product decisions around deferred editor and grounding bets. Source:
PROJECT_STATUS.md. - The GBrain profile is based on a project-specific corpus, and the checked-in Palace adapter sample has only two cases. Source:
docs/benchmark-comparison-2026-05-09.md. - Precision remains a visible gap for the GBrain-style comparison:
P@5 = 0.4166versus target metadata of0.491. Source:docs/benchmark-comparison-2026-05-09.md. - The authority-support result is advisory evidence, not a full authority-grounding guarantee. Stronger claims require a defined target corpus and validation model. Sources:
PROJECT_STATUS.md,docs/benchmark-comparison-2026-05-09.md. - Benchmark cleanup remains a human-approved operation. This public evidence page does not authorize or instruct production or staging data deletion. Source:
docs/benchmark-comparison-2026-05-09.md. - Public repository availability is not itself a benchmark claim. Do not present the plugin tooling, compiler spine, or GHCR release flow as retrieval-quality evidence.
- Source-backed wakeup, OAuth, semantic memory, retention, reflection, and Hermes plugin releases are product capabilities. They do not expand the NIST run or report-only comparison results without a new reproducible evaluation.
Source Inventory
palaceoftruth/PROJECT_STATUS.mdpalaceoftruth/docs/benchmark-comparison-2026-05-09.mdpalaceoftruth/docs/public-memory-benchmark-artifacts.mdpalaceoftruth/backend/tests/fixtures/semantic_memory_v1_eval_pack.jsonpalaceoftruth/docs/research/sar-1044-semantic-memory-v15-load-report.md