Six items from fox's de-novo synthesis (2026-05-01) landing as one atomic commit. Each item references its motivation and respects the v9.8 honesty boundary (no claims of semantic truth, no proof-path contamination by soft signals). NEW DOCS -------- docs/cti-architecture.md Maps today's modules onto the architectural layer cake fox named in his synthesis: PROMETHEUS-Σ (controller) / CTI (claim-lattice reasoning IR) / Merkle-AGI-DAG (commitment) / Reverse-RAG (evidence direction) / v9.8 Providence (admissibility ledger) / Hermes (weak proposer). Per-layer responsibility table + module map showing which existing files implement which layer. The architecture is real; the labels above name what's already there. docs/naming-deferral.md Argues why we keep internal terms (claim_lattice, claim_lattice_pointer, verify_claim_lattice) instead of renaming to architectural labels (CTI, PROMETHEUS-Σ). The rename costs cache invalidation, ~150 test fixture references, schema CHECK migration, blame-history disconnect, mesh peer coordination. The bridge is the cti-architecture mapping doc — read it once, then read code in code's vocabulary and prose in prose's vocabulary. Lists four triggers that would invert the decision. docs/self-reference-distillation-design.md Successor to docs/self-reference-thought-chains-design.md (the flat-source MVP). Maps STRICT claims onto the existing Distiller / Core / derivations infrastructure: each STRICT record becomes a Fact-Core via a new ProvidenceDistiller, with per-cited-chunk Merkle inclusion proofs back to Wikipedia source_roots. The fact-graph grows; new claims attach by inheriting the proof chain. CompositionDistiller (combining facts into new claims) is explicitly deferred — that's reasoning machinery, not infrastructure. docs/test-coverage-audit-2026-05-01.md Maps fox's §11 test list (16 items) against the codebase. 16/16 covered. Documents item #9's manual-quote-rule lifecycle: the rule was retired from pointer mode in commit224bfd6after the bench journey identified it was rejecting factually correct claims; retained in JSON variant where the punctuation-rationale argument doesn't apply. The audit doc itself is the requirements- drift defense. CODE ---- aborist/qa/keys.py:verifier_policy_hash New pure function. Hashes the verifier-relevant subset of policy (answer_mode + claim_lattice_* verifier knobs + entity_policy fields + base_version). Folded into cache_key as an optional 9th dimension via a new keyword arg with backward-compat default None — legacy 8-dim callers continue to work and produce the pre-2026-05-01 hash result. aborist/qa/{runner,query}.py Compute verifier_policy_hash from the policy_variant and pass to cache_key as the 9th dim. Records written from this commit forward bind to the verifier-policy identity. Pre-existing 8-dim records become cache-misses on next lookup — same migration cost as any governance_policy_hash bump. The 9-dim form makes the question "did the verifier rules change?" answerable from cache_key diff alone, separate from "did the prompt change?" CLAUDE.md Adds "old maps vs runtime maps" engineering discipline to the bench-maxing section. Codifies the principle: every base model carries old maps from training time; the runtime carries the fresh map; when they disagree, the runtime wins. Move authority OUT of the model's prior and INTO runtime artifacts (pointer IDs the runtime mints, source spans the runtime interpolates by offset, evidence maps assembled per query, policy hashes that fold prompt/verifier/retrieval into cache_key, hard checks run by the verifier). Hermes' content-addressed-evidence-id hallucination (commitbb8450d) is the canonical case study. TESTS ----- tests/test_keys.py Six new tests for verifier_policy_hash: - only hashes verifier subset (non-verifier fields don't change it) - changes when verifier-relevant field changes - empty-policy → stable - 9-dim cache_key distinct from 8-dim - 9-dim distinct under different verifier hashes - 8-dim form preserved for legacy callers (None == omit) DEFERRED (per fox + naming-deferral.md) --------------------------------------- - PROMETHEUS-Σ as an extracted controller module — the dispatch logic in runner.py + query.py already IS PROMETHEUS-Σ; an explicit prometheus.py is a refactor with no behavior change. Defer until a §5-rename-trigger fires. 507 tests pass (was 501 before, +6 from verifier_policy_hash coverage).
1989 lines
85 KiB
Python
1989 lines
85 KiB
Python
"""Multi-source corpus Q&A.
|
||
|
||
Pose a question, the tree finds related cached docs, assembles them as context,
|
||
asks Hermes, caches the answer.
|
||
|
||
The flow:
|
||
1. FTS5 search across all shards (chunks_fts can't be UNION'd in views,
|
||
so we query each shard's index independently and merge by score).
|
||
2. Title-boost rerank: hits whose title contains query tokens get a
|
||
score bump. Title is a strong topical signal that BM25 alone misses
|
||
(BM25 favors short docs with rare body tokens — Tell_(poker) outranks
|
||
Back_to_the_Future without a title boost).
|
||
3. Pick top-K distinct documents within a character budget.
|
||
4. Compute `context_root` = Merkle root over the sorted source
|
||
document_roots — that's the "source" dimension of v9.8's 8-dim
|
||
cache_key for this multi-source answer.
|
||
5. Cache lookup; hit returns the persisted audit_mode.
|
||
6. Miss calls Hermes via the OpenAI-compatible client, then runs the
|
||
faithfulness check (`verify_quotes`) — every double-quoted span in
|
||
the answer is verbatim-matched against the assembled context.
|
||
Result classifies the answer:
|
||
STRICT every quote (>=1) verified against context
|
||
HYBRID some claims sourced, some emergent (training-derived)
|
||
UNGROUNDED no quotes verify — purely emergent
|
||
7. Persist record with merkle_proof = {context_root, sources: [...]},
|
||
audit_mode, and unverified_quotes (the spans the model produced
|
||
that didn't appear in any source — corpus-growth signal).
|
||
|
||
Per-source proofs are not bundled here (the source roots themselves are
|
||
already content-addressed). A verifier asks the shards for any specific
|
||
chunk's proof on demand.
|
||
"""
|
||
|
||
from __future__ import annotations
|
||
|
||
import json
|
||
import re
|
||
import time
|
||
from dataclasses import dataclass
|
||
from pathlib import Path
|
||
|
||
from aborist import (
|
||
CANONICALIZATION_VERSION,
|
||
CHUNKING_VERSION,
|
||
SCHEMA_VERSION,
|
||
)
|
||
from aborist.compress import unpack_chunk
|
||
from aborist.merkle import MerkleTree
|
||
from aborist.qa.client import ChatClient
|
||
from aborist.qa.concepts import (
|
||
has_compare_phrasing,
|
||
rivalry_excluded,
|
||
synonym_expand,
|
||
)
|
||
from aborist.qa.keys import (
|
||
DEFAULT_FIDELITY,
|
||
DEFAULT_QUESTION_DEDUP,
|
||
FIDELITY_MODES,
|
||
QUESTION_DEDUP_MODES,
|
||
cache_key,
|
||
canonical_question,
|
||
conversation_hash,
|
||
governance_policy_hash,
|
||
model_profile_hash,
|
||
question_hash,
|
||
verifier_policy_hash,
|
||
)
|
||
from aborist.qa.dag import build_run_dag
|
||
from aborist.qa.evidence import (
|
||
build_evidence_map,
|
||
evidence_map_root,
|
||
render_evidence_map,
|
||
render_evidence_map_for_json,
|
||
)
|
||
from aborist.qa.repair import mechanical_repair, reprompt_repair
|
||
from aborist.qa.verify import (
|
||
ANSWER_MODES,
|
||
CLAIM_LATTICE_JSON_SCHEMA,
|
||
verify_claim_lattice_json,
|
||
DEFAULT_ANSWER_MODE,
|
||
verify_claim_lattice,
|
||
verify_quotes,
|
||
)
|
||
|
||
try:
|
||
from aborist.wikitext import BASE_VERSION as _WIKITEXT_BASE_VERSION
|
||
from aborist.wikitext import to_base as _wikitext_to_base
|
||
except ImportError: # pragma: no cover
|
||
_WIKITEXT_BASE_VERSION = None
|
||
_wikitext_to_base = None
|
||
from aborist.search import FTS5Backend
|
||
from aborist.store import (
|
||
append_audit,
|
||
connect,
|
||
discover_shards,
|
||
transaction,
|
||
)
|
||
|
||
|
||
_TITLE_TOKEN_RE = re.compile(r"[A-Za-z][A-Za-z0-9]*")
|
||
# Kept in sync with FTS5 stopwords in aborist.search.fts5 — both filter
|
||
# the same set of question-shaping words. "tell" leaking into title-LIKE
|
||
# search caused "tell me about permacomputer" to pull Tell_(poker), the
|
||
# Tell-Tale_Heart movie, Tell_City Indiana, etc.
|
||
_TITLE_STOPWORDS = frozenset(
|
||
"""
|
||
the a an is are was were be been being of to in on at for with by from
|
||
as about into through during and or but not no nor so yet too very also just
|
||
what who where when why how which this that these those such i you he she
|
||
it we they me him her us them do does did have has had can could should
|
||
would will may might
|
||
tell show describe explain summarize say give list find make please
|
||
all there know everything anything something
|
||
""".split()
|
||
)
|
||
|
||
|
||
def _title_query_tokens(s: str) -> set[str]:
|
||
return {
|
||
t.lower()
|
||
for t in _TITLE_TOKEN_RE.findall(s)
|
||
if t.lower() not in _TITLE_STOPWORDS and len(t) > 1
|
||
}
|
||
|
||
|
||
def _rerank_by_title(
|
||
hits: list,
|
||
question: str,
|
||
boost: float = 10.0,
|
||
) -> list:
|
||
"""Boost hits whose title overlaps query tokens. Pure ordering aid."""
|
||
qtokens = _title_query_tokens(question)
|
||
if not qtokens:
|
||
return hits
|
||
for h in hits:
|
||
if not h.title:
|
||
continue
|
||
ttokens = _title_query_tokens(h.title.replace("_", " "))
|
||
overlap = len(qtokens & ttokens)
|
||
if overlap:
|
||
h.score += overlap * boost
|
||
hits.sort(key=lambda h: -h.score)
|
||
return hits
|
||
|
||
|
||
def _filter_by_title_relevance(
|
||
hits: list,
|
||
question: str,
|
||
*,
|
||
core_match_roots: set[str] | None = None,
|
||
body_density_check: callable | None = None,
|
||
fallback_top_n: int = 5,
|
||
) -> list:
|
||
"""Concept-aware relevance filter with three accept paths:
|
||
|
||
1. Title-token overlap (after synonym expansion). Strongest signal.
|
||
2. TF-IDF core keyword overlap — `core_match_roots` is a precomputed
|
||
set of source document_roots whose TF-IDF cores contain query
|
||
tokens. Closes the gap for neologisms like "permacomputer" that
|
||
never appear in titles but are distinctive enough to be TF-IDF
|
||
keywords of conversation bodies.
|
||
3. Body density — docs mentioning the query token >= N times pass
|
||
even without title or core match. Cheap proxy for "actually about
|
||
the topic." `body_density_check(hit)` returns bool.
|
||
|
||
Rivalry exclusion (Intel-titled docs in AMD queries) still applies on
|
||
every accept path.
|
||
|
||
If all three accept paths together produce nothing, fall back to the
|
||
top `fallback_top_n` body-BM25 hits — the LLM gets enough context to
|
||
say "I don't know" rather than fabricating from a single tangential
|
||
source.
|
||
"""
|
||
qtokens = _title_query_tokens(question)
|
||
if not qtokens:
|
||
return hits
|
||
qtokens_stem = {_stem_token_for_match(t) for t in qtokens}
|
||
accept = synonym_expand(qtokens)
|
||
exclude = rivalry_excluded(qtokens, compare_phrasing=has_compare_phrasing(question))
|
||
core_roots = core_match_roots or set()
|
||
# Title-overlap breadth threshold scales with query length, mirroring
|
||
# _body_density_passes: ≤2 tokens require ALL, 3+ require N-1. Without
|
||
# this, a 2-token query like "supermans girlfriend" admits docs that
|
||
# share only ONE token with the title (e.g. `Girlfriends` the TV show)
|
||
# — title-overlap fires first & body-density never gets to reject.
|
||
title_breadth = len(qtokens) if len(qtokens) <= 2 else len(qtokens) - 1
|
||
|
||
kept: list = []
|
||
for h in hits:
|
||
ttokens = _title_query_tokens(h.title.replace("_", " ")) if h.title else set()
|
||
ttokens_stem = {_stem_token_for_match(t) for t in ttokens}
|
||
if exclude & ttokens:
|
||
continue # rivalry: opposing-side title, drop it
|
||
# Direct stem-aware match against query tokens (each qtoken must
|
||
# be present, possessive/plural-tolerant). Strict signal.
|
||
direct_matches = len(qtokens_stem & ttokens_stem)
|
||
if direct_matches >= title_breadth:
|
||
kept.append(h)
|
||
continue
|
||
# Synonym fallback only for 1-token queries — otherwise a single
|
||
# synonym hit (e.g. "amd" matching an "intel"-titled doc via the
|
||
# Intel/AMD group) would over-recall.
|
||
if len(qtokens) == 1 and accept & ttokens:
|
||
kept.append(h)
|
||
continue
|
||
if h.document_root in core_roots:
|
||
kept.append(h)
|
||
continue
|
||
if body_density_check is not None and body_density_check(h):
|
||
kept.append(h)
|
||
continue
|
||
if not kept:
|
||
return hits[: max(1, fallback_top_n)] if hits else []
|
||
return kept
|
||
|
||
|
||
DEFAULT_QUERY_POLICY = {
|
||
"system_prompt": (
|
||
"You are answering a question using ONLY the sources provided below. "
|
||
"Each source is delimited by '=== Source: <URI> ===' headers.\n\n"
|
||
"GROUNDING RULE (most important):\n"
|
||
"For EVERY factual claim in your answer, include a verbatim quote "
|
||
"from a source enclosed in double quotes (\"...\"). The quoted span "
|
||
"must appear word-for-word in one of the sources. Make a claim "
|
||
"only when a verbatim quote in one of the sources supports it. "
|
||
"Quote the source word-for-word inside the double quotes.\n\n"
|
||
"ATTRIBUTION RULES:\n"
|
||
"1. A single source may discuss MULTIPLE products, companies, or "
|
||
"competitors. Read carefully and attribute facts only to the entity "
|
||
"the source explicitly names for that fact.\n"
|
||
"2. When the question asks about ONE specific entity (e.g., 'fastest "
|
||
"AMD CPU'), keep facts about that entity in the answer; leave facts "
|
||
"about competitors (Intel, Pentium) out, even when the source "
|
||
"discusses both — they are different products that happen to share "
|
||
"an article.\n"
|
||
"3. When referencing model numbers like 'Athlon XP 3200+', remember "
|
||
"that AMD's PR-rating numbers (3200+, 2500+) are model labels, not "
|
||
"the clock speed in MHz. Quote the source's own wording for clock "
|
||
"speeds; let the source state the speed.\n"
|
||
"4. When the sources contain the answer, write it. When the answer "
|
||
"is absent from the sources, say 'I don't know based on the "
|
||
"provided sources.' and stop there. Stay inside the sources at "
|
||
"all times — let the corpus speak."
|
||
),
|
||
# Restated rule fired as a user message immediately before the sources
|
||
# turn. Hermes (and most instruction-tuned 8B models) follow recent
|
||
# user-turn instructions more reliably than a system-turn rule that
|
||
# decays under long context. Repetition is not redundancy — it raises
|
||
# the prior on the response shape we want.
|
||
"grounding_reminder": (
|
||
"REMINDER: wrap every factual claim in double quotes (\"...\") "
|
||
"and the quoted span must appear word-for-word in one of the "
|
||
"sources. Each claim earns a verbatim quote. "
|
||
"Example:\n\n"
|
||
" Q: who founded Apple?\n"
|
||
" Sources: ...Apple Inc. was founded by Steve Jobs, Steve "
|
||
"Wozniak, and Ronald Wayne in 1976...\n"
|
||
" A: Apple's founders are named in the source: \"Apple Inc. was "
|
||
"founded by Steve Jobs, Steve Wozniak, and Ronald Wayne\".\n\n"
|
||
"Now answer the question on the next message."
|
||
),
|
||
"temperature": 0.1,
|
||
"top_p": 1.0,
|
||
"max_tokens": 768,
|
||
# Entity-path policy for the faithfulness verifier. See
|
||
# aborist/qa/verify.py:ENTITY_POLICIES. Default 'proximity' promotes
|
||
# to STRICT only if N verified entities cluster within W chars of
|
||
# each other in source — separating "source documents these entities
|
||
# as a group" (cast list / infobox / roster) from "source mentions
|
||
# them incidentally in scattered prose". Lives in policy so changing
|
||
# it bumps governance_policy_hash and invalidates cache cleanly.
|
||
"entity_policy": "proximity",
|
||
"entity_proximity_n": 3,
|
||
"entity_proximity_window": 300,
|
||
# Strip wikitext markup before the LLM ever sees the context. Lets
|
||
# Hermes quote prose verbatim and shrinks token bills (~43% on
|
||
# Wikipedia chunks). Bumps governance_policy_hash so prior cached
|
||
# answers under raw-wikitext policy stay distinct on lookup. No-op
|
||
# if mwparserfromhell isn't installed.
|
||
"base_version": _WIKITEXT_BASE_VERSION,
|
||
# Mechanical answer repair after first verify. Off by default so
|
||
# existing callers don't see answer text mutate. When on:
|
||
# synthetic_elision splits, trailing_artifact trims, and no_overlap
|
||
# claim drops are applied deterministically; the repaired answer is
|
||
# re-verified & persisted (cache_key inputs unchanged, only
|
||
# answer_text differs from what the LLM produced). One audit event
|
||
# `providence_repair` records the pre→post transition. Bumps
|
||
# governance_policy_hash so on/off agents share no cache silos.
|
||
"repair_enabled": False,
|
||
# Maximum re-prompt iterations after mechanical repair. 0 = no
|
||
# re-prompt (mechanical only). 1 = at most one extra LLM call
|
||
# asking the model to rewrite around the failed quotes.
|
||
"repair_max_reprompts": 0,
|
||
# G0 / CTI — claim-lattice-pointer answer mode. See
|
||
# aborist/qa/runner.py:DEFAULT_POLICY for full semantics. Default
|
||
# "quote" preserves existing behavior; "claim_lattice_pointer"
|
||
# instructs the runtime to build an evidence map, show short
|
||
# pointer ids (E1, E2, …) to the model, and accept pointer-line
|
||
# output ("Claim. [E12]") that the verifier maps back to
|
||
# content-addressed evidence_ids for the cache & run-DAG.
|
||
"answer_mode": DEFAULT_ANSWER_MODE,
|
||
"claim_lattice_system_prompt": (
|
||
"You will see numbered EVIDENCE blocks tagged E1, E2, E3, etc. "
|
||
"Answer using natural-language pointer-lines: one claim per "
|
||
"line, followed by a bracket tag with the pointer IDs that "
|
||
"directly support that claim.\n\n"
|
||
"WORKED EXAMPLE 1 — narrow factoid\n"
|
||
"---------------------------------\n"
|
||
"EVIDENCE:\n\n"
|
||
"=== E1 (Apple_Inc | primary_answer_source) ===\n"
|
||
"Apple Inc. was founded by Steve Jobs, Steve Wozniak, and "
|
||
"Ronald Wayne in April 1976.\n\n"
|
||
"=== E2 (Steve_Wozniak | secondary_context_source) ===\n"
|
||
"Steve Wozniak co-founded Apple Computer Company alongside "
|
||
"Steve Jobs in 1976 and designed the Apple I.\n\n"
|
||
"QUESTION: who founded Apple?\n\n"
|
||
"ANSWER:\n"
|
||
"Steve Jobs co-founded Apple. [E1]\n"
|
||
"Steve Wozniak co-founded Apple. [E1,E2]\n"
|
||
"Ronald Wayne co-founded Apple. [E1]\n\n"
|
||
"WORKED EXAMPLE 2 — broad descriptive\n"
|
||
"------------------------------------\n"
|
||
"EVIDENCE:\n\n"
|
||
"=== E1 (Mars | primary_answer_source) ===\n"
|
||
"Mars is the fourth planet from the Sun. It has two moons, "
|
||
"Phobos and Deimos. Mars has a thin atmosphere of carbon "
|
||
"dioxide. Average surface temperature is around -60 "
|
||
"degrees Celsius.\n\n"
|
||
"QUESTION: tell me about Mars\n\n"
|
||
"ANSWER:\n"
|
||
"Mars is the fourth planet from the Sun. [E1]\n"
|
||
"Mars has two moons, Phobos and Deimos. [E1]\n"
|
||
"Mars has a thin atmosphere of carbon dioxide. [E1]\n"
|
||
"The average surface temperature on Mars is around -60 "
|
||
"degrees Celsius. [E1]\n\n"
|
||
"END OF EXAMPLES\n\n"
|
||
"RULES (each rule says what TO do):\n"
|
||
"1. Reference evidence by pointer ID. The runtime displays "
|
||
"the literal source span beside each claim — referencing is "
|
||
"your job; quoting is the runtime's job.\n"
|
||
"2. Use only pointer IDs that appear in the EVIDENCE blocks "
|
||
"above.\n"
|
||
"3. Cite 1 or 2 pointers per claim — the blocks whose text "
|
||
"directly contains the claim's key terms.\n"
|
||
"4. End every claim line with [E#] or [E#,E#].\n"
|
||
"5. Make a claim only when an EVIDENCE block textually "
|
||
"supports it. Stop when the evidence runs out — a short "
|
||
"answer is the right answer when only short evidence "
|
||
"exists.\n"
|
||
"6. Write each claim as one plain-prose sentence on its own "
|
||
"line.\n"
|
||
"7. Each line begins with the claim text and ends with a "
|
||
"bracket tag of pointer ids. The headers above the QUESTION "
|
||
"(`EVIDENCE:`, `=== E1 (Title) ===`) are scaffolding for "
|
||
"the input only — your answer is the claim lines themselves.\n"
|
||
"8. Each claim line cites at most two pointers. Lines with "
|
||
"three or more pointers are rejected. When more than two "
|
||
"evidence blocks support a claim, pick the two that most "
|
||
"directly contain the claim's key terms, or split the "
|
||
"claim into two lines."
|
||
),
|
||
"claim_lattice_grounding_reminder": (
|
||
"REMINDER: format = pointer-line — `Claim text. [E1]` per "
|
||
"line, plain prose with bracket tags. Pointer IDs come from "
|
||
"the EVIDENCE blocks above. At most two pointers per claim. "
|
||
"Begin each line with the claim text. Now answer the question "
|
||
"on the next message."
|
||
),
|
||
"claim_lattice_allowed_source_roles": [
|
||
"primary_answer_source",
|
||
"secondary_context_source",
|
||
"background_source",
|
||
"unclassified",
|
||
# Self-reference: STRICT-trusted-as-fact unless falsified.
|
||
# See docs/self-reference-thought-chains-design.md.
|
||
"self_reference_source",
|
||
],
|
||
"claim_lattice_max_pointers_per_claim": 2,
|
||
# Cap on evidence blocks (chunks) per retrieved source. Default 2.
|
||
# Without this cap, a long Wikipedia article alone can split into
|
||
# ~20 chunks, each becoming a separate E* — the model then sees
|
||
# E1-E26 for what is really 5 sources and writes a mega-claim
|
||
# citing all of them. With the cap, 5 sources × 2 = 10 evidence
|
||
# blocks max. Chunks within each source are still relevance-ranked
|
||
# (distinct_query_tokens DESC, total_mentions DESC, chunk_idx ASC)
|
||
# so the cap keeps the most-relevant 2 chunks. Folds into
|
||
# governance_policy_hash; changing the cap invalidates prior
|
||
# cached records.
|
||
"claim_lattice_max_chunks_per_source": 2,
|
||
"claim_lattice_min_citation_coverage": 0.30,
|
||
"claim_lattice_min_claim_content_tokens": 2,
|
||
"claim_lattice_lazy_anchor_demote_threshold": 0.5,
|
||
"claim_lattice_lazy_anchor_demote_min_pairs": 3,
|
||
# Claim-count ceiling — see runner.DEFAULT_POLICY for rationale.
|
||
# Bench finding (york-england "tell me all there is to know")
|
||
# caught the runaway shape; cap of 12 admits entity-list
|
||
# questions while flagging encyclopedic spam. Folds into
|
||
# governance_policy_hash on change.
|
||
"claim_lattice_max_claims_per_answer": 12,
|
||
# JSON variant — `answer_mode="claim_lattice"`. Pairs with grammar-
|
||
# constrained inference (vLLM guided_json, Claude/GPT-4 native
|
||
# JSON, Qwen 3.6 reasoner). Lenient pre-parser keeps the path
|
||
# survivable on inference paths without grammar guidance. Toggling
|
||
# `claim_lattice_use_guided_json=False` disables the extra_body
|
||
# pass for endpoints that 400 on unknown fields.
|
||
"claim_lattice_json_system_prompt": (
|
||
"You will see numbered EVIDENCE blocks tagged with short "
|
||
"pointer IDs: E1, E2, E3, etc. Answer as a single JSON "
|
||
"object using EXACTLY this schema. Your output begins with "
|
||
"`{` and ends with `}`:\n\n"
|
||
' {"claims":[{"text":"<claim text>","evidence_ids":["E1"]}]}\n\n'
|
||
"RULES:\n"
|
||
"1. Your output is exactly one JSON object — first character "
|
||
"is `{`, last character is `}`.\n"
|
||
"2. `text` is plain prose. Punctuation (including any quoted "
|
||
"names from the source) appears in the source span the "
|
||
"runtime renders for you, so write the text using only "
|
||
"letters, digits, spaces, and basic sentence punctuation "
|
||
"(`,`, `.`, `;`, `:`, `?`, `!`, `(`, `)`, `-`).\n"
|
||
"3. `evidence_ids` are pointer IDs that appear verbatim in "
|
||
"the EVIDENCE block headers above (E1, E2, …). Use only IDs "
|
||
"you can see in the headers above. At most two IDs per claim.\n"
|
||
"4. Each claim references at least one evidence_id.\n"
|
||
"5. Emit a claim only when an EVIDENCE block directly "
|
||
"supports it; stop emitting claims when supporting evidence "
|
||
"runs out — a short answer is the right answer when only "
|
||
"short evidence exists.\n"
|
||
"6. Each claim is one focused fact citing the source that "
|
||
"directly contains it. Multiple distinct facts become "
|
||
"multiple claim entries, each citing the source for that "
|
||
"fact. Three short claims, each citing one pointer, beat "
|
||
"one long claim citing five."
|
||
),
|
||
"claim_lattice_json_grounding_reminder": (
|
||
"REMINDER: emit one JSON object — first character `{`, last "
|
||
"character `}` — with the schema "
|
||
'`{"claims":[{"text":"...","evidence_ids":["E1"]}]}`. '
|
||
"evidence_ids are pointer IDs (E1, E2, …) that appear "
|
||
"verbatim in the EVIDENCE block headers above. At most two "
|
||
"per claim. Each claim is one focused fact. Now answer the "
|
||
"question on the next message."
|
||
),
|
||
"claim_lattice_use_guided_json": True,
|
||
"claim_lattice_json_stop_sequences": ["\n\n"],
|
||
}
|
||
|
||
|
||
@dataclass
|
||
class _Hit:
|
||
document_root: str
|
||
document_uri: str
|
||
title: str | None
|
||
score: float
|
||
shard_path: str
|
||
chunk_idx: int
|
||
source_role: str = "unclassified"
|
||
|
||
|
||
# Heuristic role classification + per-role budget multiplier. Lets the
|
||
# primary answer page (e.g. `Jurassic Park (film)` for a JP film query)
|
||
# claim a wider context slice than peripheral pages (`Jurassic Park (film
|
||
# score)`, `Jurassic Park video games`). The running `char_budget` check
|
||
# still bounds total context to `max_context_chars`; weights just shift
|
||
# how the budget gets divided.
|
||
SOURCE_ROLE_BUDGET_WEIGHTS = {
|
||
"primary_answer_source": 2.0,
|
||
"secondary_context_source": 1.0,
|
||
"noisy_background_source": 0.5,
|
||
"sequel_background_source": 0.5,
|
||
"background_source": 1.0,
|
||
"unclassified": 1.0,
|
||
# Self-promoted providence records (STRICT live, past kindergarten
|
||
# window). Same budget weight as background — Wikipedia stays the
|
||
# canonical primary; self-reference is supplementary anchoring.
|
||
# Trust model: STRICT-as-fact unless the verifier falsifies it.
|
||
"self_reference_source": 1.0,
|
||
}
|
||
|
||
# Title patterns that demote a source's role. Lower-cased substring match.
|
||
# 2026-04-30: extended to catch tie-in spinoff titles. The JP-dinosaurs
|
||
# query lazy-anchored on "Jurassic Park: Operation Genesis" (a video game)
|
||
# whose enumerative dinosaur tables pattern-matched the question shape
|
||
# more cleanly than the actual film article's prose. Adding the explicit
|
||
# game subtitle plus generic markers ("the game", "video games") so
|
||
# similar tie-ins classify as noisy and drop out of the evidence map.
|
||
_NOISY_TITLE_MARKERS = (
|
||
"score", "music", "soundtrack", "video game", "video games",
|
||
"merchandise", "discography", "operation genesis", "the game",
|
||
)
|
||
_SEQUEL_TITLE_MARKERS = (
|
||
" iii", " ii)", " ii ", " iv", " v ", " v)", "lost world", "sequel",
|
||
" 2)", " 3)", " 4)",
|
||
)
|
||
_SECONDARY_TITLE_MARKERS = (
|
||
"list of", "characters", "franchise", "history of", "people",
|
||
"timeline of",
|
||
)
|
||
|
||
|
||
def _classify_source_role(
|
||
title: str | None,
|
||
qtokens_stem: set[str],
|
||
*,
|
||
document_uri: str | None = None,
|
||
) -> str:
|
||
"""Tag a source by its likely role for an N-token query.
|
||
|
||
URI-scheme classification fires first: documents whose URI starts
|
||
with ``aborist://providence/`` are self-promoted providence
|
||
records (per ``aborist/sources/providence.py``) and classify as
|
||
``self_reference_source`` regardless of title shape — that role
|
||
captures the trust model "STRICT-as-fact unless verifier
|
||
falsifies."
|
||
|
||
Order matters for the title-based fallback: noisy/sequel/secondary
|
||
markers fire first because they catch peripheral pages whose
|
||
titles otherwise overlap query tokens fully (e.g. `Jurassic Park
|
||
(film score)` shares 3 stems with `{dinosaur, jurassic, park,
|
||
film}` but is not the primary answer source for a dinosaurs
|
||
question). Primary requires the strongest title coverage (N-1
|
||
of N stems present).
|
||
"""
|
||
if document_uri and document_uri.startswith("aborist://providence/"):
|
||
return "self_reference_source"
|
||
if not title:
|
||
return "unclassified"
|
||
t = title.lower()
|
||
if any(k in t for k in _NOISY_TITLE_MARKERS):
|
||
return "noisy_background_source"
|
||
if any(k in t for k in _SEQUEL_TITLE_MARKERS):
|
||
return "sequel_background_source"
|
||
if any(k in t for k in _SECONDARY_TITLE_MARKERS):
|
||
return "secondary_context_source"
|
||
title_tokens = _title_query_tokens(t.replace("_", " "))
|
||
title_stems = {_stem_token_for_match(tok) for tok in title_tokens}
|
||
if qtokens_stem and len(title_stems & qtokens_stem) >= max(1, len(qtokens_stem) - 1):
|
||
return "primary_answer_source"
|
||
return "background_source"
|
||
|
||
|
||
def _search_titles(conn, qtokens: list[str], limit: int) -> list[tuple]:
|
||
"""SQL LIKE over documents.title — finds the HTTP article that FTS5
|
||
misses because list-pages with many URLs have higher 'http' term
|
||
frequency than the actual protocol article. Returns rows shaped
|
||
to match the FTS5 hit tuple."""
|
||
if not qtokens:
|
||
return []
|
||
clauses = " OR ".join(["LOWER(title) LIKE ?"] * len(qtokens))
|
||
params = [f"%{t.lower()}%" for t in qtokens]
|
||
params.append(limit)
|
||
rows = conn.execute(
|
||
f"SELECT document_root, document_uri, title FROM documents "
|
||
f"WHERE {clauses} LIMIT ?",
|
||
params,
|
||
).fetchall()
|
||
return rows
|
||
|
||
|
||
def _docs_with_core_keyword_match(
|
||
conn, qtokens: list[str], limit: int
|
||
) -> list[tuple]:
|
||
"""Find SOURCE docs whose TF-IDF core keywords contain a query token.
|
||
|
||
TF-IDF cores act as enriched titles: a doc's distinctive low-frequency
|
||
terms get distilled into the core's content as comma-separated keywords.
|
||
A query for "permacomputer" (a neologism that never makes it into a
|
||
title) can match the TF-IDF core of a Grok conversation that mentioned
|
||
it, because permacomputer is a rare term that ranks high under TF-IDF.
|
||
|
||
Returns rows of (document_root, document_uri, title) for the SOURCE docs
|
||
(not the cores) — the source is what gets fed to the LLM as context.
|
||
"""
|
||
if not qtokens:
|
||
return []
|
||
# Word-boundary match against the comma-separated TF-IDF keyword list.
|
||
# Prepending/appending ", " lets one LIKE pattern (`%, token, %`) check
|
||
# for the token regardless of its position in the keyword string.
|
||
# Without this, naive `LIKE '%intel%'` would match "intelligence",
|
||
# "intellectual", "intellivision" — drowning real hits like Pentium_4
|
||
# (whose TF-IDF core has "intel" as an exact keyword) in noise.
|
||
#
|
||
# Per-row `match_count` tallies how many distinct query tokens hit
|
||
# this doc's TF-IDF core. Multi-token coverage is a strong relevance
|
||
# signal — a doc whose core has "intel" + "cpu" + "faster" beats a
|
||
# doc whose only signal is "intel" appearing in its TITLE. The
|
||
# caller boosts the score by match_count.
|
||
case_clauses = " + ".join(
|
||
[
|
||
"(CASE WHEN LOWER(', ' || c.content || ', ') LIKE ? THEN 1 ELSE 0 END)"
|
||
]
|
||
* len(qtokens)
|
||
)
|
||
where_clauses = " OR ".join(
|
||
["LOWER(', ' || c.content || ', ') LIKE ?"] * len(qtokens)
|
||
)
|
||
patterns = [f"%, {t.lower()}, %" for t in qtokens]
|
||
params = patterns + patterns + [limit]
|
||
rows = conn.execute(
|
||
f"""
|
||
SELECT
|
||
src.document_root,
|
||
src.document_uri,
|
||
src.title,
|
||
MAX({case_clauses}) AS match_count
|
||
FROM chunks c
|
||
JOIN documents core ON core.document_root = c.document_root
|
||
JOIN derivations der ON der.core_root = core.document_root
|
||
JOIN documents src ON src.document_root = der.src_root
|
||
WHERE core.source_type LIKE 'core:tfidf-%'
|
||
AND ({where_clauses})
|
||
GROUP BY src.document_root, src.document_uri, src.title
|
||
ORDER BY match_count DESC
|
||
LIMIT ?
|
||
""",
|
||
params,
|
||
).fetchall()
|
||
return rows
|
||
|
||
|
||
def _stem_token_for_match(t: str) -> str:
|
||
"""Light suffix-strip for query-token vs body matching.
|
||
|
||
Two normalizations:
|
||
possessive ``"superman's" -> "supermans" -> "superman"`` (the apostrophe
|
||
is already gone via _TITLE_TOKEN_RE; we drop the trailing
|
||
``s`` here so the lookup matches plain ``superman`` in body).
|
||
plural ``"powers" -> "power"``, ``"girlfriends" -> "girlfriend"``
|
||
so plural questions match singular source mentions.
|
||
|
||
Both are the same operation: strip trailing ``s`` for tokens > 4 chars.
|
||
Conservative on short tokens (``"is"``, ``"as"``, ``"us"`` would lose
|
||
meaning) and on tokens that don't end in ``s`` (no-op).
|
||
"""
|
||
if len(t) > 4 and t.endswith("s") and not t.endswith("ss"):
|
||
return t[:-1]
|
||
return t
|
||
|
||
|
||
def _body_count_with_stem(body: str, t: str) -> int:
|
||
"""Count mentions of ``t`` in ``body``, falling back to the lite-stemmed
|
||
form if the literal didn't match. Returns the LARGER of the two counts
|
||
so a query token that appears under both forms (rare) still scores."""
|
||
n_literal = body.count(t)
|
||
if n_literal:
|
||
return n_literal
|
||
stem = _stem_token_for_match(t)
|
||
if stem != t:
|
||
return body.count(stem)
|
||
return 0
|
||
|
||
|
||
def _chunk_query_relevance(
|
||
span: str, qtokens_stem: set[str]
|
||
) -> tuple[int, int]:
|
||
"""Rank one chunk by query-token overlap. Returns
|
||
``(distinct_present, total_mentions)``.
|
||
|
||
Stem-aware: same ``_body_count_with_stem`` we use elsewhere, so
|
||
``"supermans"`` matches ``"superman"`` in the chunk text. Soft
|
||
signal — the score never enters the proof path; it only orders
|
||
chunks within a source so the most-relevant chunk gets the lowest
|
||
pointer id (and the model's lazy-anchor habit lands on a useful
|
||
chunk by accident).
|
||
|
||
Sort callers should walk ``(-distinct, -total, chunk_idx_asc)``
|
||
to break ties stably toward document order.
|
||
"""
|
||
if not qtokens_stem:
|
||
return (0, 0)
|
||
body = span.lower()
|
||
counts = {t: _body_count_with_stem(body, t) for t in qtokens_stem}
|
||
distinct_present = sum(1 for n in counts.values() if n > 0)
|
||
total = sum(counts.values())
|
||
return (distinct_present, total)
|
||
|
||
|
||
def _body_density_passes(
|
||
conn, document_root: str, qtokens: set[str], min_mentions: int = 3
|
||
) -> bool:
|
||
"""Body-token CO-OCCURRENCE relevance.
|
||
|
||
Pre-2026-04-27 this counted any token's mentions in the body (so
|
||
Intel_8086, with 30 "Intel" mentions, would pass for a query of
|
||
{intel, fastest, cpu} despite never mentioning "fastest"). Pre-2026-
|
||
04-29 the threshold was "at least HALF" — too lenient for 2-token
|
||
queries: ``"supermans girlfriend"`` admitted ``Girlfriends`` (TV
|
||
show) which had ``girlfriend`` in body but no ``superman`` whatsoever.
|
||
|
||
Current rules — breadth scales with query length:
|
||
|
||
≤ 2 tokens require ALL of them present in body
|
||
3+ tokens require N - 1 (allow one weak signal token to miss)
|
||
|
||
Plus depth (``total_mentions >= min_mentions``) is still enforced.
|
||
|
||
Token matching is stem-tolerant (``_stem_token_for_match``): so a query
|
||
token ``"supermans"`` matches body ``"superman"`` and ``"girlfriends"``
|
||
matches body ``"girlfriend"``.
|
||
"""
|
||
if not qtokens:
|
||
return False
|
||
rows = conn.execute(
|
||
"SELECT content FROM chunks WHERE document_root = ? AND content IS NOT NULL",
|
||
(document_root,),
|
||
).fetchall()
|
||
if not rows:
|
||
return False
|
||
body = " ".join(unpack_chunk(r["content"]) or "" for r in rows).lower()
|
||
counts = {t: _body_count_with_stem(body, t.lower()) for t in qtokens}
|
||
distinct_present = sum(1 for n in counts.values() if n > 0)
|
||
total_mentions = sum(counts.values())
|
||
if len(qtokens) <= 2:
|
||
breadth_threshold = len(qtokens)
|
||
else:
|
||
breadth_threshold = len(qtokens) - 1
|
||
return distinct_present >= breadth_threshold and total_mentions >= min_mentions
|
||
|
||
|
||
def _search_corpus(
|
||
shards_dir: Path | None,
|
||
single_db: Path | None,
|
||
question: str,
|
||
over_fetch: int,
|
||
) -> list[_Hit]:
|
||
"""Two parallel searches across shards, merged:
|
||
|
||
- FTS5 body search (BM25 over chunk content).
|
||
- SQL title-LIKE search (catches articles whose body is short on the
|
||
query terms but whose title is the literal topic — e.g., the HTTP
|
||
protocol article doesn't out-frequency-score URL-heavy list pages
|
||
but it IS the topic).
|
||
|
||
Dedupe by document_root. Title hits get a baseline score that
|
||
out-ranks FTS5 body hits so the actual topic article rises to the top.
|
||
"""
|
||
qtokens = _title_query_tokens(question)
|
||
# Synonym expansion: a query for "athlon" also fetches AMD-titled docs.
|
||
accept_tokens = synonym_expand(qtokens)
|
||
paths: list[Path]
|
||
if shards_dir is not None:
|
||
paths = discover_shards(shards_dir)
|
||
elif single_db is not None:
|
||
paths = [Path(single_db)]
|
||
else:
|
||
paths = []
|
||
|
||
raw: list[tuple] = []
|
||
# Roots whose TF-IDF cores contain a query token — collected across
|
||
# shards. Used downstream by _filter_by_title_relevance as accept-path 2.
|
||
core_match_roots: set[str] = set()
|
||
# Per-shard mapping of doc root -> shard path, for body-density lookups
|
||
# in accept-path 3. Lets us reach back to the source shard cheaply.
|
||
root_to_shard: dict[str, str] = {}
|
||
|
||
for p in paths:
|
||
conn = connect(p)
|
||
try:
|
||
backend = FTS5Backend(conn)
|
||
for h in backend.search(question, limit=over_fetch):
|
||
raw.append(
|
||
(
|
||
h.score,
|
||
h.document_root,
|
||
h.document_uri,
|
||
h.title,
|
||
h.chunk_idx,
|
||
str(p.resolve()),
|
||
)
|
||
)
|
||
root_to_shard[h.document_root] = str(p.resolve())
|
||
# Parallel title search using synonym-expanded tokens.
|
||
# The score starts deliberately low so this signal can't drown
|
||
# FTS5 BM25 + body relevance. _rerank_by_title later adds
|
||
# `overlap*10` to every hit (FTS5 and title-search alike) that
|
||
# has title-token overlap, so a doc whose title genuinely IS
|
||
# the topic ends up rewarded twice (once here, once in rerank).
|
||
# Single-token title matches against generic terms ("intel",
|
||
# "cpu") used to score 60+ standalone, dominating top-K with
|
||
# legacy 80486-era articles for queries like "fastest intel
|
||
# CPU?". Now that contribution is the same scale as FTS5 body
|
||
# BM25, which lets Pentium_4 (high body relevance, no title
|
||
# overlap) win on its actual topical fit.
|
||
for r in _search_titles(conn, list(accept_tokens), over_fetch):
|
||
title_lower = (r["title"] or "").lower().replace("_", " ")
|
||
overlap = sum(1 for t in accept_tokens if t in title_lower)
|
||
if overlap == 0:
|
||
continue
|
||
title_score = overlap * 10.0
|
||
raw.append(
|
||
(
|
||
title_score,
|
||
r["document_root"],
|
||
r["document_uri"],
|
||
r["title"],
|
||
0,
|
||
str(p.resolve()),
|
||
)
|
||
)
|
||
root_to_shard[r["document_root"]] = str(p.resolve())
|
||
# Core-keyword search: docs whose TF-IDF core keywords match.
|
||
# The query token doesn't need to be in title or even in body —
|
||
# being a TF-IDF keyword of the doc's core is enough signal.
|
||
for r in _docs_with_core_keyword_match(
|
||
conn, list(accept_tokens), over_fetch
|
||
):
|
||
core_match_roots.add(r["document_root"])
|
||
# Score scales with how many query tokens hit this doc's
|
||
# TF-IDF core. A 3-token coverage (e.g. Pentium_4's core
|
||
# carries "intel", "cpu", "faster" for the query "fastest
|
||
# intel CPU?") beats single-token title boosts (~80) that
|
||
# otherwise saturate the top with Intel_80486DX,
|
||
# Intel_8086, etc. — articles that share *one* word with
|
||
# the query but aren't the topical answer.
|
||
match_count = r["match_count"] or 1
|
||
kw_score = 40.0 + 25.0 * match_count
|
||
raw.append(
|
||
(
|
||
kw_score,
|
||
r["document_root"],
|
||
r["document_uri"],
|
||
r["title"],
|
||
0,
|
||
str(p.resolve()),
|
||
)
|
||
)
|
||
root_to_shard[r["document_root"]] = str(p.resolve())
|
||
finally:
|
||
conn.close()
|
||
|
||
raw.sort(key=lambda r: -r[0])
|
||
seen: set[str] = set()
|
||
out: list[_Hit] = []
|
||
for score, root, uri, title, idx, sp in raw:
|
||
if root in seen:
|
||
continue
|
||
seen.add(root)
|
||
out.append(
|
||
_Hit(
|
||
document_root=root,
|
||
document_uri=uri,
|
||
title=title,
|
||
score=score,
|
||
shard_path=sp,
|
||
chunk_idx=idx,
|
||
)
|
||
)
|
||
return out
|
||
|
||
|
||
def _rerank(
|
||
hits: list[_Hit],
|
||
question: str,
|
||
*,
|
||
core_match_roots: set[str] | None = None,
|
||
body_density_check: callable | None = None,
|
||
) -> list[_Hit]:
|
||
"""Filter off-topic, then layer in body-coverage, title-overlap, and
|
||
source-role rank boosts.
|
||
|
||
Order matters: filter first (drops noise), body-coverage rerank next
|
||
(counters BM25's short-doc bias by rewarding topical density across
|
||
distinct query tokens), title-overlap rerank (breaks ties when a doc
|
||
IS the named topic), source-role rank-boost last so a primary
|
||
answer source outranks a list-page even when the list-page won on
|
||
BM25 + title-overlap (caught the "where is florida" defect:
|
||
``List_of_places_in_Florida`` and ``List_of_State_Roads_in_Florida``
|
||
each contain "Florida" hundreds of times in row markup, scoring
|
||
above the actual ``Florida`` article on body density).
|
||
"""
|
||
hits = _filter_by_title_relevance(
|
||
hits,
|
||
question,
|
||
core_match_roots=core_match_roots,
|
||
body_density_check=body_density_check,
|
||
)
|
||
hits = _rerank_by_body_coverage(hits, question)
|
||
hits = _rerank_by_title(hits, question)
|
||
hits = _rerank_by_source_role(hits, question)
|
||
return _rerank_by_title_purity(hits, question)
|
||
|
||
|
||
# Per-role rank multiplier. Affects sort order, NOT just per-source
|
||
# context cap (the latter is SOURCE_ROLE_BUDGET_WEIGHTS, applied later).
|
||
# Defaults skew strongly toward primary so a real topic article beats
|
||
# list-pages and franchise/sequel siblings even when the list-page wins
|
||
# on body-density. Tuned against the JP-dinosaurs and "where is florida"
|
||
# defects.
|
||
SOURCE_ROLE_RANK_WEIGHTS = {
|
||
"primary_answer_source": 2.0,
|
||
"secondary_context_source": 0.7,
|
||
"background_source": 0.9,
|
||
"noisy_background_source": 0.3,
|
||
"sequel_background_source": 0.3,
|
||
"unclassified": 1.0,
|
||
# Self-reference: STRICT live providence records past kindergarten
|
||
# window. Treated like background — Wikipedia stays canonical
|
||
# primary; self-reference is supplementary anchoring trusted as
|
||
# fact unless the verifier falsifies the underlying record.
|
||
"self_reference_source": 0.9,
|
||
}
|
||
|
||
|
||
def _rerank_by_source_role(hits: list[_Hit], question: str) -> list[_Hit]:
|
||
"""Classify each hit by source role and rescale score by role weight.
|
||
|
||
Mutates ``h.source_role`` so the classification happens once and
|
||
downstream context-build code can reuse the value (instead of
|
||
re-classifying at cap time). Stable sort by score desc.
|
||
"""
|
||
qtokens_stem = {
|
||
_stem_token_for_match(t.lower())
|
||
for t in _title_query_tokens(question)
|
||
}
|
||
for h in hits:
|
||
h.source_role = _classify_source_role(
|
||
h.title, qtokens_stem, document_uri=h.document_uri
|
||
)
|
||
weight = SOURCE_ROLE_RANK_WEIGHTS.get(h.source_role, 1.0)
|
||
h.score = h.score * weight
|
||
hits.sort(key=lambda h: -h.score)
|
||
return hits
|
||
|
||
|
||
def _rerank_by_title_purity(hits: list[_Hit], question: str) -> list[_Hit]:
|
||
"""Boost titles by both purity AND multi-token-match breadth.
|
||
|
||
Two signals combine here:
|
||
|
||
- **Purity** = ``|title_tokens ∩ query_tokens| / |title_tokens|``.
|
||
Rewards titles that ARE the topic without off-topic suffix tokens.
|
||
``Jurassic Park (film)`` (purity 1.0) beats ``Jurassic Park
|
||
(NES game)`` (purity 0.5).
|
||
- **Overlap count** = ``|title_tokens ∩ query_tokens|``. Rewards
|
||
titles that match more of the query's content tokens. For a
|
||
query ``{dawson, creek}``: ``List of Dawson's Creek episodes``
|
||
(overlap 2) beats ``Clinton Creek, Yukon`` (overlap 1) even
|
||
when both have similar purity.
|
||
|
||
Multiplier: ``(1 + overlap_count) * (1 + purity)``:
|
||
|
||
overlap=2, purity=0.5 (e.g. ``Dawson's Creek episodes``) → 4.5×
|
||
overlap=1, purity=1.0 (bare-token-title match) → 4.0×
|
||
overlap=2, purity=0.4 (e.g. ``List of ... Dawson Creek``) → 4.2×
|
||
overlap=1, purity=0.5 (e.g. ``Dawson Leery``) → 3.0×
|
||
overlap=1, purity=0.33 (e.g. ``Clinton Creek, Yukon``) → 2.67×
|
||
overlap=0 → 1.0× (no change)
|
||
|
||
Pre-2026-05-01 the multiplier was ``1 + 2 * purity`` — purity
|
||
only, indifferent to overlap-count. That let ``Clinton Creek,
|
||
Yukon`` (purity 0.33 → 1.67× boost) outrank ``List of Dawson's
|
||
Creek episodes`` (purity 0.5 → 2.0×) on a query like "in
|
||
dawsons creek who is the girl across the creek?" once BM25's
|
||
short-title bias is folded in. The 2-token-match should beat
|
||
the 1-token-match cleanly.
|
||
|
||
Original use case (JP-dinosaurs lazy-anchor) still served:
|
||
``Jurassic Park (film)`` (overlap 2, purity 1.0) → 6.0× sits
|
||
well above ``Jurassic Park (NES game)`` (overlap 2, purity 0.5)
|
||
→ 4.5×, and far above ``Jurassic Park (franchise)`` (overlap 2,
|
||
purity 0.67) → 5.0×.
|
||
"""
|
||
qtokens = _title_query_tokens(question)
|
||
if not qtokens:
|
||
return hits
|
||
# Stem-aware matching so possessive / plural variants match. The
|
||
# 2026-05-01 Dawson's Creek defect: question "dawsons creek" with
|
||
# title "List of Dawson's Creek episodes" — raw set intersection
|
||
# treated `dawsons` and `dawson` as distinct → overlap=1 (only
|
||
# `creek`) and the multi-token title bonus didn't fire. Stemming
|
||
# both sides via `_stem_token_for_match` (trailing-s strip on
|
||
# tokens >4 chars, skipping ss-enders) collapses both forms onto
|
||
# `dawson`, the overlap goes to 2, and the title beats single-
|
||
# token ``Clinton Creek, Yukon`` matches.
|
||
qstems = {_stem_token_for_match(t) for t in qtokens}
|
||
for h in hits:
|
||
if not h.title:
|
||
continue
|
||
ttokens = _title_query_tokens(h.title.replace("_", " "))
|
||
if not ttokens:
|
||
continue
|
||
tstems = {_stem_token_for_match(t) for t in ttokens}
|
||
overlap = tstems & qstems
|
||
if not overlap:
|
||
continue
|
||
purity = len(overlap) / len(tstems)
|
||
overlap_count = len(overlap)
|
||
h.score = h.score * (1.0 + overlap_count) * (1.0 + purity)
|
||
hits.sort(key=lambda h: -h.score)
|
||
return hits
|
||
|
||
|
||
def _rerank_by_body_coverage(
|
||
hits: list,
|
||
question: str,
|
||
*,
|
||
weight: float = 0.6,
|
||
) -> list:
|
||
"""Boost score by per-token body coverage to counteract BM25's short-doc bias.
|
||
|
||
BM25 normalizes by document length, but its `b` parameter under-penalizes
|
||
*very* short docs that happen to mention every query token. Result: a
|
||
1-paragraph stub on Intel 4040 (1974 microcontroller) outscored a 30-page
|
||
Intel Core i7 article for "what is the fastest intel CPU?".
|
||
|
||
Approach: for each hit's full body, sum sqrt(count) per query token.
|
||
Sqrt scaling lets long topical articles meaningfully out-score short
|
||
tangential ones without runaway domination by enumerative list pages.
|
||
|
||
stub article: 5 intel + 1 fastest + 4 cpu -> sqrt(5)+sqrt(1)+sqrt(4) ~ 5.2
|
||
long topical: 200 + 10 + 80 -> sqrt(200)+sqrt(10)+sqrt(80) ~ 26.2
|
||
enumeration: 1000 + 50 + 300 -> sqrt(1000)+sqrt(50)+sqrt(300) ~ 55.8
|
||
|
||
Multiply by `weight` (default 0.6) and add to the existing score. The
|
||
differentiation in the example above is decisive but bounded: long topical
|
||
articles get +15-16, enumerations get +33, stubs get +3. Combined with
|
||
FTS5 base scores in the 40s, the long topical article wins comfortably.
|
||
|
||
Cost: one body fetch per surviving candidate (typically 24-32). Bounded.
|
||
"""
|
||
import math
|
||
|
||
qtokens_lower = {t.lower() for t in _title_query_tokens(question)}
|
||
if not qtokens_lower:
|
||
return hits
|
||
for h in hits:
|
||
body = _load_doc_text(h.shard_path, h.document_root)
|
||
if body is None:
|
||
continue
|
||
body_lower = body.lower()
|
||
coverage = sum(
|
||
math.sqrt(body_lower.count(t)) for t in qtokens_lower
|
||
)
|
||
h.score += coverage * weight
|
||
hits.sort(key=lambda h: -h.score)
|
||
return hits
|
||
|
||
|
||
def _load_doc_text(shard_path: str, document_root: str) -> str | None:
|
||
"""Concatenate all hot chunks of a document. Returns None if cold or missing."""
|
||
conn = connect(shard_path)
|
||
try:
|
||
rows = conn.execute(
|
||
"SELECT content FROM chunks "
|
||
"WHERE document_root = ? AND content IS NOT NULL "
|
||
"ORDER BY idx ASC",
|
||
(document_root,),
|
||
).fetchall()
|
||
finally:
|
||
conn.close()
|
||
if not rows:
|
||
return None
|
||
return "\n\n".join(unpack_chunk(r["content"]) or "" for r in rows)
|
||
|
||
|
||
def _load_doc_chunks(
|
||
shard_path: str, document_root: str
|
||
) -> list[tuple[int, str, str]] | None:
|
||
"""Per-chunk hot rows of a document.
|
||
|
||
Returns ``[(chunk_idx, leaf_hash, span), ...]`` in chunk order, or
|
||
``None`` if the document is cold or absent. Cold individual chunks
|
||
are skipped (the WHERE clause filters NULL content). Used by the
|
||
claim-lattice-pointer evidence-map builder to emit one
|
||
``EvidenceObject`` per chunk so the model can cite the specific
|
||
paragraph that supports a claim instead of lazy-anchoring the
|
||
whole article.
|
||
"""
|
||
conn = connect(shard_path)
|
||
try:
|
||
rows = conn.execute(
|
||
"SELECT idx, leaf_hash, content FROM chunks "
|
||
"WHERE document_root = ? AND content IS NOT NULL "
|
||
"ORDER BY idx ASC",
|
||
(document_root,),
|
||
).fetchall()
|
||
finally:
|
||
conn.close()
|
||
if not rows:
|
||
return None
|
||
out: list[tuple[int, str, str]] = []
|
||
for r in rows:
|
||
span = unpack_chunk(r["content"]) or ""
|
||
if not span:
|
||
continue
|
||
out.append((r["idx"], r["leaf_hash"], span))
|
||
return out or None
|
||
|
||
|
||
def _context_root(source_roots: list[str]) -> str:
|
||
"""Merkle root over sorted source document_roots — the v9.8 'source' dim
|
||
for multi-source answers. Sorting makes the root deterministic regardless
|
||
of search ranking order."""
|
||
if not source_roots:
|
||
return "00" * 32
|
||
sorted_roots = sorted(source_roots)
|
||
if len(sorted_roots) == 1:
|
||
return sorted_roots[0]
|
||
leaves = [bytes.fromhex(r) for r in sorted_roots]
|
||
return MerkleTree.build(leaves).root.hex()
|
||
|
||
|
||
def query(
|
||
*,
|
||
question: str,
|
||
qa_db: Path,
|
||
chat_client: ChatClient,
|
||
model_id: str,
|
||
revision: str = "",
|
||
quantization: str = "",
|
||
shards_dir: Path | None = None,
|
||
single_db: Path | None = None,
|
||
top_k: int = 8,
|
||
over_fetch: int = 32,
|
||
max_context_chars: int = 60000,
|
||
policy: dict | None = None,
|
||
chain: str = "private",
|
||
fidelity: str | None = None,
|
||
burn_existing: bool = False,
|
||
retrieval_keywords: str | None = None,
|
||
) -> dict:
|
||
"""Answer `question` using the corpus. Cache to qa_db. Returns a result dict.
|
||
|
||
`fidelity` controls lookup tolerance — see ``FIDELITY_MODES`` in
|
||
``aborist.qa.keys``. ``"strict"`` checks only the cache_key
|
||
matching this call's ``policy["question_dedup"]``. ``"equivalence_class"``
|
||
(default) tries the primary cache_key first, then the alternate
|
||
dedup-mode cache_key as a fallback so a fast-cache agent can reuse
|
||
a record written under either mode. Result includes ``lookup_path``
|
||
naming which key matched (or ``"miss"`` when the LLM ran).
|
||
|
||
`burn_existing=True` deletes the matching live providence_cache row
|
||
(under the primary dedup-mode cache_key) BEFORE the cache lookup,
|
||
forcing a fresh inference. Each burn writes a ``providence_burn``
|
||
audit event. Test-ergonomic: run `make query Q=... BURN=1` after
|
||
tweaking a knob to see the new behavior without finding cache_keys
|
||
by hand. Result includes ``burned_existing`` reporting how many
|
||
rows were deleted (0 or 1 for the primary key; the equivalence-
|
||
class fallback key is left alone so prior alt-mode records stay
|
||
historic).
|
||
|
||
`retrieval_keywords` augments the FTS5 search and title-filter
|
||
token set with operator-supplied keywords WITHOUT changing what
|
||
the LLM sees as its question or what the verifier checks.
|
||
Empirically observed 2026-05-01: long discursive questions like
|
||
'what technology is currently or soon available which may enable
|
||
one person to reconstruct another person's thoughts...' under-
|
||
retrieve because their content tokens get diluted by template
|
||
phrasing. Appending domain keywords ('transcranial knowledge
|
||
acquisition') narrows OR-mode FTS5 to the topical article
|
||
(Neurotechnology) and lifts the verdict from HYBRID to STRICT.
|
||
|
||
Keywords do NOT enter ``question_hash`` directly, but they DO
|
||
change which sources get chosen — and that re-routes the
|
||
``context_root`` and ``conversation_hash`` components of
|
||
``cache_key``. Two calls with the same question and different
|
||
keywords therefore land under different cache_keys (different
|
||
contexts, different cached records — correctly so). Pair with
|
||
``burn_existing=True`` to force fresh inference when iterating
|
||
on keyword sets.
|
||
"""
|
||
policy = policy or DEFAULT_QUERY_POLICY
|
||
if fidelity is None:
|
||
fidelity = policy.get("fidelity", DEFAULT_FIDELITY)
|
||
if fidelity not in FIDELITY_MODES:
|
||
raise ValueError(
|
||
f"fidelity must be one of {FIDELITY_MODES}, got {fidelity!r}"
|
||
)
|
||
answer_mode = policy.get("answer_mode", DEFAULT_ANSWER_MODE)
|
||
if answer_mode not in ANSWER_MODES:
|
||
raise ValueError(
|
||
f"policy['answer_mode'] must be one of {ANSWER_MODES}, got {answer_mode!r}"
|
||
)
|
||
t_start = time.monotonic()
|
||
|
||
# 1. Search.
|
||
#
|
||
# Retrieval-only query string: question + operator-supplied
|
||
# ``retrieval_keywords`` (a hint, never part of the cache_key /
|
||
# LLM prompt / verifier surface). When the user passes
|
||
# `--retrieval-keywords "transcranial knowledge acquisition"`,
|
||
# only the FTS5 MATCH and title-filter token set see those
|
||
# tokens; the question text fed to the LLM and to question_hash
|
||
# stays untouched.
|
||
retrieval_query = question
|
||
if retrieval_keywords and retrieval_keywords.strip():
|
||
retrieval_query = f"{question} {retrieval_keywords.strip()}"
|
||
t_phase = time.monotonic()
|
||
hits = _search_corpus(shards_dir, single_db, retrieval_query, over_fetch)
|
||
if not hits:
|
||
return {
|
||
"status": "no_sources",
|
||
"msg": "FTS5 search returned no hits",
|
||
"timings": {
|
||
"search_ms": _ms_since(t_phase),
|
||
"total_ms": _ms_since(t_start),
|
||
},
|
||
}
|
||
core_match_roots = getattr(hits, "_core_match_roots", set())
|
||
root_to_shard = getattr(hits, "_root_to_shard", {})
|
||
qtokens_lower = {t.lower() for t in _title_query_tokens(retrieval_query)}
|
||
|
||
def _body_density_check(h) -> bool:
|
||
# Lazy per-hit check: open the shard, count token mentions in this doc.
|
||
sp = root_to_shard.get(h.document_root) or h.shard_path
|
||
if not sp:
|
||
return False
|
||
c = connect(sp)
|
||
try:
|
||
return _body_density_passes(c, h.document_root, qtokens_lower)
|
||
finally:
|
||
c.close()
|
||
|
||
hits = _rerank(
|
||
hits,
|
||
retrieval_query,
|
||
core_match_roots=core_match_roots,
|
||
body_density_check=_body_density_check,
|
||
)
|
||
search_ms = _ms_since(t_phase)
|
||
|
||
# 2. Pull doc texts within budget.
|
||
#
|
||
# Per-source cap so a single huge document can't monopolize the
|
||
# context window. Without this, a top-ranked bibliography page
|
||
# (e.g. List_of_Batman_comics at 80 KB) consumes the entire 60 KB
|
||
# budget at hit #1 and every subsequent doc is dropped with
|
||
# char_budget <= 0 — even when the bio article is hit #2 with
|
||
# the actual answer. By default we cap each source at
|
||
# `max_context_chars // top_k` so all top_k hits land in context.
|
||
# Total context ≤ max_context_chars by construction.
|
||
t_phase = time.monotonic()
|
||
chosen: list[_Hit] = []
|
||
context_parts: list[str] = []
|
||
per_source_cap = max(1, max_context_chars // max(1, top_k))
|
||
char_budget = max_context_chars
|
||
for h in hits[:top_k]:
|
||
text = _load_doc_text(h.shard_path, h.document_root)
|
||
if not text:
|
||
continue
|
||
# source_role is already set by _rerank_by_source_role; reuse
|
||
# it for the per-source cap. Primary answer source gets 2× the
|
||
# baseline cap, noisy/sequel get 0.5×, secondary & background
|
||
# get 1×. Total context still bounded by `char_budget`.
|
||
weight = SOURCE_ROLE_BUDGET_WEIGHTS.get(h.source_role, 1.0)
|
||
hit_cap = max(1, int(per_source_cap * weight))
|
||
if len(text) > hit_cap:
|
||
text = text[:hit_cap]
|
||
if len(text) > char_budget:
|
||
text = text[:char_budget]
|
||
if not text:
|
||
continue
|
||
context_parts.append(
|
||
f"=== Source: {h.document_uri} ===\n{text}"
|
||
)
|
||
chosen.append(h)
|
||
char_budget -= len(text)
|
||
if char_budget <= 0:
|
||
break
|
||
|
||
if not chosen:
|
||
return {
|
||
"status": "no_sources",
|
||
"msg": "top-k hits had cold or empty content",
|
||
"timings": {
|
||
"search_ms": search_ms,
|
||
"context_ms": _ms_since(t_phase),
|
||
"total_ms": _ms_since(t_start),
|
||
},
|
||
}
|
||
context = "\n\n".join(context_parts)
|
||
|
||
# Wikitext → prose before the LLM sees it. The model can then quote
|
||
# verbatim against the prose form; the verifier compares like-against-
|
||
# like. Idempotent if context is already plain prose. Gated on
|
||
# policy["base_version"] so this is part of governance_policy_hash.
|
||
if policy.get("base_version") and _wikitext_to_base is not None:
|
||
context = _wikitext_to_base(context)
|
||
context_ms = _ms_since(t_phase)
|
||
|
||
# 3. Build messages + hashes. Branch on answer_mode:
|
||
# "quote" (default) raw sources block + verbatim-quote rules.
|
||
# "claim_lattice" one evidence object per chosen source, each
|
||
# labeled with a content-addressed evidence_id;
|
||
# model emits JSON referencing IDs.
|
||
evidence_map = []
|
||
if answer_mode == "claim_lattice_pointer":
|
||
# G0.1 — per-chunk evidence granularity. Each retrieved source
|
||
# contributes ONE evidence object per chunk (up to the
|
||
# role-weighted per_source_cap budget) instead of one
|
||
# whole-doc span.
|
||
#
|
||
# G0.3 — query-relevance chunk ordering. Within each source,
|
||
# chunks are ranked by (distinct_query_tokens_present,
|
||
# total_mentions, chunk_idx_asc) so the chunk that most
|
||
# textually supports the question gets the lowest pointer id
|
||
# and lands at the top of the per-source evidence stack. Without
|
||
# this re-rank, Hermes-3-8B lazy-anchors on the first few
|
||
# chunks regardless of relevance — burying the actual answer
|
||
# paragraph behind irrelevant article-header text. Soft signal
|
||
# only (token overlap; no embeddings) so it stays out of the
|
||
# proof path; the verifier still runs the same hard checks
|
||
# against whatever the model picked.
|
||
qtokens_stem_for_chunks = {
|
||
_stem_token_for_match(t.lower())
|
||
for t in _title_query_tokens(question)
|
||
}
|
||
chunks_for_map: list[dict] = []
|
||
for h in chosen:
|
||
doc_chunks = _load_doc_chunks(h.shard_path, h.document_root)
|
||
if not doc_chunks:
|
||
continue
|
||
weight = SOURCE_ROLE_BUDGET_WEIGHTS.get(h.source_role, 1.0)
|
||
hit_cap = max(1, int(per_source_cap * weight))
|
||
# Score each chunk by query-token overlap. Tuple sort:
|
||
# distinct present DESC, total mentions DESC, doc order ASC.
|
||
scored = []
|
||
for chunk_idx, leaf_hash, span in doc_chunks:
|
||
distinct, total = _chunk_query_relevance(
|
||
span, qtokens_stem_for_chunks
|
||
)
|
||
scored.append(
|
||
(chunk_idx, leaf_hash, span, distinct, total)
|
||
)
|
||
scored.sort(key=lambda r: (-r[3], -r[4], r[0]))
|
||
# Greedy: fill the per-source budget with relevance-ranked
|
||
# chunks. If a single chunk exceeds what's left, truncate
|
||
# that one chunk and stop. Total context across the source
|
||
# stays bounded by hit_cap, same shape as the prose path.
|
||
# Per-source chunk cap also bounds chunk COUNT (in addition
|
||
# to char budget) so an encyclopedic article doesn't inflate
|
||
# the evidence catalog into E1-E26 territory and induce
|
||
# mega-claim failures.
|
||
max_chunks = max(1, int(policy.get(
|
||
"claim_lattice_max_chunks_per_source", 2
|
||
)))
|
||
spent = 0
|
||
chunks_used = 0
|
||
for chunk_idx, leaf_hash, span, _d, _t in scored:
|
||
if spent >= hit_cap or chunks_used >= max_chunks:
|
||
break
|
||
if policy.get("base_version") and _wikitext_to_base is not None:
|
||
span = _wikitext_to_base(span)
|
||
remaining = hit_cap - spent
|
||
if len(span) > remaining:
|
||
span = span[:remaining]
|
||
if not span:
|
||
break
|
||
chunks_for_map.append({
|
||
"source_root": h.document_root,
|
||
"document_uri": h.document_uri,
|
||
"title": h.title,
|
||
"chunk_idx": chunk_idx,
|
||
"chunk_root": leaf_hash,
|
||
"span": span,
|
||
"source_role": h.source_role,
|
||
})
|
||
spent += len(span)
|
||
chunks_used += 1
|
||
evidence_map = build_evidence_map(chunks_for_map)
|
||
sys_prompt = policy["claim_lattice_system_prompt"]
|
||
grounding_reminder = policy.get("claim_lattice_grounding_reminder")
|
||
rendered_evidence = render_evidence_map(evidence_map)
|
||
|
||
def _user_payload(q: str) -> str:
|
||
return f"EVIDENCE:\n\n{rendered_evidence}\n\n---\n\nQUESTION: {q}"
|
||
elif answer_mode == "claim_lattice":
|
||
# JSON variant — same evidence-map construction as the pointer
|
||
# path, but blocks are labeled with content-addressed
|
||
# ``evidence_id`` (long hex) since the model emits IDs in JSON.
|
||
# The lenient pre-parser in verify_claim_lattice_json keeps the
|
||
# path survivable on inference paths without grammar guidance;
|
||
# vLLM ``guided_json`` (passed via extra_body below) eliminates
|
||
# SCHEMA_INVALID failures at sampling time when available.
|
||
qtokens_stem_for_chunks = {
|
||
_stem_token_for_match(t.lower())
|
||
for t in _title_query_tokens(question)
|
||
}
|
||
chunks_for_map: list[dict] = []
|
||
for h in chosen:
|
||
doc_chunks = _load_doc_chunks(h.shard_path, h.document_root)
|
||
if not doc_chunks:
|
||
continue
|
||
weight = SOURCE_ROLE_BUDGET_WEIGHTS.get(h.source_role, 1.0)
|
||
hit_cap = max(1, int(per_source_cap * weight))
|
||
scored = []
|
||
for chunk_idx, leaf_hash, span in doc_chunks:
|
||
distinct, total = _chunk_query_relevance(
|
||
span, qtokens_stem_for_chunks
|
||
)
|
||
scored.append((chunk_idx, leaf_hash, span, distinct, total))
|
||
scored.sort(key=lambda r: (-r[3], -r[4], r[0]))
|
||
max_chunks = max(1, int(policy.get(
|
||
"claim_lattice_max_chunks_per_source", 2
|
||
)))
|
||
spent = 0
|
||
chunks_used = 0
|
||
for chunk_idx, leaf_hash, span, _d, _t in scored:
|
||
if spent >= hit_cap or chunks_used >= max_chunks:
|
||
break
|
||
if policy.get("base_version") and _wikitext_to_base is not None:
|
||
span = _wikitext_to_base(span)
|
||
remaining = hit_cap - spent
|
||
if len(span) > remaining:
|
||
span = span[:remaining]
|
||
if not span:
|
||
break
|
||
chunks_for_map.append({
|
||
"source_root": h.document_root,
|
||
"document_uri": h.document_uri,
|
||
"title": h.title,
|
||
"chunk_idx": chunk_idx,
|
||
"chunk_root": leaf_hash,
|
||
"span": span,
|
||
"source_role": h.source_role,
|
||
})
|
||
spent += len(span)
|
||
chunks_used += 1
|
||
evidence_map = build_evidence_map(chunks_for_map)
|
||
sys_prompt = policy.get(
|
||
"claim_lattice_json_system_prompt",
|
||
policy["claim_lattice_system_prompt"],
|
||
)
|
||
grounding_reminder = policy.get(
|
||
"claim_lattice_json_grounding_reminder",
|
||
policy.get("claim_lattice_grounding_reminder"),
|
||
)
|
||
rendered_evidence = render_evidence_map_for_json(evidence_map)
|
||
|
||
def _user_payload(q: str) -> str:
|
||
return f"EVIDENCE:\n\n{rendered_evidence}\n\n---\n\nQUESTION: {q}"
|
||
else:
|
||
sys_prompt = policy["system_prompt"]
|
||
grounding_reminder = policy.get("grounding_reminder")
|
||
|
||
def _user_payload(q: str) -> str:
|
||
return f"Sources:\n\n{context}\n\n---\n\nQuestion: {q}"
|
||
|
||
messages = [{"role": "system", "content": sys_prompt}]
|
||
if grounding_reminder:
|
||
messages.append({"role": "user", "content": grounding_reminder})
|
||
messages.append({"role": "user", "content": _user_payload(question)})
|
||
|
||
# Capacity metrics. Char-level for now — a fast model-agnostic proxy
|
||
# for prompt size (rule of thumb: ~4 chars/token for English prose,
|
||
# ~2.5 for JSON-evidence-block heavy contexts). Surfaced in the
|
||
# result dict so the bench can correlate strict-rate with input
|
||
# size and the operator can tell at a glance whether a STRICT
|
||
# verdict came from a tight 5KB prompt or a 50KB context-stuffed
|
||
# one. Never enters cache_key — these are runtime measurements,
|
||
# not policy.
|
||
if answer_mode in ("claim_lattice_pointer", "claim_lattice"):
|
||
evidence_or_context_chars = len(rendered_evidence)
|
||
else:
|
||
evidence_or_context_chars = len(context)
|
||
prompt_chars = {
|
||
"system_prompt": len(sys_prompt or ""),
|
||
"grounding_reminder": len(grounding_reminder or ""),
|
||
"user_question": len(question or ""),
|
||
"evidence_or_context": evidence_or_context_chars,
|
||
"messages_total": sum(len(m["content"]) for m in messages),
|
||
}
|
||
|
||
context_root = _context_root([h.document_root for h in chosen])
|
||
mhash = model_profile_hash(model_id, revision, quantization)
|
||
|
||
# Dedup-mode-aware cache_key build. For each mode we substitute the
|
||
# mode's canonical question form into the user message used for
|
||
# `conversation_hash` (LLM still sees the verbatim question), AND we
|
||
# vary `policy["question_dedup"]` to match the mode so
|
||
# `governance_policy_hash` matches what an agent under that mode
|
||
# would have written. This makes cross-silo fallback work: a
|
||
# strict-policy agent looking up with equivalence_class fidelity can
|
||
# find a record written by an equivalence_class-policy agent.
|
||
def _ckey_for_mode(mode: str) -> str:
|
||
canon_q = canonical_question(question, mode=mode)
|
||
canon_msgs = list(messages[:-1]) + [
|
||
{"role": "user", "content": _user_payload(canon_q)},
|
||
]
|
||
policy_variant = dict(policy, question_dedup=mode)
|
||
return cache_key(
|
||
context_root,
|
||
question_hash(question, mode=mode),
|
||
mhash,
|
||
conversation_hash(canon_msgs),
|
||
governance_policy_hash(policy_variant),
|
||
SCHEMA_VERSION,
|
||
CANONICALIZATION_VERSION,
|
||
CHUNKING_VERSION,
|
||
verifier_policy_hash(policy_variant),
|
||
)
|
||
|
||
ghash = governance_policy_hash(policy) # for the legacy INSERT below
|
||
|
||
primary_dedup = policy.get("question_dedup", DEFAULT_QUESTION_DEDUP)
|
||
if primary_dedup not in QUESTION_DEDUP_MODES:
|
||
raise ValueError(
|
||
f"policy['question_dedup'] must be one of {QUESTION_DEDUP_MODES}, "
|
||
f"got {primary_dedup!r}"
|
||
)
|
||
# Re-derive the per-mode hashes for use in the INSERT below. The legacy
|
||
# INSERT references qhash/chash by name; _ckey_for_mode already builds
|
||
# them but doesn't expose the intermediates.
|
||
qhash = question_hash(question, mode=primary_dedup)
|
||
canonical_q_primary = canonical_question(question, mode=primary_dedup)
|
||
canonical_messages_primary = list(messages[:-1]) + [
|
||
{"role": "user", "content": _user_payload(canonical_q_primary)},
|
||
]
|
||
chash = conversation_hash(canonical_messages_primary)
|
||
primary_ckey = _ckey_for_mode(primary_dedup)
|
||
ckey = primary_ckey # keep the legacy name in the rest of the function
|
||
|
||
qa_conn = connect(qa_db)
|
||
burned_existing = 0
|
||
try:
|
||
# 3.5. Optional pre-lookup burn: deletes the matching live row
|
||
# under the primary cache_key so the lookup misses & a fresh
|
||
# inference runs. Test-ergonomic — pass `burn_existing=True`
|
||
# (or `make query Q=... BURN=1`) after tweaking a knob to see
|
||
# the new behavior. The equivalence-class fallback key is
|
||
# deliberately NOT touched: prior alt-mode records stay as
|
||
# historic witnesses.
|
||
if burn_existing:
|
||
existing = qa_conn.execute(
|
||
"SELECT cache_key, audit_mode, n_verified, "
|
||
" falsification_state, question_text "
|
||
"FROM providence_cache WHERE cache_key = ?",
|
||
(primary_ckey,),
|
||
).fetchone()
|
||
if existing is not None:
|
||
with transaction(qa_conn):
|
||
qa_conn.execute(
|
||
"DELETE FROM providence_cache WHERE cache_key = ?",
|
||
(primary_ckey,),
|
||
)
|
||
append_audit(
|
||
qa_conn,
|
||
event_type="providence_burn",
|
||
subject_root=primary_ckey,
|
||
body={
|
||
"cache_key": primary_ckey,
|
||
"burned_audit_mode": existing["audit_mode"],
|
||
"burned_n_verified": int(existing["n_verified"] or 0),
|
||
"burned_state": existing["falsification_state"],
|
||
"question_text": existing["question_text"],
|
||
"reason": "query --burn (test-ergonomic mid-query bust)",
|
||
},
|
||
)
|
||
burned_existing = 1
|
||
|
||
# 4. Cache lookup. Try the primary dedup-mode cache_key first.
|
||
# If fidelity allows fallback AND the alternate dedup mode
|
||
# produces a different cache_key, try that too — lets a
|
||
# fast-cache agent reuse a record written under either mode.
|
||
t_phase = time.monotonic()
|
||
cached = qa_conn.execute(
|
||
"SELECT * FROM providence_cache "
|
||
"WHERE cache_key = ? AND falsification_state = 'live'",
|
||
(primary_ckey,),
|
||
).fetchone()
|
||
hit_ckey = primary_ckey
|
||
lookup_path = primary_dedup if cached is not None else None
|
||
if cached is None and fidelity == "equivalence_class":
|
||
other_mode = (
|
||
"equivalence_class" if primary_dedup == "strict" else "strict"
|
||
)
|
||
other_ckey = _ckey_for_mode(other_mode)
|
||
if other_ckey != primary_ckey:
|
||
cached = qa_conn.execute(
|
||
"SELECT * FROM providence_cache "
|
||
"WHERE cache_key = ? AND falsification_state = 'live'",
|
||
(other_ckey,),
|
||
).fetchone()
|
||
if cached is not None:
|
||
hit_ckey = other_ckey
|
||
lookup_path = f"{other_mode}_fallback"
|
||
cache_lookup_ms = _ms_since(t_phase)
|
||
if cached is not None:
|
||
now = int(time.time())
|
||
with transaction(qa_conn):
|
||
qa_conn.execute(
|
||
"UPDATE providence_cache "
|
||
"SET hit_count = hit_count + 1, last_hit_at = ? "
|
||
"WHERE cache_key = ?",
|
||
(now, hit_ckey),
|
||
)
|
||
return {
|
||
"status": "cache_hit",
|
||
"audit_mode": cached["audit_mode"],
|
||
"cache_key": hit_ckey,
|
||
"lookup_path": lookup_path,
|
||
"burned_existing": burned_existing,
|
||
"context_root": context_root,
|
||
"answer_text": cached["answer_text"],
|
||
"sources": json.loads(cached["merkle_proof"])["sources"],
|
||
"n_quotes": cached["n_quotes"],
|
||
"n_verified": cached["n_verified"],
|
||
"verifier_method": cached["verifier_method"],
|
||
"unverified_quotes": (
|
||
json.loads(cached["unverified_quotes"])
|
||
if cached["unverified_quotes"]
|
||
else []
|
||
),
|
||
# Cache schema doesn't carry the partially-verified split —
|
||
# those claims were folded into unverified_quotes pre-2026-
|
||
# 04-30. Cache hits surface an empty partial list; new
|
||
# writes populate it correctly. Acceptable degradation
|
||
# since governance_policy_hash invalidated prior records.
|
||
"partially_verified_quotes": [],
|
||
"prompt_chars": prompt_chars,
|
||
"answer_chars": len(cached["answer_text"] or ""),
|
||
"timings": {
|
||
"search_ms": search_ms,
|
||
"context_ms": context_ms,
|
||
"cache_lookup_ms": cache_lookup_ms,
|
||
"llm_ms": None,
|
||
"persist_ms": None,
|
||
"total_ms": _ms_since(t_start),
|
||
},
|
||
}
|
||
|
||
# 5. Cache miss — call LLM. JSON mode passes guided_json via
|
||
# extra_body so vLLM constrains output to the schema at sampling
|
||
# time. Non-vLLM endpoints silently drop the field; the lenient
|
||
# pre-parser in the verifier handles whatever drift remains.
|
||
extra_body: dict | None = None
|
||
stop_seqs: list[str] | None = None
|
||
if answer_mode == "claim_lattice" and policy.get(
|
||
"claim_lattice_use_guided_json", True
|
||
):
|
||
extra_body = {"guided_json": CLAIM_LATTICE_JSON_SCHEMA}
|
||
if answer_mode == "claim_lattice":
|
||
# JSON-mode token-runaway guard — see runner.py for the
|
||
# full rationale. Stops generation on a blank line so
|
||
# post-JSON whitespace spam doesn't blow max_tokens.
|
||
stop_seqs = list(policy.get(
|
||
"claim_lattice_json_stop_sequences", ["\n\n"]
|
||
))
|
||
t_phase = time.monotonic()
|
||
raw_answer = chat_client.chat_completion(
|
||
messages,
|
||
model=model_id,
|
||
temperature=policy["temperature"],
|
||
max_tokens=policy["max_tokens"],
|
||
top_p=policy.get("top_p", 1.0),
|
||
extra_body=extra_body,
|
||
stop=stop_seqs,
|
||
)
|
||
llm_ms = _ms_since(t_phase)
|
||
|
||
# 5b. Faithfulness check. Branch on answer_mode:
|
||
# "quote" substring-verify quoted spans against context.
|
||
# Optional repair loop (mechanical + reprompt).
|
||
# "claim_lattice" deterministic checks on the JSON output:
|
||
# evidence_id resolution, source_role allowlist,
|
||
# manual-quote prohibition. NO repair loop —
|
||
# one-shot benchmark discipline.
|
||
repair_changes: list[dict] = []
|
||
pre_repair_verdict: dict | None = None
|
||
|
||
if answer_mode == "claim_lattice_pointer":
|
||
verdict = verify_claim_lattice(
|
||
raw_answer,
|
||
evidence_map,
|
||
allowed_source_roles=tuple(
|
||
policy.get(
|
||
"claim_lattice_allowed_source_roles",
|
||
[
|
||
"primary_answer_source",
|
||
"secondary_context_source",
|
||
"background_source",
|
||
"unclassified",
|
||
],
|
||
)
|
||
),
|
||
max_pointers_per_claim=int(policy.get(
|
||
"claim_lattice_max_pointers_per_claim", 2
|
||
)),
|
||
min_citation_coverage=float(policy.get(
|
||
"claim_lattice_min_citation_coverage", 0.30
|
||
)),
|
||
min_claim_content_tokens=int(policy.get(
|
||
"claim_lattice_min_claim_content_tokens", 3
|
||
)),
|
||
lazy_anchor_demote_threshold=float(policy.get(
|
||
"claim_lattice_lazy_anchor_demote_threshold", 0.5
|
||
)),
|
||
lazy_anchor_demote_min_pairs=int(policy.get(
|
||
"claim_lattice_lazy_anchor_demote_min_pairs", 3
|
||
)),
|
||
max_claims_per_answer=int(policy.get(
|
||
"claim_lattice_max_claims_per_answer", 12
|
||
)),
|
||
)
|
||
rendered = verdict["rendered_text"]
|
||
answer_text = rendered if rendered else raw_answer
|
||
elif answer_mode == "claim_lattice":
|
||
verdict = verify_claim_lattice_json(
|
||
raw_answer,
|
||
evidence_map,
|
||
allowed_source_roles=tuple(
|
||
policy.get(
|
||
"claim_lattice_allowed_source_roles",
|
||
[
|
||
"primary_answer_source",
|
||
"secondary_context_source",
|
||
"background_source",
|
||
"unclassified",
|
||
],
|
||
)
|
||
),
|
||
max_evidence_per_claim=int(policy.get(
|
||
"claim_lattice_max_pointers_per_claim", 2
|
||
)),
|
||
min_citation_coverage=float(policy.get(
|
||
"claim_lattice_min_citation_coverage", 0.30
|
||
)),
|
||
max_claims_per_answer=int(policy.get(
|
||
"claim_lattice_max_claims_per_answer", 12
|
||
)),
|
||
)
|
||
rendered = verdict["rendered_text"]
|
||
answer_text = rendered if rendered else raw_answer
|
||
else:
|
||
answer_text = raw_answer
|
||
verdict = verify_quotes(
|
||
answer_text,
|
||
context,
|
||
entity_policy=policy.get("entity_policy", "hybrid"),
|
||
proximity_n=policy.get("entity_proximity_n", 3),
|
||
proximity_window=policy.get("entity_proximity_window", 300),
|
||
)
|
||
|
||
def _verify(text: str) -> dict:
|
||
return verify_quotes(
|
||
text,
|
||
context,
|
||
entity_policy=policy.get("entity_policy", "hybrid"),
|
||
proximity_n=policy.get("entity_proximity_n", 3),
|
||
proximity_window=policy.get("entity_proximity_window", 300),
|
||
)
|
||
|
||
if (
|
||
policy.get("repair_enabled")
|
||
and verdict["audit_mode"] != "STRICT"
|
||
and verdict.get("unverified_quotes")
|
||
):
|
||
# Tier 1: mechanical (deterministic, no extra LLM call).
|
||
repair_result = mechanical_repair(
|
||
answer_text, verdict["unverified_quotes"], context
|
||
)
|
||
if repair_result["changes"]:
|
||
new_verdict = _verify(repair_result["repaired_text"])
|
||
if new_verdict["n_verified"] >= verdict["n_verified"]:
|
||
pre_repair_verdict = verdict
|
||
answer_text = repair_result["repaired_text"]
|
||
verdict = new_verdict
|
||
repair_changes = list(repair_result["changes"])
|
||
|
||
# Tier 2: re-prompt feedback (one extra LLM call max).
|
||
max_reprompts = int(policy.get("repair_max_reprompts", 0))
|
||
for _ in range(max_reprompts):
|
||
if (
|
||
verdict["audit_mode"] == "STRICT"
|
||
or not verdict.get("unverified_quotes")
|
||
):
|
||
break
|
||
new_text = reprompt_repair(
|
||
chat_client=chat_client,
|
||
model_id=model_id,
|
||
original_messages=messages,
|
||
original_answer=answer_text,
|
||
failed_quotes=verdict["unverified_quotes"],
|
||
policy=policy,
|
||
)
|
||
if not new_text:
|
||
break
|
||
new_verdict = _verify(new_text)
|
||
if new_verdict["n_verified"] > verdict["n_verified"]:
|
||
if pre_repair_verdict is None:
|
||
pre_repair_verdict = verdict
|
||
answer_text = new_text
|
||
verdict = new_verdict
|
||
repair_changes.append({
|
||
"action": "reprompt_rewrite",
|
||
"diagnosis": "model_feedback_loop",
|
||
})
|
||
else:
|
||
break
|
||
|
||
unverified_blob = (
|
||
json.dumps(verdict["unverified_quotes"], separators=(",", ":"))
|
||
if verdict["unverified_quotes"]
|
||
else None
|
||
)
|
||
|
||
# 6. Persist record + audit event.
|
||
t_phase = time.monotonic()
|
||
proof_obj = {
|
||
"context_root": context_root,
|
||
"sources": [
|
||
{
|
||
"document_root": h.document_root,
|
||
"document_uri": h.document_uri,
|
||
"title": h.title,
|
||
"score": h.score,
|
||
"chunk_idx": h.chunk_idx,
|
||
"shard": Path(h.shard_path).name,
|
||
"source_role": h.source_role,
|
||
}
|
||
for h in chosen
|
||
],
|
||
}
|
||
proof_blob = json.dumps(proof_obj, separators=(",", ":"))
|
||
|
||
# Per-run Merkle-DAG. Quote mode keeps the original 7-stage
|
||
# shape (question / retrieval / context / prompt / answer /
|
||
# verify / final_label). Pointer mode swaps to the 9-stage CTI
|
||
# shape — context drops out, answer splits into raw_answer /
|
||
# parsed_claim_lattice / render — and threads violations +
|
||
# answer_mode into the verify/final_label payloads.
|
||
ev_root = evidence_map_root(evidence_map) if evidence_map else None
|
||
parsed_lattice = None
|
||
is_lattice_mode = answer_mode in ("claim_lattice_pointer", "claim_lattice")
|
||
if is_lattice_mode:
|
||
evidence_id_pairs = verdict.get("evidence_id_pairs") or []
|
||
parsed_lattice = [
|
||
{
|
||
"claim_text": cs.get("text", ""),
|
||
"evidence_ids": evidence_id_pairs[i] if i < len(evidence_id_pairs) else [],
|
||
}
|
||
for i, cs in enumerate(verdict.get("claim_statuses") or [])
|
||
]
|
||
run_dag = build_run_dag(
|
||
question_hash=qhash,
|
||
sources=proof_obj["sources"],
|
||
context_root=context_root,
|
||
conversation_hash=chash,
|
||
answer_text=answer_text,
|
||
audit_mode=verdict["audit_mode"],
|
||
verifier_method=verdict["verifier_method"],
|
||
n_quotes=verdict["n_quotes"],
|
||
n_verified=verdict["n_verified"],
|
||
claim_statuses=verdict.get("claim_statuses", []),
|
||
lookup_path="miss",
|
||
evidence_map_root=ev_root,
|
||
answer_mode=answer_mode if answer_mode != "quote" else None,
|
||
violations=verdict.get("violations"),
|
||
raw_answer_text=raw_answer if is_lattice_mode else None,
|
||
parsed_lattice=parsed_lattice,
|
||
rendered_text=answer_text if is_lattice_mode else None,
|
||
)
|
||
run_dag_blob = json.dumps(run_dag, separators=(",", ":"))
|
||
|
||
now = int(time.time())
|
||
with transaction(qa_conn):
|
||
# Record the repair event BEFORE the providence_query event so
|
||
# the audit chain shows: repair-happened, THEN we wrote the
|
||
# final record. Repair body links pre→post verdicts so an
|
||
# auditor can reconstruct what changed.
|
||
if repair_changes and pre_repair_verdict is not None:
|
||
append_audit(
|
||
qa_conn,
|
||
event_type="providence_repair",
|
||
subject_root=ckey,
|
||
body={
|
||
"kind": "mechanical",
|
||
"n_changes": len(repair_changes),
|
||
"changes": repair_changes,
|
||
"pre_audit_mode": pre_repair_verdict["audit_mode"],
|
||
"post_audit_mode": verdict["audit_mode"],
|
||
"pre_n_verified": pre_repair_verdict["n_verified"],
|
||
"post_n_verified": verdict["n_verified"],
|
||
},
|
||
ts=now,
|
||
)
|
||
event_hash = append_audit(
|
||
qa_conn,
|
||
event_type="providence_query",
|
||
subject_root=ckey,
|
||
body={
|
||
"context_root": context_root,
|
||
"n_sources": len(chosen),
|
||
"model_id": model_id,
|
||
"revision": revision,
|
||
"quantization": quantization,
|
||
"answer_chars": len(answer_text),
|
||
"context_chars": len(context),
|
||
"audit_mode": verdict["audit_mode"],
|
||
"n_quotes": verdict["n_quotes"],
|
||
"n_verified": verdict["n_verified"],
|
||
"verifier_method": verdict["verifier_method"],
|
||
},
|
||
ts=now,
|
||
)
|
||
qa_conn.execute(
|
||
"INSERT INTO providence_cache "
|
||
"(cache_key, source_root, document_uri, question_hash, question_text, "
|
||
" answer_text, merkle_proof, model_profile_hash, conversation_hash, "
|
||
" governance_policy_hash, schema_version, canonicalization_version, "
|
||
" chunking_version, falsification_state, chain, audit_event_hash, "
|
||
" created_at, hit_count, audit_mode, n_quotes, n_verified, "
|
||
" unverified_quotes, verifier_method, run_dag_root, run_dag_blob) "
|
||
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, 'live', ?, ?, ?, 0, "
|
||
" ?, ?, ?, ?, ?, ?, ?)",
|
||
(
|
||
ckey,
|
||
context_root,
|
||
"corpus://multi-source",
|
||
qhash,
|
||
question,
|
||
answer_text,
|
||
proof_blob,
|
||
mhash,
|
||
chash,
|
||
ghash,
|
||
SCHEMA_VERSION,
|
||
CANONICALIZATION_VERSION,
|
||
CHUNKING_VERSION,
|
||
chain,
|
||
event_hash,
|
||
now,
|
||
verdict["audit_mode"],
|
||
verdict["n_quotes"],
|
||
verdict["n_verified"],
|
||
unverified_blob,
|
||
verdict["verifier_method"],
|
||
run_dag["root"],
|
||
run_dag_blob,
|
||
),
|
||
)
|
||
finally:
|
||
qa_conn.close()
|
||
|
||
persist_ms = _ms_since(t_phase)
|
||
# Pull the verify node's failure_stage out of the run_dag for the result.
|
||
failure_stage = next(
|
||
(n.get("hash") for n in run_dag["nodes"] if n["stage"] == "verify"),
|
||
None,
|
||
)
|
||
# The actual label is in the verify_payload, which we computed in
|
||
# localize_failure earlier — recompute for the result dict.
|
||
from aborist.qa.dag import localize_failure as _localize
|
||
failure_stage = _localize(
|
||
audit_mode=verdict["audit_mode"],
|
||
n_sources=len(chosen),
|
||
n_quotes=verdict["n_quotes"],
|
||
n_verified=verdict["n_verified"],
|
||
)
|
||
return {
|
||
"status": "cache_miss_then_written",
|
||
"audit_mode": verdict["audit_mode"],
|
||
"cache_key": ckey,
|
||
"run_dag_root": run_dag["root"],
|
||
"lookup_path": "miss",
|
||
"failure_stage": failure_stage,
|
||
"repair_changes": repair_changes,
|
||
"pre_repair_audit_mode": (
|
||
pre_repair_verdict["audit_mode"] if pre_repair_verdict else None
|
||
),
|
||
"burned_existing": burned_existing,
|
||
"context_root": context_root,
|
||
"answer_text": answer_text,
|
||
"sources": proof_obj["sources"],
|
||
"n_quotes": verdict["n_quotes"],
|
||
"n_verified": verdict["n_verified"],
|
||
"verifier_method": verdict["verifier_method"],
|
||
"unverified_quotes": verdict["unverified_quotes"],
|
||
"partially_verified_quotes": verdict.get("partially_verified_quotes") or [],
|
||
# Sidecar smell signals (claim_lattice mode only) — surfaced
|
||
# for the renderer; never persisted in providence_cache and
|
||
# never threaded into run_dag_root.
|
||
"pointer_id_distribution": verdict.get("pointer_id_distribution"),
|
||
"lazy_anchor_ratio": verdict.get("lazy_anchor_ratio"),
|
||
"prompt_chars": prompt_chars,
|
||
"answer_chars": len(answer_text or ""),
|
||
"timings": {
|
||
"search_ms": search_ms,
|
||
"context_ms": context_ms,
|
||
"cache_lookup_ms": cache_lookup_ms,
|
||
"llm_ms": llm_ms,
|
||
"persist_ms": persist_ms,
|
||
"total_ms": _ms_since(t_start),
|
||
},
|
||
}
|
||
|
||
|
||
def _ms_since(t: float) -> float:
|
||
"""Wall-time elapsed in milliseconds, rounded to 1 decimal."""
|
||
return round((time.monotonic() - t) * 1000, 1)
|