Phase 2 of the concepts/ layer. The Apr 27 commit (c6182ae) shipped hand-curated frozensets in aborist/qa/concepts.py with a TODO to "derive from Wikipedia's category graph or 'See also' sections" — that's this commit. Fox's 2026-05-01 critique landed it: a 7-entry list of arbitrary frozensets won't scale to a 3.47M-doc corpus, & adding domains shouldn't require a Python edit + commit + redeploy. Architecture: (1) Per-shard concept_relations SQLite table. Append-only, with UNIQUE (source_root, relation_kind, token, target, evidence_kind) so re-derivation is idempotent. Lives next to documents in each shard so mesh sync moves relations alongside the docs that derived them. A SECONDARY index — writes here NEVER affect document_root / chunk_root / cache_key, so backfilling is safe across the entire corpus. (2) aborist/concepts/ package: - store.py: add_concept_relation, concept_relations_for_token, purge_by_evidence_kind, list_evidence_kinds - query.py: cross-shard synonym_expand, rivalry_excluded; union-find collapse on synonym edges so partial pairs build full equivalence classes; mtime-keyed per-process LRU so retrieval doesn't re-walk shards on hot loops - seed.py: one-shot migration of legacy frozensets (8 groups incl. brain-tech) to evidence_kind='manual_legacy' rows under source_root='__legacy__concepts__' - extract.py: pluggable extractor registry. Built-in: link_reciprocity_synonym — for any reciprocal edge pair (A→B AND B→A) in the existing edges table, emit synonym edges between the docs' title-tokens. Works for Wikipedia (See-also bidirectional), HTML site internal-link clusters (russell.ballestrini.net pattern), or any link graph the corpus already encodes — no new crawler needed; the html_page parser already populates `edges` rows on ingest. (3) aborist/qa/concepts.py rewritten as a backwards-compat shim — same public API (synonym_expand, rivalry_excluded, has_compare_phrasing) so query.py call sites unchanged. shards_dir threaded through _filter_by_title_relevance & _search_corpus' synonym_expand calls. Without shards_dir (legacy 2-arg call shape), helpers degenerate to no-op — matches the behavior the frozenset code had when no group hit. (4) Cross-shard UNION view: concept_relations added to _SHARDABLE_TABLES in store.py so connect_query() exposes a unified view across all shards (same pattern as documents, chunks, providence_cache, etc.). Live-verified on the 4-shard 3.47M-doc corpus + the crawl_russell_ballestrini_net.db shard: Q: what technology are currently or soon available which may enable one person to reconstruct and understand some or a portion of another persons thoughts or ideas without speaking or sign language? Result: same as thebde1bd6in-memory frozenset version — Telepathy E4 cited alongside Videoconferencing E1/E2, POINTER-LINKED-PARTIAL 6/10. The DB-backed lookup reproduces the frozenset behavior byte-for-byte. Tests: 633 passed (was 624). 9 new concept tests covering DB-backed synonym/rivalry lookup, shards_dir=None degenerate behavior, brain-tech group seed, cache invalidation. Old tests that called helpers directly without shards_dir kept as no-op assertions (synonym_expand({"athlon"}) without shards_dir returns {"athlon"} unchanged). Backfill mechanics: existing live shards needed a one-shot `connect()` to auto-create the new concept_relations table (SCHEMA_SQL has CREATE TABLE IF NOT EXISTS). Re-derivation never mutates the Merkle tree — concept_relations is fully orthogonal to document_root / chunk_root. Cache_key dimensions are unaffected. Deferred (follow-on): - CLI commands: aborist concepts {seed,list,add,derive,purge} (currently fox runs the helpers via python -c) - Wikipedia See-also extractor (requires parsing wikitext sections beyond what's already in edges) - Wikipedia category extractor (requires reading Category: links from chunked wikitext)
50 lines
1.7 KiB
Python
50 lines
1.7 KiB
Python
"""Corpus-derived concept relations: synonyms, antonyms, rivalries, categories.
|
|
|
|
Replaces the hand-curated frozensets that lived in ``aborist.qa.concepts``
|
|
through April 2026 (commit c6182ae). The frozensets were Phase 1; this is
|
|
Phase 2.
|
|
|
|
The concept-relations layer is a **secondary index** over the existing
|
|
Merkle-committed corpus. Writes to ``concept_relations`` NEVER affect
|
|
``document_root``, ``chunk_root``, or ``cache_key`` — backfilling
|
|
relations is safe across the entire corpus without invalidating cached
|
|
answers or breaking audit chains.
|
|
|
|
Architecture:
|
|
|
|
- ``store.py`` — DB read/write helpers, append-only with UNIQUE-key idempotency
|
|
- ``extract.py`` — Extractor ABC + registry; per-source extractors plug in here
|
|
- ``query.py`` — Cross-shard ``synonyms_for(token)`` & ``rivalries_for(token)``
|
|
- ``seed.py`` — One-time migration of the legacy frozensets to manual rows
|
|
|
|
Public API for retrieval-time use (matches the legacy
|
|
``aborist.qa.concepts`` shape, so call sites in ``query.py`` keep working):
|
|
|
|
synonym_expand(tokens, *, shards_dir) -> set[str]
|
|
rivalry_excluded(tokens, *, shards_dir, compare_phrasing=False) -> set[str]
|
|
has_compare_phrasing(question) -> bool
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from aborist.concepts.query import (
|
|
has_compare_phrasing,
|
|
invalidate_cache,
|
|
rivalry_excluded,
|
|
synonym_expand,
|
|
)
|
|
from aborist.concepts.store import (
|
|
add_concept_relation,
|
|
concept_relations_for_token,
|
|
purge_by_evidence_kind,
|
|
)
|
|
|
|
__all__ = [
|
|
"add_concept_relation",
|
|
"concept_relations_for_token",
|
|
"has_compare_phrasing",
|
|
"invalidate_cache",
|
|
"purge_by_evidence_kind",
|
|
"rivalry_excluded",
|
|
"synonym_expand",
|
|
]
|