Pure-Python (stdlib only). Uses chunk-level corpus baseline so terms concentrated in fewer chunks outrank common ones. Default top-K = 16. Output cores are comma-separated keyword lists — extreme compression toward the tweet/haiku end of the planet metaphor. Same source can now carry both a first-sentence-v1 core AND a tfidf-keywords-v1 core, each derived independently and Merkle-signed back to the same surface. The 'contributing_chunk_indices' for TF-IDF is every chunk that contains at least one of the top-K keywords — proof binding remains honest and cryptographically tight. |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| test_distill.py | ||
| test_distill_recursive.py | ||
| test_evict.py | ||
| test_html_source.py | ||
| test_ingest.py | ||
| test_merkle.py | ||
| test_tfidf.py | ||