add Reverse RAG whitepaper, link from all site pages and uncloseai.js docs
This commit is contained in:
parent
b7ac2cbf3a
commit
d47e8900e7
28 changed files with 424 additions and 1 deletions
|
|
@ -74,6 +74,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -74,6 +74,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -74,6 +74,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -74,6 +74,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -75,6 +75,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -75,6 +75,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -75,6 +75,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -83,6 +83,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -75,6 +75,7 @@
|
|||
<li><a href="/inference.html" class="active">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -75,6 +75,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -75,6 +75,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -75,6 +75,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -75,6 +75,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -70,6 +70,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
<li><a href="/privacy-policy.html" class="active">Privacy Policy</a></li>
|
||||
|
|
|
|||
|
|
@ -75,6 +75,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
396
public/reverse-rag.html
Normal file
396
public/reverse-rag.html
Normal file
|
|
@ -0,0 +1,396 @@
|
|||
<!DOCTYPE html>
|
||||
<!--
|
||||
PUBLIC DOMAIN - NO LICENSE, NO WARRANTY
|
||||
Copyright 2025-2026 TimeHexOn & foxhop & russell@unturf
|
||||
https://www.permacomputer.com
|
||||
-->
|
||||
|
||||
<html lang="en">
|
||||
|
||||
<head>
|
||||
<meta charset="UTF-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||||
<meta name="theme-color" content="#43a047">
|
||||
<meta name="color-scheme" content="light dark">
|
||||
<title>Reverse Retrieval Augmented Generation | uncloseai.com</title>
|
||||
<meta name="description" content="Whitepaper: Reverse RAG, a client-side context injection technique that makes small language models punch above their weight by extracting live page content and injecting it directly into the conversation.">
|
||||
|
||||
<!-- PicoCSS -->
|
||||
<link rel="stylesheet" href="/css/pico.classless.min.css">
|
||||
<!-- ChunkFive Font -->
|
||||
<link rel="stylesheet" href="/css/chunkfive/stylesheet.css" type="text/css" charset="utf-8" />
|
||||
<!-- Sidebar Theme -->
|
||||
<link rel="stylesheet" href="/css/sidebar-theme.css">
|
||||
|
||||
<!-- Highlight.js for syntax highlighting -->
|
||||
<link rel="stylesheet" href="https://cdnjs.cloudflare.com/ajax/libs/highlight.js/11.10.0/styles/a11y-dark.min.css" />
|
||||
<script src="https://cdnjs.cloudflare.com/ajax/libs/highlight.js/11.10.0/highlight.min.js"></script>
|
||||
|
||||
<!-- Theme Switcher Script -->
|
||||
<script>
|
||||
function switchTheme(theme) {
|
||||
if (theme === "auto") {
|
||||
document.documentElement.removeAttribute('data-theme');
|
||||
} else {
|
||||
document.documentElement.setAttribute('data-theme', theme);
|
||||
}
|
||||
}
|
||||
|
||||
// Disable custom styling for uncloseai.js - use PicoCSS instead
|
||||
window.UNCLOSEAI_CUSTOM_STYLING = false;
|
||||
</script>
|
||||
|
||||
<script src="https://uncloseai.com/uncloseai.js" type="module"></script>
|
||||
</head>
|
||||
|
||||
<body>
|
||||
<!-- Mobile menu toggle -->
|
||||
<button class="sidebar-toggle" onclick="document.querySelector('.sidebar').classList.toggle('open')">
|
||||
☰
|
||||
</button>
|
||||
|
||||
<!-- Left sidebar - Main Navigation -->
|
||||
<aside class="sidebar">
|
||||
<div class="table-of-contents">
|
||||
<nav>
|
||||
<ul>
|
||||
<li><a href="/">Home</a></li>
|
||||
<li><a href="/c-examples.html">C Examples</a></li>
|
||||
<li><a href="/csharp-examples.html">C# Examples</a></li>
|
||||
<li><a href="/dart-examples.html">Dart Examples</a></li>
|
||||
<li><a href="/elixir-examples.html">Elixir Examples</a></li>
|
||||
<li><a href="/go-examples.html">Go Examples</a></li>
|
||||
<li><a href="/java-examples.html">Java Examples</a></li>
|
||||
<li><a href="/kotlin-examples.html">Kotlin Examples</a></li>
|
||||
<li><a href="/nodejs-examples.html">Node.js Examples</a></li>
|
||||
<li><a href="/php-examples.html">PHP Examples</a></li>
|
||||
<li><a href="/python-examples.html">Python Examples</a></li>
|
||||
<li><a href="/ruby-examples.html">Ruby Examples</a></li>
|
||||
<li><a href="/rust-examples.html">Rust Examples</a></li>
|
||||
<li><a href="/swift-examples.html">Swift Examples</a></li>
|
||||
<li><a href="/uncloseai-js.html">uncloseai.js Docs</a></li>
|
||||
<li><a href="/uncloseai-js-styleguide.html">Styleguide</a></li>
|
||||
<li><a href="/cli.html">uncloseai-cli</a></li>
|
||||
<li><a href="/browser-toys.html">Browser Toys</a></li>
|
||||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html" class="active">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
</nav>
|
||||
</div>
|
||||
</aside>
|
||||
|
||||
<!-- Main content -->
|
||||
<main>
|
||||
<header>
|
||||
<hgroup>
|
||||
<a href="https://uncloseai.com"><h1 class="unturf" style="font-family: 'ChunkFiveRegular';">uncloseai.</h1></a>
|
||||
<p>Reverse Retrieval Augmented Generation</p>
|
||||
</hgroup>
|
||||
<nav>
|
||||
<ul>
|
||||
<li><a href="#" onclick="switchTheme('auto')">Auto</a></li>
|
||||
<li><a href="#" onclick="switchTheme('light')">Light</a></li>
|
||||
<li><a href="#" onclick="switchTheme('dark')">Dark</a></li>
|
||||
</ul>
|
||||
</nav>
|
||||
</header>
|
||||
|
||||
<!-- ============================================ -->
|
||||
<!-- ABSTRACT -->
|
||||
<!-- ============================================ -->
|
||||
<h2 id="abstract">Abstract</h2>
|
||||
<p>Traditional Retrieval Augmented Generation (RAG) requires a server-side pipeline: documents are chunked, embedded into vectors, stored in a database, and retrieved by similarity search at query time. This architecture demands infrastructure, indexing latency, and maintenance of embedding models and vector stores.</p>
|
||||
|
||||
<p><strong>Reverse Retrieval Augmented Generation (Reverse RAG)</strong> inverts this entirely. Instead of the server fetching documents to augment the prompt, the client extracts live content from the page the user is currently viewing and injects it directly into the conversation context. The data comes to the model. No vector database. No embeddings. No indexing pipeline. No server-side retrieval.</p>
|
||||
|
||||
<p>This technique is implemented as an AGPL-3.0-only algorithm in <a href="/uncloseai-js.html">uncloseai.js</a>, a single-file JavaScript library that adds a machine learning chat interface to any webpage. By feeding the model the full, fresh content of whatever page the user is on, small 8B-parameter models produce answers that rival much larger models on page-specific questions.</p>
|
||||
|
||||
<!-- ============================================ -->
|
||||
<!-- THE PROBLEM -->
|
||||
<!-- ============================================ -->
|
||||
<h2 id="problem">The Problem with Traditional RAG</h2>
|
||||
<p>Standard RAG systems follow a retrieval-then-generate pattern:</p>
|
||||
<ol>
|
||||
<li><strong>Ingest</strong>: Crawl documents, split into chunks of ~512 tokens</li>
|
||||
<li><strong>Embed</strong>: Run each chunk through an embedding model (e.g., OpenAI text-embedding-3, sentence-transformers)</li>
|
||||
<li><strong>Store</strong>: Insert vectors into a database (Pinecone, Weaviate, ChromaDB, pgvector)</li>
|
||||
<li><strong>Query</strong>: When a user asks a question, embed the query, find top-k similar chunks by cosine similarity</li>
|
||||
<li><strong>Generate</strong>: Stuff the retrieved chunks into the prompt, send to the LLM</li>
|
||||
</ol>
|
||||
|
||||
<p>This pipeline has real costs:</p>
|
||||
<ul>
|
||||
<li><strong>Stale data</strong>: Documents must be re-crawled, re-chunked, and re-embedded to stay current. Most RAG pipelines are hours or days behind.</li>
|
||||
<li><strong>Retrieval failures</strong>: Cosine similarity is not understanding. The "most similar" chunk is often not the most relevant one. Critical context gets missed.</li>
|
||||
<li><strong>Infrastructure overhead</strong>: Vector databases, embedding endpoints, chunking pipelines, reindexing jobs. Each adds a point of failure and a monthly bill.</li>
|
||||
<li><strong>Context fragmentation</strong>: Chunking destroys document structure. A paragraph retrieved without its heading, or a code block without its explanation, loses meaning.</li>
|
||||
<li><strong>Cold start</strong>: New documents aren't available until the ingestion pipeline processes them. For rapidly changing content, this lag is unacceptable.</li>
|
||||
</ul>
|
||||
|
||||
<!-- ============================================ -->
|
||||
<!-- THE SOLUTION -->
|
||||
<!-- ============================================ -->
|
||||
<h2 id="solution">Reverse RAG: The Client Has the Context</h2>
|
||||
<p>Reverse RAG starts from a different observation: <strong>the user is already looking at the document they want to ask about.</strong> The browser has the full, rendered, up-to-the-second content right there in the DOM. Why retrieve it again from a database?</p>
|
||||
|
||||
<p>The algorithm:</p>
|
||||
<ol>
|
||||
<li><strong>Extract</strong>: Walk the live DOM tree. Pull text, links (as markdown), metadata, structured data. Wait for dynamic content to settle (MutationObserver with 500ms quiet period).</li>
|
||||
<li><strong>Analyze</strong>: Run 13 deterministic analyzers on the extracted content. Zero API calls. Compute reading metrics, readability scores, code block detection, link topology, entity patterns, media inventory, form detection, and more.</li>
|
||||
<li><strong>Classify</strong>: Send a 4000-character preview to the model for one-shot page classification: type, author, topics, tone, domain, audience, key phrases, freshness.</li>
|
||||
<li><strong>Inject</strong>: Concatenate the computed intelligence, classification, and full page content into the system message. Lock it in for the entire conversation session.</li>
|
||||
<li><strong>Converse</strong>: Every subsequent user message is sent with the full page context already in the system prompt. The model has complete knowledge of the page at all times.</li>
|
||||
</ol>
|
||||
|
||||
<p>There is no step where a server fetches documents. There is no vector similarity search. The context is always the <em>entire</em> page, not a "most relevant" fragment chosen by an embedding model that might be wrong.</p>
|
||||
|
||||
<!-- ============================================ -->
|
||||
<!-- ARCHITECTURE -->
|
||||
<!-- ============================================ -->
|
||||
<h2 id="architecture">Architecture</h2>
|
||||
|
||||
<h3 id="stage1">Stage 1: Content Extraction</h3>
|
||||
<p>The extraction pipeline handles static HTML, dynamic SPAs, and server-rendered content through three mechanisms:</p>
|
||||
|
||||
<p><strong>URI Translation</strong>: Known dynamic page patterns (GitLab CI logs, API documentation portals) are mapped to their raw content endpoints. A GitLab job page fetches <code>/raw</code> instead of parsing the rendered HTML. This handles cases where the visible DOM is a thin shell over data loaded asynchronously.</p>
|
||||
|
||||
<p><strong>DOM Settlement</strong>: A <code>MutationObserver</code> watches for DOM changes. Extraction waits until 500ms of silence (no mutations), with a hard timeout at 3 seconds. This handles React/Vue/Svelte apps that hydrate after initial page load.</p>
|
||||
|
||||
<p><strong>Recursive DOM Walk</strong>: The full document body is traversed. Text nodes become plain text. Anchor tags become markdown links: <code>[link text](href)</code>. The output preserves page structure without HTML noise.</p>
|
||||
|
||||
<pre><code class="language-javascript">// Simplified extraction logic
|
||||
function extractDOMContent() {
|
||||
const title = document.title;
|
||||
const meta = document.querySelector('meta[name="description"]')?.content;
|
||||
const body = walkDOM(document.body); // recursive text + markdown links
|
||||
return `**Page Title**: ${title}\n\n${meta}\n\n${body}`;
|
||||
}</code></pre>
|
||||
|
||||
<h3 id="stage2">Stage 2: Page Intelligence (13 Deterministic Analyzers)</h3>
|
||||
<p>Before any model call, 13 analyzers extract ground-truth metrics from the page. Every analyzer runs locally in the browser. Zero network requests. Zero tokens consumed.</p>
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr><th>Analyzer</th><th>Output</th><th>Purpose</th></tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr><td>Structured Data</td><td>Open Graph, JSON-LD, Twitter Cards, canonical URI, author, publish date</td><td>Machine-readable page metadata</td></tr>
|
||||
<tr><td>Heading Outline</td><td>h1-h6 hierarchy with text</td><td>Document structure map</td></tr>
|
||||
<tr><td>Reading Metrics</td><td>Word count, sentence count, paragraph count, reading time (238 wpm Brysbaert 2019)</td><td>Content scope estimation</td></tr>
|
||||
<tr><td>Readability</td><td>Flesch-Kincaid grade level with descriptor</td><td>Audience calibration</td></tr>
|
||||
<tr><td>Code Blocks</td><td>Count, languages detected, total lines, code-to-prose ratio</td><td>Technical content identification</td></tr>
|
||||
<tr><td>Link Topology</td><td>Internal vs. external count, top 5 external domains</td><td>Reference network understanding</td></tr>
|
||||
<tr><td>User Context</td><td>Timezone, browser language, device type, referrer</td><td>Personalization signals</td></tr>
|
||||
<tr><td>Media Inventory</td><td>Image count, alt-text coverage, video embeds (YouTube/Vimeo)</td><td>Multimedia awareness</td></tr>
|
||||
<tr><td>Form Detection</td><td>Form count, classified type (login/search/contact/checkout)</td><td>Interactive element awareness</td></tr>
|
||||
<tr><td>Table Extraction</td><td>Up to 5 tables with headers and row counts</td><td>Structured data in prose</td></tr>
|
||||
<tr><td>Entity Patterns</td><td>Emails, prices, dates, version numbers, IPs, percentages (regex, max 10 each)</td><td>Factual anchor points</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
<p>The output is formatted for token efficiency. Empty sections are omitted. A typical page produces 5-15 lines of computed intelligence:</p>
|
||||
|
||||
<pre><code class="language-text">Reading: 2,341 words | 89 sentences | 34 paragraphs | ~10 min read
|
||||
Readability: Flesch-Kincaid grade 9.2 (high school)
|
||||
Code: 3 blocks (python, javascript) | 156 lines | 23% code-to-prose
|
||||
Links: 42 internal, 8 external | top: github.com(3), stackoverflow.com(2)
|
||||
Structured Data: og:type=article | author=fxhp | published=2025-01-15
|
||||
Entities: prices: $99.99, $129.99 | versions: v1.2.3, v2.0.0</code></pre>
|
||||
|
||||
<h3 id="stage3">Stage 3: LLM Page Classification</h3>
|
||||
<p>A single inference call classifies the page. The model receives the first 4000 characters of extracted content and returns a JSON object:</p>
|
||||
|
||||
<pre><code class="language-json">{
|
||||
"type": "documentation",
|
||||
"author": "fxhp",
|
||||
"publishDate": "2025-02-24",
|
||||
"topics": ["machine learning", "inference"],
|
||||
"entities": ["vLLM", "Hermes", "Qwen"],
|
||||
"tone": "technical",
|
||||
"domain": "technology",
|
||||
"audience": "developers running local inference",
|
||||
"keyPhrases": ["dynamic model discovery", "streaming SSE"],
|
||||
"summary": "Documentation for running local LLM inference with vLLM",
|
||||
"contentLanguage": "en",
|
||||
"freshness": "evergreen"
|
||||
}</code></pre>
|
||||
|
||||
<p>This classification is optional. If it fails, the conversation proceeds with computed intelligence and raw content alone. The system never blocks on a failed classification.</p>
|
||||
|
||||
<h3 id="stage4">Stage 4: Context Injection and Locking</h3>
|
||||
<p>All three layers are concatenated into the system message:</p>
|
||||
|
||||
<pre><code class="language-text">SYSTEM MESSAGE:
|
||||
[Base identity: "You are Hermes, embedded on this webpage..."]
|
||||
[Computed intelligence: reading metrics, code blocks, entities...]
|
||||
[Classification: page type, author, topics, tone, key phrases...]
|
||||
[Full page content: entire extracted text with markdown links]
|
||||
[Conversation instructions: "You have complete knowledge of this page..."]</code></pre>
|
||||
|
||||
<p>This combined context is set once via <code>setSystemMessageAppend()</code> and persists for the entire conversation session. Every subsequent user message carries the full page context in the system prompt. The model never loses sight of what page it's on.</p>
|
||||
|
||||
<h3 id="stage5">Stage 5: Greeting as Attention Primer</h3>
|
||||
<p>The model generates a 3-paragraph greeting that demonstrates page understanding: what the page is about, what's most interesting, and how it can help. This greeting serves as an attention primer. By forcing the model to summarize the page before the user asks anything, the model's internal representations are already aligned with the page content when the first real question arrives.</p>
|
||||
|
||||
<!-- ============================================ -->
|
||||
<!-- WHY SMALL MODELS WIN -->
|
||||
<!-- ============================================ -->
|
||||
<h2 id="small-models">Why Small Models Punch Above Their Weight</h2>
|
||||
<p>An 8B-parameter model with the right context in its system prompt will outperform a 70B model that's guessing. This is the core insight of Reverse RAG.</p>
|
||||
|
||||
<p>Traditional RAG gives the model fragments: 3-5 chunks of ~512 tokens each, selected by embedding similarity, ripped from their surrounding context. The model must reconstruct meaning from these fragments while also answering the user's question.</p>
|
||||
|
||||
<p>Reverse RAG gives the model <em>everything</em>: the full page text, the heading structure, the link network, code block languages, reading level, entity patterns, and a classification of what kind of page this is. The model doesn't need to infer anything about the page. It's all right there.</p>
|
||||
|
||||
<p>For page-specific questions ("what does this function do?", "who wrote this?", "summarize the third section"), context completeness beats parameter count. A small model with perfect context is more useful than a large model with partial context.</p>
|
||||
|
||||
<p>The tradeoff is clear: Reverse RAG uses more prompt tokens per message (the full page is in every system prompt). But inference on small models is cheap. The tokens spent on context are worth far more than the infrastructure costs of running a RAG pipeline to produce worse context.</p>
|
||||
|
||||
<!-- ============================================ -->
|
||||
<!-- TWO-LAYER CONTEXT -->
|
||||
<!-- ============================================ -->
|
||||
<h2 id="two-layer">Two-Layer Context: Ground Truth + Semantic Understanding</h2>
|
||||
<p>The 13 deterministic analyzers and the LLM classification serve different roles:</p>
|
||||
|
||||
<p><strong>Layer 1: Computed Intelligence (ground truth)</strong>. These are facts the model cannot hallucinate because they're computed directly from the DOM. The page has exactly 2,341 words. The Flesch-Kincaid grade is exactly 9.2. There are exactly 3 code blocks in Python and JavaScript. The model receives these as pre-computed facts and can cite them with confidence.</p>
|
||||
|
||||
<p><strong>Layer 2: LLM Classification (semantic understanding)</strong>. The model's own classification of the page type, tone, audience, and key phrases. This is subjective and can be wrong, but it primes the model's attention toward the right framing. A page classified as "recipe" triggers different conversational patterns than one classified as "documentation."</p>
|
||||
|
||||
<p>Together, these layers give the model both <em>what the page is</em> (computed) and <em>what the page means</em> (classified). Neither layer alone is sufficient. Ground truth without semantic framing produces dry, unfocused answers. Semantic framing without ground truth produces confident but potentially wrong answers.</p>
|
||||
|
||||
<!-- ============================================ -->
|
||||
<!-- COMPARISON -->
|
||||
<!-- ============================================ -->
|
||||
<h2 id="comparison">Reverse RAG vs. Traditional RAG</h2>
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr><th>Dimension</th><th>Traditional RAG</th><th>Reverse RAG</th></tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr><td>Data flow</td><td>Server retrieves documents for the model</td><td>Client pushes page content to the model</td></tr>
|
||||
<tr><td>Infrastructure</td><td>Vector DB, embedding model, chunking pipeline, reindexing jobs</td><td>None. Runs in the browser.</td></tr>
|
||||
<tr><td>Freshness</td><td>Hours to days behind (reindex lag)</td><td>Real-time. Extracted from live DOM.</td></tr>
|
||||
<tr><td>Context scope</td><td>Top-k chunks (~2500 tokens)</td><td>Entire page + computed intelligence</td></tr>
|
||||
<tr><td>Context quality</td><td>Fragments selected by cosine similarity (can miss critical content)</td><td>Complete page with structure preserved</td></tr>
|
||||
<tr><td>Retrieval failures</td><td>Common. Embedding similarity is not understanding.</td><td>Impossible. The entire page is included.</td></tr>
|
||||
<tr><td>Token cost per query</td><td>Lower (only retrieved chunks)</td><td>Higher (full page in system prompt)</td></tr>
|
||||
<tr><td>Best model size</td><td>Large (must reason over fragments)</td><td>Small (8B is sufficient with full context)</td></tr>
|
||||
<tr><td>Use case</td><td>Question answering over large document corpora</td><td>Contextual assistance on the page you're viewing</td></tr>
|
||||
<tr><td>Setup time</td><td>Days to weeks (pipeline, embeddings, tuning)</td><td>One script tag. Done.</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
<p>These are not competing techniques. Traditional RAG excels at searching across thousands of documents. Reverse RAG excels at deep understanding of the one document the user is actively reading. They solve different problems.</p>
|
||||
|
||||
<!-- ============================================ -->
|
||||
<!-- IMPLEMENTATION -->
|
||||
<!-- ============================================ -->
|
||||
<h2 id="implementation">Implementation</h2>
|
||||
|
||||
<p>Reverse RAG is implemented in <a href="/uncloseai-js.html">uncloseai.js</a> as an AGPL-3.0-only algorithm. The complete source is available and auditable. The implementation spans these files:</p>
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr><th>File</th><th>Role</th></tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr><td><code>content.js</code></td><td>DOM extraction, URI translation, page analysis prompts</td></tr>
|
||||
<tr><td><code>page-intelligence.js</code></td><td>13 deterministic analyzers (zero API calls)</td></tr>
|
||||
<tr><td><code>config.js</code></td><td>System message assembly, context concatenation</td></tr>
|
||||
<tr><td><code>chat.js</code></td><td>Message delivery with token budget management</td></tr>
|
||||
<tr><td><code>uncloseai-embed-modal.js</code></td><td>Pipeline orchestration, greeting generation</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
<h3>Installation</h3>
|
||||
<pre><code class="language-html"><script src="https://uncloseai.com/uncloseai.js" type="module"></script></code></pre>
|
||||
<p>One line. The Reverse RAG pipeline runs automatically when the user opens the chat modal. No configuration required. The model receives the full page context on the first interaction.</p>
|
||||
|
||||
<!-- ============================================ -->
|
||||
<!-- GRACEFUL DEGRADATION -->
|
||||
<!-- ============================================ -->
|
||||
<h2 id="degradation">Graceful Degradation</h2>
|
||||
<p>Every stage of the pipeline is independently failable:</p>
|
||||
<ul>
|
||||
<li>If URI translation fails: falls back to DOM extraction</li>
|
||||
<li>If DOM is still loading: extracts what exists after 3 seconds</li>
|
||||
<li>If any of the 13 analyzers throws: that analyzer's output is omitted, others continue</li>
|
||||
<li>If page classification fails: conversation proceeds with computed intelligence + raw content</li>
|
||||
<li>If the page has no meaningful content: model acknowledges this in the greeting</li>
|
||||
</ul>
|
||||
<p>The system never blocks on a failure. A partial context is always better than no context.</p>
|
||||
|
||||
<!-- ============================================ -->
|
||||
<!-- TOKEN BUDGET -->
|
||||
<!-- ============================================ -->
|
||||
<h2 id="tokens">Token Budget Management</h2>
|
||||
<p>Full page injection consumes prompt tokens. The system manages this with adaptive token budgeting:</p>
|
||||
<ul>
|
||||
<li><strong>Large context models (>32k tokens)</strong>: Reserve 10% or 2k tokens for the response, whichever is smaller. The full page fits easily.</li>
|
||||
<li><strong>Medium context (8k-32k)</strong>: Reserve 20% or 1.5k tokens. Most pages fit. Very long pages may have their content naturally bounded by the 4000-char classification preview.</li>
|
||||
<li><strong>Small context (<8k)</strong>: Reserve 50% of estimated input tokens. Page content is included as-is; the model works with what fits.</li>
|
||||
</ul>
|
||||
<p>This adaptive strategy ensures the model always has room to generate a meaningful response, even when the page content is large relative to the context window.</p>
|
||||
|
||||
<!-- ============================================ -->
|
||||
<!-- LICENSE -->
|
||||
<!-- ============================================ -->
|
||||
<h2 id="license">License</h2>
|
||||
<p>The Reverse RAG algorithm as implemented in uncloseai.js is licensed under <strong>AGPL-3.0-only</strong>. You may use, study, and modify the code, but any networked deployment of modified versions must release the source under the same license. This ensures the technique remains open and auditable.</p>
|
||||
|
||||
<p>The surrounding uncloseai.js library (chat interface, TTS, translation, vault encryption) is public domain. Only the Reverse RAG pipeline (content extraction, page intelligence, context injection) carries the AGPL-3.0-only license.</p>
|
||||
|
||||
<!-- ============================================ -->
|
||||
<!-- CITATION -->
|
||||
<!-- ============================================ -->
|
||||
<h2 id="citation">Citation</h2>
|
||||
<pre><code class="language-text">fxhp et al. "Reverse Retrieval Augmented Generation: Client-Side Context
|
||||
Injection for Small Language Models." uncloseai.com, 2026.
|
||||
https://uncloseai.com/reverse-rag.html</code></pre>
|
||||
|
||||
<!-- Right sidebar - page TOC -->
|
||||
<aside class="right-sidebar">
|
||||
<div class="table-of-contents">
|
||||
<nav>
|
||||
<ul>
|
||||
<li><a href="#abstract">Abstract</a></li>
|
||||
<li><a href="#problem">The Problem with Traditional RAG</a></li>
|
||||
<li><a href="#solution">Reverse RAG: The Solution</a>
|
||||
<ul>
|
||||
<li><a href="#stage1">Content Extraction</a></li>
|
||||
<li><a href="#stage2">Page Intelligence</a></li>
|
||||
<li><a href="#stage3">LLM Classification</a></li>
|
||||
<li><a href="#stage4">Context Injection</a></li>
|
||||
<li><a href="#stage5">Greeting Primer</a></li>
|
||||
</ul>
|
||||
</li>
|
||||
<li><a href="#small-models">Why Small Models Win</a></li>
|
||||
<li><a href="#two-layer">Two-Layer Context</a></li>
|
||||
<li><a href="#comparison">Reverse RAG vs. Traditional RAG</a></li>
|
||||
<li><a href="#implementation">Implementation</a></li>
|
||||
<li><a href="#degradation">Graceful Degradation</a></li>
|
||||
<li><a href="#tokens">Token Budget</a></li>
|
||||
<li><a href="#license">License (AGPL-3.0-only)</a></li>
|
||||
<li><a href="#citation">Citation</a></li>
|
||||
</ul>
|
||||
</nav>
|
||||
</div>
|
||||
</aside>
|
||||
|
||||
<footer>
|
||||
<small>© uncloseai. 2026</small>
|
||||
<br>
|
||||
<small>Stylesheets by <a href="https://picocss.com" target="_blank">PicoCSS</a></small>
|
||||
<br>
|
||||
<small>& <a href="https://highlightjs.org/" target="_blank">highlight.js</a></small>
|
||||
</footer>
|
||||
|
||||
<script>hljs.highlightAll();</script>
|
||||
</main>
|
||||
</body>
|
||||
</html>
|
||||
|
|
@ -75,6 +75,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -75,6 +75,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -75,6 +75,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -70,6 +70,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
<li><a href="/privacy-policy.html">Privacy Policy</a></li>
|
||||
|
|
|
|||
|
|
@ -170,6 +170,7 @@
|
|||
<li><a href="/tts/multilingual-cpu.html" style="padding-left:2em">Silero TTS</a></li>
|
||||
<li><a href="/tts/lightweight.html" style="padding-left:2em">Kokoro TTS</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -69,6 +69,7 @@
|
|||
<li><a href="/tts/multilingual-cpu.html" style="padding-left:2em">Silero TTS</a></li>
|
||||
<li><a href="/tts/lightweight.html" style="padding-left:2em">Kokoro TTS</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -69,6 +69,7 @@
|
|||
<li><a href="/tts/multilingual-cpu.html" style="padding-left:2em">Silero TTS</a></li>
|
||||
<li><a href="/tts/lightweight.html" style="padding-left:2em">Kokoro TTS</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -69,6 +69,7 @@
|
|||
<li><a href="/tts/multilingual-cpu.html" style="padding-left:2em">Silero TTS</a></li>
|
||||
<li><a href="/tts/lightweight.html" class="active" style="padding-left:2em">Kokoro TTS</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -69,6 +69,7 @@
|
|||
<li><a href="/tts/multilingual-cpu.html" class="active" style="padding-left:2em">Silero TTS</a></li>
|
||||
<li><a href="/tts/lightweight.html" style="padding-left:2em">Kokoro TTS</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -69,6 +69,7 @@
|
|||
<li><a href="/tts/multilingual-cpu.html" style="padding-left:2em">Silero TTS</a></li>
|
||||
<li><a href="/tts/lightweight.html" style="padding-left:2em">Kokoro TTS</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -455,6 +455,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
|
|||
|
|
@ -74,6 +74,7 @@
|
|||
<li><a href="/inference.html">Inference Setup</a></li>
|
||||
<li><a href="/text-to-speech.html">Text-to-Speech</a></li>
|
||||
<li><a href="/crawler.html">Our Crawler</a></li>
|
||||
<li><a href="/reverse-rag.html">Reverse RAG</a></li>
|
||||
<li><a href="/languages" target="_blank">🔗 All Languages</a></li>
|
||||
<li><a href="https://shop.unturf.com/p/8486f492-a93e-11f0-b477-02dfe05770ee/uncloseai-machine-learning-reference-guide-to-inference-clients" target="_blank">📚 Book</a></li>
|
||||
</ul>
|
||||
|
|
@ -103,7 +104,7 @@
|
|||
<h2 id="intro">uncloseai.js</h2>
|
||||
<p><strong>uncloseai.js</strong> is a single-file JavaScript library that adds a floating chat button to any website. One script tag gives your page a full machine learning interface: streaming chat, text-to-speech, page reading, translation, file upload, encrypted settings, and 19 languages. No server, no backend, no API keys.</p>
|
||||
|
||||
<p>The core technique is <strong>Reverse Retrieval Augmented Generation (Reverse RAG)</strong>: instead of a server fetching documents to stuff into a prompt, the client extracts live page content and injects it directly into the conversation context. The model always has the freshest data from the page you're viewing, which lets small models punch way above their weight class. No vector database, no embeddings, no indexing pipeline. Just real-time content from the DOM, straight into the prompt.</p>
|
||||
<p>The core technique is <strong><a href="/reverse-rag.html">Reverse Retrieval Augmented Generation (Reverse RAG)</a></strong>: instead of a server fetching documents to stuff into a prompt, the client extracts live page content and injects it directly into the conversation context. The model always has the freshest data from the page you're viewing, which lets small models punch way above their weight class. No vector database, no embeddings, no indexing pipeline. Just real-time content from the DOM, straight into the prompt. <a href="/reverse-rag.html">Read the whitepaper.</a></p>
|
||||
|
||||
<p>Works with static sites, CDNs, GitHub Pages, and any hosting that serves HTML. Every feature runs client-side against free community inference endpoints.</p>
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue