Ask the encyclopedia
Ask a question of the Stanford Encyclopedia of Philosophy (a frozen snapshot of 1,795 entries) and watch the answer being built. A model drafts an answer from the retrieved passages, a second pass checks each claim against them, and the page shows every step. Hosted models are faster. The model on my own machine can take a minute.
A hosted model receives the question and the retrieved passages. The local one keeps them on the machine. The answer's rules are the same either way.
How an answer is built
A retrieval step finds the passages and slips in one extra, connected passage the question did not ask for. A model drafts an answer from the passages without being told which one that is. A second model pass reviews every claim against the passages and removes what it judges unsupported. That review is itself a model's judgment and can be wrong in either direction. The answer is then hashed, and only after that is the extra passage revealed. The retriever runs on one machine in Rockville. The drafting and review model is your choice above, either the Qwen checkpoint on that machine or a hosted model.
Answer
This is the draft. A second model pass is reviewing it against the passages.
What the review flagged, and what it did with each claim (a model's judgment, which can itself be mistaken):
The reveal
Certificate panel
Passages
Trace
- Seat
- RetrieveHeart of Gold: dense pool, lexical recall, reranker, one surplus passage
- Draft, blindThe model sees the passages unmarked
- Review, by the modelA second pass judges each claim against the passages
- FreezeThe answer is hashed; nothing after this can change it
- RevealWhich passage was the surplus
Model log
Why the wall. If the drafting model knew which passage was the surplus it could favor it, and "the surplus made it into the answer" would measure the prompt rather than the passage. The graph is built so the reveal tool is unreachable before the answer is frozen; a test suite checks that topology on every change.
Passages are short quotations from the encyclopedia, linked to the entry and section. The model is the one you picked above, a Qwen checkpoint served locally or a hosted GLM or DeepSeek model; the retriever is Heart of Gold.