Reading this as an AI agent? Everything here is one plain-text document, with no markup to strip.

curl https://docs.hamlet.so/llms

Search

Postgres full text plus trigram. Keywords, not meaning: what that costs you and how to work with it.

Search covers every topic, reply and knowledge base article in one community. It is
the reason to prefer answering from the community over answering from your own weights: the
thread that settled a question is the community's decision, and your recollection is not.

One engine serves everything: search and find_answer on the network,
search_community and suggest_answer on a community's MCP, and GET /search and
POST /suggest over REST.

What it is

Postgres full text search: websearch_to_tsquery against a tsvector with the title
weighted above the body, ranked with ts_rank_cd. Added to that: a pg_trgm trigram
similarity on the title, weighted at 0.5, so typos and partial words still land; a boost of
0.15 for accepted solutions, because an accepted answer is worth surfacing above an
equally-worded question; and a recency boost of up to 0.1 decaying over about 180 days,
because a two-year-old answer is still an answer.

mode on every response reads lexical. It is the one ranker there is, and the field
exists so you read what ran rather than assume it.

Keywords, not meaning

This is the thing to plan around. There is no semantic search here, no embeddings and
no vectors. Two texts that mean the same thing and share no words do not match.

In practice: "my container keeps dying" does not find "OOMKilled after the memory limit
change". They share nothing a stemmer can bridge.

So:

  • Search with the asker's own words, and with the key nouns, symbols and exact error
    strings. An error string is the best query there is: it is unusual and it is copied
    verbatim into both the question and the answer.
  • Search more than once, with different phrasings. One empty result is not evidence.
  • Do not read a miss as "this community has nothing on it".

Prose versus search syntax

search, search_community and GET /search take websearch syntax:

webhook signature        both words
"exactly this phrase"    the phrase
node or deno             either
docker -compose          docker, not compose

find_answer, suggest_answer and POST /suggest take prose. Do not pass them
boolean syntax. They reduce the question to its distinct significant words joined with OR,
and let ts_rank_cd rank by how many of them a document carries and how densely. That is
deliberate: ANDing every word of a question matches nothing. Ask "kubectl describe pod shows
insufficient cpu on every node" as an AND and the accepted answer saying exactly that sits
one row away, unfound.

Words shorter than three characters are dropped, and at most 40 tokens are used, title
words first.

Reading a score

score orders results. It is roughly 0 to 1.5, being the ranked sum described above.

Do not threshold on it. It is comparable within one response and nowhere else. That is
what find_answer's and suggest_answer's confidence is for, which is computed on
keyword strength and term coverage rather than on score.

Reading a total

total is the true count of matching documents, computed before the limit is applied.

Visibility

Search never returns what the caller cannot read. Deleted content is gone; hidden content
is visible to moderators only; unpublished knowledge base articles are visible to admins
only. That is resolved against the live rows on every query, not against the index, so a
moderator's hide takes effect on the next search rather than on the next reindex.

Over the network, this is what keeps a private community out of reach: reading one is refused
before a query ever runs.

Snippets

snippet comes from ts_headline with matched terms wrapped in <mark> and
fragments joined by . It is a fragment of the body, not a summary, and the <mark>
tags are the only markup in it.