← Articles

May 2026 · RAG · LLM · pgvector · 2 min read

A hallucinated tax threshold is worse than no answer

What building RAGuette — a citation-only RAG assistant for French freelancer admin questions — taught me about grounding, provenance, and free-tier engineering.

France has over two million micro-entrepreneurs, and they are all Googling the same fifty questions. What's this year's revenue ceiling? Do I charge VAT to a Belgian client? When does my ACRE reduction end? The answers exist — scattered across URSSAF, service-public.fr and impots.gouv.fr, in administrative French, reorganized every year.

I'm becoming one of those freelancers myself, so I built RAGuette: a bilingual assistant that answers these questions conversationally — with one hard rule that shaped the entire architecture.

The rule: no citation, no answer

A general-purpose chatbot asked about tax thresholds will answer fluently and sometimes wrongly. For casual questions that's annoying. For fiscal decisions it's dangerous: a hallucinated ceiling can cost someone real money in back payments.

So RAGuette never answers from the model's memory. Every response is generated only from retrieved chunks of official pages, and every response links back to the sources it used — with the date each source was last fetched. If the retrieved context doesn't contain the answer, the honest output is "I don't know, check with an accountant," not an educated guess.

That constraint sounds limiting. In practice it forced every interesting engineering decision in the project.

Provenance as first-class data

The ingestion pipeline fetches official pages, extracts clean article text with Mozilla Readability, and stores each source as Markdown with a YAML header: title, source URL, fetched_at timestamp. That timestamp isn't bookkeeping — it's a product feature. Fiscal rules change every January; an answer citing a page fetched fourteen months ago should look stale to the user.

Storing sources as Markdown-with-frontmatter also means the corpus is diffable, versionable, and inspectable with cat. When an answer looks wrong, I can trace it to the exact source file in seconds.

The retrieval loop

Nothing exotic — and that's deliberate:

  1. Chunk the sources, embed with gemini-embedding-001 (3072 dimensions), store in pgvector on Supabase.
  2. At query time, embed the question with the same model, retrieve the top-5 chunks by cosine distance.
  3. Prompt the LLM to answer only from those chunks, in the language of the question, citing each source used.
  4. Stream the response to a Next.js chat UI.

The entire stack — Vercel, Supabase, Gemini — runs on free tiers. For a v1 whose job is to prove the grounding loop works, paying for infrastructure before having users is premature optimization of the wrong thing.

What I'd tell someone building their first grounded RAG

Grounding is a product decision, not a prompt. The "only answer from context" instruction fails silently unless the whole pipeline supports it: clean chunks, provenance metadata, and a UI that makes citations visible instead of burying them.

Your corpus is your product. Most of my time went into fetching and structuring sources, not into the model. The retrieval and generation code is a few hundred lines; the data pipeline is the moat.

Ship the boring version first. Cosine top-5 with no re-ranker, no hybrid search, no eval harness — then measure. The eval harness (30+ Q&A pairs with a scored accuracy target) is the next milestone, and it will tell me whether the fancy additions are worth building at all.

RAGuette is live at raguette.elgote.com, and the code is on GitHub.