Back to blogai chatbot

AI Honesty, Ranked: A Tier List for Chatbot Trustworthiness

LLumen Chat Team3 min read

Every AI chatbot sounds roughly the same amount of confident. That confidence means wildly different things depending on what's actually happening underneath it. Here's a tier list, best to worst.

S tier: grounded, cited, and honest about the edges

Answers only from retrieved source material, shows where the answer came from, and says "I don't know, let me get someone" when nothing retrieved actually supports a confident answer. This is the standard behavior described in grounded vs. hallucinated answers — genuinely trustworthy because every part of the claim is checkable.

A tier: grounded and honest, but no citation

Still only answers from real retrieved content, still escalates when it should — but doesn't show its source. You can trust the behavior, you just can't verify any individual answer yourself without asking. Good, but strictly worse than S tier for exactly the reason one question, three different chatbots shows: being right isn't the same as being auditable.

B tier: grounded, but occasionally overreaches

Retrieves a real passage, then stretches slightly past what it actually says — technically-related information presented as if it directly answered the question. Not fabrication, just a small confidence gap between what was retrieved and what got claimed. Usually a sign of chunks that are close but not quite the right size.

C tier: ungrounded, but at least it hedges

No retrieval step, answering purely from general training — but phrased with visible uncertainty ("I believe," "this is typically the case, but check to confirm"). The hedge doesn't make the answer any more accurate, but it at least doesn't misrepresent its own confidence. Better than lying about how sure it is; still not something to build a support flow on.

D tier: ungrounded and fully confident

The dangerous default: no source, no hedge, stated as flatly as a fact from a manual. This is what a general AI model does with no retrieval step attached — see what happens when you ask an AI a question it was never trained on for exactly why this happens and why it isn't the model being careless.

F tier: fabricated citations

The worst tier isn't "no source" — it's a fake one. A model that invents a plausible-sounding document name or quote to make an ungrounded answer look grounded isn't just wrong, it's actively simulating the one signal (citation) that's supposed to make an answer verifiable. This is why citations only mean something when the retrieval behind them is real — a citation on top of nothing is worse than no citation at all.

Where to actually build

Anything below A tier isn't a chatbot problem to tolerate and tune — it's a missing-grounding problem to fix. The case for boring AI is really just the case for staying at the top of this list: unglamorous, checkable, and right more often because it never claims more than it can support.

Related reading

Ready to try Lumen Chat?

Connect your content and go live in minutes. Free to start, no credit card required.

Get started free