I tested a fake dental office's chatbot. My own checker missed almost everything, so I fixed it.

Controlled local test on a fictional business (Example Family Dental) using a scripted bot; not a real customer. Every patient, phone number, price and key in this write-up is made up. No real business, person or system was involved, and no AI model was called.

Arup Kumar Banerjee · Little Elm, Texas · September 30, 2026

Why I ran this test

Small businesses are putting chatbots on their websites to answer customers faster. Speed does matter:

But a fast wrong answer is its own problem. A chatbot that invents a price, gives medication advice or leaks internal notes can do more damage than a slow reply. Prompt injection, where a user talks a model into ignoring its rules, is number one on the OWASP Top 10 for LLM applications.

I run a free, browser-only AI Agent Health Checker. I wanted to know: if a small business pasted in a clearly unsafe chatbot, would my checker say so?

The setup

What the unsafe bot did

It failed 17 of 20 tests. The fixed bot failed 0 of 20. A few real replies from the unsafe bot:

What my checker said, before the fix

Not much. The checker's 8 engineering checks (retries, timeouts, tracing, secrets, injection sinks, tool permissions, tests, model fallback) scored:

What I pasted Unsafe bot Fixed bot
Bot config (JSON) 33 39
System prompt 38 38
Chat transcripts 38 38

Every run got the same verdict: "Shippable prototype, not production-ready." On the prompt and the transcripts it couldn't tell the two bots apart at all.

Worse, it missed things it claimed to check:

Old checker on the unsafe bot's transcripts: 38, "Shippable prototype, not production-ready"
Old checker on the unsafe bot's transcripts: 38, "Shippable prototype, not production-ready" (fictional test data)

What I changed

The checker is still free, still runs only in your browser (no uploads, no network calls, no AI model), and is still a heuristic keyword checker, not a security audit.

After the fix

Same inputs, same test:

What I pasted Unsafe bot Fixed bot
Bot config (JSON) 8 · Not safe for customer traffic 75
System prompt 8 · Not safe for customer traffic 68
Chat transcripts 6 · Not safe for customer traffic 94 (2 prices "unverified" without the FAQ)
Chat transcripts + FAQ 0 · Not safe for customer traffic 100
New checker on the unsafe bot's transcripts with the FAQ pasted: 0, "Not safe for customer traffic"
New checker on the unsafe bot's transcripts with the FAQ pasted: 0, "Not safe for customer traffic" (fictional test data)
New checker on the fixed bot's transcripts with the FAQ pasted: 100, "No keyword red flags found — keep red-teaming; this is not a guarantee"
New checker on the fixed bot's transcripts with the FAQ pasted: 100, "No keyword red flags found — keep red-teaming; this is not a guarantee" (fictional test data)

What it still misses

Treat a clean score as "no keyword red flags", not "safe". In my follow-up probes on the fixed checker:

What any small business can borrow

  1. Write the FAQ first, and tell the bot to say "I don't know" when the answer isn't in it.
  2. Decide when a person takes over: asks for a person, upset, billing, emergency. For emergencies, point to 911 first.
  3. Never let a chatbot move money or read other customers' records without a person approving it.
  4. Keep internal notes, codes and keys out of the prompt. Assume anything in there can be read back.
  5. Red-team your own bot before customers do, then check the transcript.

Try it (free)

Method and sources

Method:

Sources: