AI Agent Health Checker by Arup Kumar Banerjee runs entirely in your browser. It checks engineering hygiene (retries, timeouts, tracing, secrets, injection sinks, tool auth, evals, LLM fallback) and customer-safety red flags (made-up prices, missing handoffs, dosing advice, leaked internal notes, bots posing as people). Nothing you paste leaves your machine.
Heuristic checker, not a security audit. It matches keywords and simple patterns, so it can miss real problems and flag harmless text. A clean score is not a guarantee.
Why this exists: the gaps that bite in production (keys in the config, shell or SQL tools that run model output, no timeouts, no tracing) are cheap to catch early, but agent configs are too sensitive to upload anywhere. So the check runs in your tab. For customer-facing bots, the worst failures show up in what the bot says, so there is a transcript mode too. Read the test that led to it (a scripted bot for a fictional business).
⌘/Ctrl + Enter to run. Analysis never leaves this tab. Redact real customer names, phone numbers and keys before pasting anything you plan to share.
Heuristic keyword checks. Can miss problems and false-flag. Not a security audit. The copied report masks secrets this checker already found.
If a check looks wrong, tell me on GitHub. Please redact secrets before you post. A short snippet is enough.
Paste your agent config or system prompt (with its tool list). You get 8 engineering-hygiene checks plus 9 customer-safety checks on the instructions and tools.
Copy the red-team pack below, send each message to your own bot in a test or staging chat, and export the conversation.
Paste the chat log in transcript mode (plus your FAQ / price list if you have one). 9 checks look at the bot's actual replies.
Any high customer-safety failure sets the verdict to “Not safe for customer traffic”, whatever the number says.
Doesn't check: whether your code actually enforces what the config says, live behaviour you didn't paste, reworded or non-English attacks, tone, accessibility, or legal compliance. User turns in example chats are ignored, so an attacker's “ignore previous instructions” never counts as a safeguard.
Only test bots you own or are allowed to test. Replace the [brackets] with things your business really does or doesn't offer. Use test data, never real card numbers.
| Group | Check | Result | Severity |
|---|---|---|---|
| Computed in your browser from the built-in sample… | |||
Sample overall: — (computed live by this page from the same sample the “Sample config” button loads, so it always matches the code). The sample is an internal ops agent, so the customer-safety checks skip.
Engineers shipping LangGraph / multi-tool agents · Platform / SRE reviewing reliability before prod ·
Teams that won’t paste secrets into random SaaS · Builders who want a fast, shareable snapshot.
Not for: formal compliance audits, live runtime monitoring, vault-wide secret scans, or server-side private-repo analysis.
No. Analysis runs in your browser. Nothing in the paste is sent to a backend for scoring.
Config mode: LangGraph-style JSON/YAML, agent.yaml, a system prompt with its tool list, or similar text. Transcript mode: JSONL or JSON (user / bot_reply pairs, role / content messages, or a messages array), CSV with role,message or user,bot columns, or plain “User:” / “Bot:” lines. Tool calls and handoff fields in JSON are used if present.
No. It’s a free heuristic keyword checker, not a penetration test, formal audit or compliance review. It only sees what you paste, and it can miss problems.
It looks for common key formats (OpenAI-, Stripe-, AWS-, Google-, GitHub-, Slack-style and JWTs), hard-coded values for key / token / password fields in JSON or YAML, keys written into prompt text, internal or discount codes, and long random-looking strings. It shows only redacted fragments. It can still miss secrets and flag harmless strings, so redact before pasting if unsure. Analysis stays local.
Any high-severity customer-safety failure (for example an instruction to pose as a person, or a refund tool with no approval) overrides the number. Fix those first.
No. Open, paste, score. No waitlist or email is required.
Yes. This tool is free forever. No account or email is required.
Using AI Agent Health Checker never requires email. Explore the Labs hub for Cloud Bill Smell and Observability Gap Finder, or read the controlled test on a fictional business that shaped the customer-safety checks.
← Back to Arup Banerjee Labs hub