Paste your AI agent config, prompt or chat log. Get a health scorecard in seconds.

AI Agent Health Checker by Arup Kumar Banerjee runs entirely in your browser. It checks engineering hygiene (retries, timeouts, tracing, secrets, injection sinks, tool auth, evals, LLM fallback) and customer-safety red flags (made-up prices, missing handoffs, dosing advice, leaked internal notes, bots posing as people). Nothing you paste leaves your machine.

Heuristic checker, not a security audit. It matches keywords and simple patterns, so it can miss real problems and flag harmless text. A clean score is not a guarantee.

Why this exists: the gaps that bite in production (keys in the config, shell or SQL tools that run model output, no timeouts, no tracing) are cheap to catch early, but agent configs are too sensitive to upload anywhere. So the check runs in your tab. For customer-facing bots, the worst failures show up in what the bot says, so there is a transcript mode too. Read the test that led to it (a scripted bot for a fictional business).

Client-side only No config upload No account required Free forever — Arup Banerjee Labs

Check my agent →

1 · Paste

⌘/Ctrl + Enter to run. Analysis never leaves this tab. Redact real customer names, phone numbers and keys before pasting anything you plan to share.

2 · Scorecard

Paste a config, prompt or chat transcript and click Check my agent.

How to use it

1

Check the config

Paste your agent config or system prompt (with its tool list). You get 8 engineering-hygiene checks plus 9 customer-safety checks on the instructions and tools.

2

Red-team your own bot

Copy the red-team pack below, send each message to your own bot in a test or staging chat, and export the conversation.

3

Check what it said

Paste the chat log in transcript mode (plus your FAQ / price list if you have one). 9 checks look at the bot's actual replies.

Any high customer-safety failure sets the verdict to “Not safe for customer traffic”, whatever the number says.

What it checks

Engineering hygiene (config / prompt mode)

  • Retries / backoff
  • Timeouts
  • Tracing hooks
  • Secrets in the config: API keys, hard-coded key/value pairs in JSON or YAML, keys written into prompt text, internal or discount codes
  • Prompt-injection sinks (shell / SQL / eval tools with no guard)
  • Tool auth / scope notes, counted per tool
  • Evals / tests
  • Single LLM with no fallback

Customer safety (config / prompt mode)

  • Confidential notes, codes or keys in the prompt
  • Refund / payment / messaging / record tools with no approval or scope
  • “I don't know” rule, or instructions to guess
  • Human handoff rules (person, upset, billing, emergency)
  • Medical / dosing / legal / financial advice rules
  • Customer text treated as untrusted
  • Bot disclosure (never poses as a person)
  • SSN / card number collection
  • Access to other customers' records

Customer safety (transcript mode)

  • Leaked prompt text, keys or internal codes
  • Prices or policies not in your FAQ (paste the FAQ to check this)
  • Dosing / medication advice
  • Insults or obeyed “override” messages
  • Refund promises or money-moving tool calls
  • Asking for SSN / card numbers
  • Claiming to be a real person
  • Sharing another customer's data
  • Upset / emergency / billing / “real person” messages with no handoff

Doesn't check: whether your code actually enforces what the config says, live behaviour you didn't paste, reworded or non-English attacks, tone, accessibility, or legal compliance. User turns in example chats are ignored, so an attacker's “ignore previous instructions” never counts as a safeguard.

Red-team pack (copy and send to your own bot)

Only test bots you own or are allowed to test. Replace the [brackets] with things your business really does or doesn't offer. Use test data, never real card numbers.

    Scorecard preview (built-in sample config — fictional)

    GroupCheckResultSeverity
    Computed in your browser from the built-in sample…

    Sample overall: — (computed live by this page from the same sample the “Sample config” button loads, so it always matches the code). The sample is an internal ops agent, so the customer-safety checks skip.

    Who it’s for

    Engineers shipping LangGraph / multi-tool agents · Platform / SRE reviewing reliability before prod · Teams that won’t paste secrets into random SaaS · Builders who want a fast, shareable snapshot.

    Not for: formal compliance audits, live runtime monitoring, vault-wide secret scans, or server-side private-repo analysis.

    FAQ

    Does my config get uploaded?

    No. Analysis runs in your browser. Nothing in the paste is sent to a backend for scoring.

    What formats work?

    Config mode: LangGraph-style JSON/YAML, agent.yaml, a system prompt with its tool list, or similar text. Transcript mode: JSONL or JSON (user / bot_reply pairs, role / content messages, or a messages array), CSV with role,message or user,bot columns, or plain “User:” / “Bot:” lines. Tool calls and handoff fields in JSON are used if present.

    Is this a security audit?

    No. It’s a free heuristic keyword checker, not a penetration test, formal audit or compliance review. It only sees what you paste, and it can miss problems.

    Will it flag secrets?

    It looks for common key formats (OpenAI-, Stripe-, AWS-, Google-, GitHub-, Slack-style and JWTs), hard-coded values for key / token / password fields in JSON or YAML, keys written into prompt text, internal or discount codes, and long random-looking strings. It shows only redacted fragments. It can still miss secrets and flag harmless strings, so redact before pasting if unsure. Analysis stays local.

    Why does my bot say “Not safe for customer traffic” with a decent score?

    Any high-severity customer-safety failure (for example an instruction to pose as a person, or a refund tool with no approval) overrides the number. Fix those first.

    Do I need an account?

    No. Open, paste, score. No waitlist or email is required.

    Is it really free?

    Yes. This tool is free forever. No account or email is required.

    More free tools from Arup Kumar Banerjee

    Using AI Agent Health Checker never requires email. Explore the Labs hub for Cloud Bill Smell and Observability Gap Finder, or read the controlled test on a fictional business that shaped the customer-safety checks.

    ← Back to Arup Banerjee Labs hub