Every "AI agents vs chatbots" article on page one runs the same table:
chatbots pattern-match, agents reason; chatbots have no memory, agents do;
chatbots follow a script, agents plan. It's not wrong, exactly. It's also not
useful, because it describes the technology instead of your decision.
The question that actually matters when you're buying or building one of
these things is narrower: does this specific workflow need a system that
takes actions, or one that gives good answers? Get that wrong in either
direction and you pay for it — over-build an agent for a task a chatbot
handled fine, or under-build a chatbot for a task that needed real tool
access, and you're rebuilding in six months either way.
Four questions, not a feature table
Ask these about the workflow you're actually trying to solve, not about AI in
the abstract.
1. Does it need to take actions across systems, or just produce an
answer?
A support assistant that tells a customer their order status is answering. A
system that cancels the order, issues the refund and updates the CRM is
acting. If every output of your system still needs a human to go do
something in another tool, you don't have an agent problem yet — you have a
retrieval and drafting problem, and a chatbot solves it for a fraction of the
engineering cost.
2. Does a wrong action cost more than a wrong answer?
A wrong answer is embarrassing and correctable — the user asks again, or a
human notices in review. A wrong action — a refund issued twice, an email
sent to the wrong list, an entitlement changed incorrectly — has to be
undone, and undoing costs more than doing. If the downside of a mistake is
irreversible or expensive, you need the guardrails and confirmation layer
that come with agent architecture, not a chatbot with a tool bolted on.
3. Does it need persistent state across sessions?
A chatbot conversation is stateless in the way that matters here: each
session starts fresh, or carries only chat history as context. An agent
managing a multi-day onboarding workflow, a claims process, or a research
task needs to remember what step it's on, what's been done, and what's still
required — as structured state in your database, not as something inferred
from scrollback. If your workflow spans more than one session or more than a
few minutes of wall-clock time, that's an agent requirement, and it changes
the architecture from day one.
4. Is there already a human in the loop who is the actual bottleneck?
If a person currently reviews every decision before it takes effect, ask
whether the AI's job is to make that person faster (a chatbot that drafts,
summarises, and surfaces context) or to remove them from the loop entirely
(an agent that acts and is audited after the fact). These are different
products with different risk profiles. Removing the human is the harder,
more expensive build — don't reach for it unless the review step is
genuinely the constraint on throughput.
If your answers land mostly on "produces an answer," "wrong answer is
recoverable," "one session," and "human review is fine" — you want a
chatbot, and a good one is not a lesser product. If two or more land on the
other side, you're buying or building an agent, and the conversation should
move to autonomy bounds and rollback paths before it moves to which model to
use.
Agent-washing, and how to catch it in a demo
"Agentic AI" is the funded line item this year, and a predictable number of
vendors relabelled their existing retrieval chatbot rather than build
anything new. The product in the demo still just answers questions — it has
simply started calling itself an agent in the deck.
Three questions expose this in a live demo, and a vendor who has actually
built agent infrastructure will answer them without flinching:
- "Show me it taking an action in a connected system, not describing one."
Ask it to actually create the ticket, update the record, or send the
message — live, in front of you. A chatbot with an agent label will
describe what it would do, or demo a single canned tool call that was
clearly scripted for the pitch.
- "What happens when it's wrong three steps into a five-step task?"
A real agent architecture has an answer involving state tracking and
rollback. A relabelled chatbot has an answer involving "the user can
correct it in the next message" — which tells you there's no multi-step
execution happening at all.
- "Walk me through the audit log of an action it took, not the chat
transcript." Agents that act need an audit trail separate from the
conversation — who did what, when, and what it changed. If the only record
is the chat log, there's no action layer underneath.
If a vendor can't clear these, you're not evaluating an agent. You're
evaluating a chatbot with a different name, and you should price it — and
scope it — as one.
Deciding to build one is step one
This framework answers whether you should build an agent at all. It does not
answer how to build one that survives production — that's a separate and
harder problem, covered in
why AI agents fail in production,
which is about the architecture choices — autonomy bounds, state
management, evaluation — once you've already decided the workflow needs one.
If you land on "agent" after running these four questions, that's where our
AI automation work starts: scoping the one workflow, the
decision points, and the guardrails before a single line of orchestration
code gets written.