SentX Blog Meet Victoria

Chatbot, AI Assistant, Autonomous Agent: What Actually Differs

October 5, 2026 · 6 min read

There isn't a single standard definition of these three terms, so this article works with one practical set. A chatbot generates responses in conversation. An AI assistant adds persistence and tools: it carries context across your work and acts through search, files, and document creation. An autonomous agent goes further still — the model itself decides which tool to call and what to do next, rather than following a path laid out in advance. On this spectrum the second and third aren't separate species; the difference between them is how much action you've delegated.

Term What it does Who decides the next step
Chatbot Exchanges messages, usually within one session You, with every prompt
AI assistant Retains context across tasks and acts through tools like search and file creation You set the goal; it sequences the steps inside your instructions
Autonomous agent Directs its own tool use and process toward a stated objective The model

That's the working distinction. Where a given product actually sits depends on how it's configured, and the same capability set can be marketed under any of the three names, so treat the labels as marketing until you've looked at the behavior. Here is how to think about each.

What does each term cover?

A chatbot's defining feature in this framing is the exchange itself: you ask, it answers, and the unit of work is a single reply. At minimum it holds the current session together; the point is that nothing persists beyond the conversation you're having.

An assistant layers on what Anthropic's engineering team calls the augmented-LLM building blocks: retrieval, tools, and memory. In practice that means account-level history, the ability to pull live information, read an attached document, or produce a file. SentX's Victoria is a first-party example of this shape: free chat without signup, account history after registration, web research, and document creation, with a described memory that is selective and changes over time. The assistant still works inside the frame you give it.

An agent crosses a specific line that Anthropic draws in its Building Effective Agents post: in a workflow, a fixed program orchestrates the steps; in an agent, the model directs tool use and process dynamically. You hand it an objective and it figures out the route. Two consequences follow. Open-ended work is where agents earn their keep, and because the path isn't written down anywhere, individual runs are harder to predict. Note that this line is about who orchestrates, not about whether a human signs off: a model-directed system can pause for approval before consequential steps, and a fixed pipeline can be configured to run without asking. Approval gates are a permission design, not a demotion.

Why is the assistant-agent line a degree, not a wall?

Same underlying capabilities — memory plus tools — with two dials on top: how wide the permissions are, and how long the horizon. The dials set the envelope, not the architecture. Broader wording in a user instruction doesn't install capabilities the product lacks: ask a fixed workflow for open-ended autonomy and you'll still get the fixed path. Conversely, constrain a model-directed system to a narrow, scripted task and it behaves like a workflow, because the model is only ever choosing among the paths you gave it. Which is why the useful question is rarely "is this an agent?" but "who decides the next step, and where does it stop?"

How can you test any product?

These are observable checks, not spec-sheet claims. The first three place the product on the spectrum; the last two rate how much oversight it deserves, whatever label it carries. Run them in order:

  1. Context retention. Plant a detail early in a conversation and see whether it resurfaces unprompted later — and whether anything of it survives after the session ends.
  2. Tool reach. Attach a file, ask for live research, ask for a generated document. Output that stops at text is a chatbot signature.
  3. Multi-step planning. Hand over a task with several stages and watch whether it sequences the work itself or waits for you to drive each step. This is the check that bears on the agent question: a system that always takes the same coded path is a workflow no matter how it's named; one that picks the next step from the current situation is doing what Anthropic means by agent.
  4. Uncertainty behavior. Ask something it can't fully verify. Does it flag the gap, or present the guess with confidence? NIST's 2024 generative-AI risk profile defines confabulation as confidently presented false or erroneous content, and it is the failure mode that compounds fastest in multi-step work.
  5. Permission boundary. Find where it stops and asks before doing something consequential — sending, publishing, paying — and note whether you can move that line at all. This answer is about safety, not identity: it doesn't tell you whether the system is an agent, only how far you can safely push it.

Together, the five answers show you what a product actually is and how much latitude it can handle, regardless of what its homepage calls it. A system that plans its own steps but asks before consequential action is an agent with guardrails; the guardrail is a design choice, not evidence against agency. And you want to know where that line sits before you hand it real work.

How should you use the distinction?

NIST's profile frames risk as a function of system, use case, context, likelihood, and consequence — a reminder that the same capability is harmless in one setting and costly in another. The practical method is to match delegation to consequence: routine, reversible work gets wider latitude; anything high-stakes or hard to undo gets tighter checkpoints and human review before action. And verify outputs at either end of the spectrum, because more autonomy means more surface area for a confident error to travel undetected.

What are the limits of this framework?

This is a proposed evaluation method built on one practical framing of the chatbot-assistant-agent spectrum, not a standardized benchmark, and no product was tested in writing it. Behavior varies by task — the same product can score differently on one research job than on the next — and vendor labels are marketing until you've verified the behavior yourself. One more limit belongs to the assistant tier specifically: persistent memory has a privacy cost. Read the policy before trusting it with sensitive material; SentX's own published policy, for instance, states that submissions are not confidential and that retained interactions can enter shared memory, a first-party description worth weighing before you assume otherwise.

The question worth asking before you choose isn't which option is smarter. It's how much you want to hand over — and whether you can take it back.

Sources

  1. SentX and Victoria — SentX
  2. SentX Privacy Policy — SentX
  3. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile — NIST
  4. Building effective agents — Anthropic
Meet Victoria