SkillBambooMenu

Articles · October 4, 2026

When to Use AI Agents: A Five-Question Test Before You Build

When to use AI agents and when a regex, SQL query or decision tree does the job cheaper, plus a five-question test to run before you build one

Arnold Explorer

An AI agent that reads customer emails, pulls out what matters and files it in the right place looks great on a laptop. What's harder is knowing when to use AI agents at all. What will it cost per month, why did it misfile three messages last week, and who gets paged when it loops? The answers depend on what an AI agent actually is, where plain code beats it, and what an agent costs once it leaves the demo. If you're weighing one now, run the five-question test near the end first.

What Is an AI Agent? Goal, Tools, Actions

The word “agent” gets stretched to cover almost anything with a language model inside. A useful definition is narrower. An agent receives a goal rather than a question, and it works toward that goal on its own: it picks tools, takes several steps in sequence and decides what to do next based on what it finds.

Three traits separate an AI agent from a chatbot or a single LLM call:

  • Goal, not prompt. A person states the outcome. The agent owns the route, the choice of tools and the response to failures.
  • Tools and real actions. A chatbot's output stays in the chat window. An agent can call APIs, update records in a CRM or ERP (the systems that hold your customer and business data), change order statuses and even move money.
  • A multi-step loop. The agent splits a large goal into subtasks, writes prompts for its own internal model calls and runs them one after another. When a step fails, it revises the plan and retries instead of stopping to wait for a human.

Think of a chatbot as a store clerk who answers one question and an agent as an assistant you send out with an errand. A Habr explainer on AI agents puts it in simple terms. Ask a chatbot whether a store has a specific phone in stock, and it says yes or no. An agent given the same request would locate the phone and, if it's unavailable, look for alternatives and for other shops that carry it.

LLM call / chatbot AI agent
Input A question A goal
Output Text in a chat window Actions in external systems
Path Set by the user's prompts Chosen by the agent
On error Stops and waits Re-plans and tries again

The third row is the one that bites. Because the agent chooses its own path, you trade predictability for flexibility. Also keep in mind that the label is cheap. According to the Habr explainer, some vendors market an ordinary model as an “agent.” If there's no goal, no tools and no loop, it isn't one.

Chatbot vs. agent across input, output, path and error handling.
Chatbot vs. agent across input, output, path and error handling.

Why Teams Build Agents They Don't Need

Most overbuilt AI projects don't start with bad engineering. They start with grabbing the most impressive option available. One engineer writing on DEV Community calls this “resume-driven AI engineering”: picking an agent framework, a vector store (a database that searches by similarity) or an orchestration layer across several models because it looks impressive on a résumé, not because the task requires it.

In a demo, that extra machinery is free. But in production, the same writer argues, it makes the system slower, pricier and harder to debug, and it adds randomness to steps that never needed any. Every component carries its own bill: extra latency, token spend, a fresh way to fail and one more thing an on-call engineer has to understand in the middle of the night.

The deeper issue is that the question gets framed wrong. Teams ask whether a language model can handle the task. It usually can. The better question, as that post puts it, is “is an LLM the simplest thing that does this reliably?” In other words, a lot of so-called AI systems turn out to be ordinary software with a model attached to a step that worked fine without one.

When Not to Use AI Agents: Three Jobs Where Boring Code Wins

The three scenarios below are composites the DEV Community writer describes as illustrative rather than specific incidents. They're still worth studying, because each one maps to a pattern you can spot in your own backlog.

Fixed-format extraction: use a regex

A team needs invoice numbers from incoming email. Every number has the same shape: “INV-” followed by eight digits. Instead of matching that pattern, they send each email to a model. In the writer's hypothetical, it gets the right answer 98% of the time. In the remaining 2%, it reformats the number or grabs a purchase order ID instead. Every call is billed and tacks on a few hundred milliseconds of delay.

A one-line regular expression (a short text pattern) such as \bINV-\d{8}\b does the same job deterministically, in microseconds, for essentially nothing, and you can write tests for it. The model doesn't have to disappear. Just call it only when the pattern comes back empty.

Factual lookup: use SQL

“Show me unpaid orders from customer 4417 over the last 30 days” sounds like a natural-language question, but it's really a filter. Pushing it through embeddings (numeric fingerprints of meaning), a vector store and a model summary of the top results risks quietly leaving out some of the matching orders. A plain SELECT with three WHERE conditions returns every match, uses your indexes and leaves an audit trail. Vector search earns its place when you're matching meaning (finding tickets similar to a complaint, for example), not when the answer is a set of facts.

Enumerable logic: use a decision tree

A support flow sends enterprise billing issues to account management, bugs to a ticket queue and everything else to an FAQ link. That's a handful of branches. Built as an autonomous agent with tools, memory and a planning loop, the routing becomes probabilistic, sometimes loops, and nobody can explain why a particular ticket landed with the wrong team. Three if statements would never surprise anyone. A good rule of thumb: if the logic fits on a whiteboard, an agent shouldn't have to work it out again for each request.

Job Tempting choice Better choice Cost of the tempting choice
Pull IDs in a fixed format LLM call per email Regex, model as fallback Occasional wrong IDs, per-call cost and latency
Answer a factual filter Vector search + LLM summary SQL query Matches silently missing
Route by known rules Autonomous agent Decision tree Unexplainable routing, loops
Regex, SQL and decision trees cover jobs teams often hand to agents.
Regex, SQL and decision trees cover jobs teams often hand to agents.

The Cost of AI Agents in Production, After the Demo

Getting a prototype is no longer the hard part. With a coding assistant, a short prompt and a few tool connections, you can have a working agent quickly. Call that stage “day one.” The real work is day two: running the thing for actual users.

Here are the questions day two raises that a demo never asks:

  • Trust. Can you rely on what it says? Take a Kubernetes troubleshooting agent, the kind DevOps educator Abhishek Veeramalla builds in his AI DevOps project.

    It might blame a memory shortage after finding real evidence, or it might find nothing and let the model guess. From the outside, both answers look identical, so you need a separate evaluation layer.

  • Runtime. Where does it run, how does it handle many users and simultaneous incidents, and what restarts it after a crash?

  • Memory. Should it recall past incidents, and where does that history live?

  • Identity and access. Each agent needs its own identity and permissions. A read-only investigator shouldn't hold the write access that an infrastructure-changing agent requires.

  • Observability and cost. You need a record of what it did and a live view of model spend. Cheaper models help, but their output may be weaker, so they fit low-reasoning steps.

All of this lands on top of the build itself. The author of the Habr explainer estimates, without citing data, that developing and rolling out an agent costs several times more than a basic LLM integration. Treat that as a rough opinion, not a benchmark. But the direction is hard to argue with.

Autonomy cuts both ways

An agent chooses its own path to the goal, which is exactly what makes it risky once it can change records or spend money. Without real oversight, agents can wander into areas nobody asked them to touch, the Habr explainer warns. In the Habr author's view, testing and controlling model behavior is also expensive, and there is still no definitive method for it.

What changes when an agent moves from a working demo to real users.
What changes when an agent moves from a working demo to real users.

The Five-Question Test Before You Build

Run your task through these questions before you pick a framework, add a model call or design an agent. They're adapted from the DEV Community checklist.

  • Does the data arrive in a predictable structure, or must the result follow a fixed format? Reach for parsing, a regex or schema validation first.
  • Are you looking up facts or matching meaning? Send facts to SQL. Only meaning justifies embeddings.
  • Can you list every branch of the logic? If so, code those branches directly. Save agents for open-ended work where listing them is impossible.
  • What does a mistake look like? If an error means bad data that nobody notices, you need determinism (the same input always giving the same output).
  • Who maintains this a year from now? Each framework adds a dependency that someone must upgrade, learn and debug.

So how do you read the results? If the first three questions point toward code, you don't need an agent, and probably not a model either. If the fourth answer is “silent bad data,” keep that step deterministic. The fifth question is the tiebreaker when two options look equally capable: pick the one your team can still understand when it breaks.

A five-question checklist to run before choosing a framework or an agent.
A five-question checklist to run before choosing a framework or an agent.
A one-page cheat sheet: tool choices, the five-question test and the cost of a “good enough” model.
A one-page cheat sheet: tool choices, the five-question test and the cost of a “good enough” model.

When to Use AI Agents: Signs an Agent Is the Right Call

None of this is an argument against AI. Language models earn their keep where the input is genuinely ambiguous: free-form text, fuzzy matching and generation. Agents go one step further and belong where the path itself is unknown.

A task is a reasonable agent candidate when all three of these hold:

  1. The goal is clear, but the steps depend on what turns up along the way, like investigating a failing service across logs, events and recent changes, or finding a substitute when a product is out of stock.
  2. Nobody can enumerate the logic in advance, so writing branches would mean guessing.
  3. The system must act through tools rather than just answer, and you're ready to fund the day-two work: evaluation, runtime, identities and monitoring.

Even then, a hybrid usually wins. Let deterministic code handle the predictable bulk and hand only the messy leftovers to a model or agent, the same pattern as the regex-first invoice extractor. And be honest about scope. Outside business process automation, the Habr explainer notes, most people are well served by a plain chatbot.

Key terms

  • AI agent — A system that pursues a goal on its own by planning steps, choosing tools and taking actions, rather than just answering a single prompt.
  • Deterministic — Behavior where the same input always produces the same output, which makes a system testable and its errors predictable.
  • Decision tree — Logic written as explicit if/then branches, suitable whenever every case can be enumerated in advance.
  • Vector search — Retrieval that matches items by similarity of meaning using embeddings, rather than by exact values.
  • Day-two operations — The work of running an agent in production after the prototype works: trust, deployment and scaling, memory, identity, observability and cost.
  • Agent identity — A distinct identity for each agent with its own access rights, so a read-only agent can't do what an infrastructure-changing one can.
  • Resume-driven AI engineering — Choosing impressive AI components because they look good on a CV rather than because the problem needs them.

Bottom Line: Make the Agent the Exception, Not the Default

Summary

An agent is defined by a goal, tools and multi-step actions, and that autonomy is exactly what makes it expensive to run and hard to trust. Before building one, ask whether a model is the simplest thing that does the job reliably. For fixed formats, factual lookups and rules you can sketch on a whiteboard, it isn't. So run the five-question test on your next project, start with the most boring solution that works, and reach for an agent only where the task is genuinely open-ended.

Check yourself

1. Incoming emails contain order IDs in the fixed format ORD- plus six digits. What's the best first approach to extract them?

  • A. Send every email to an LLM with an extraction prompt
  • B. Embed the emails and use vector search to find the IDs
  • C. Build an agent that reads the inbox and decides what to extract
  • D. A regex, with an LLM fallback only for emails where the pattern finds nothing
Show answer

Answer: D. Fixed-format input calls for deterministic parsing; the model is reserved for the messy leftovers.

2. A manager asks: "Show all unpaid invoices from client 210 over the last quarter." Which tool fits?

  • A. A SQL query with filters on client, status and date
  • B. A vector store plus an LLM summary of the top matching chunks
  • C. An autonomous agent with database access
Show answer

Answer: A. It's a question about facts, so it needs an exact, complete and auditable filter, not semantic matching.

3. Which of these tasks are reasonable candidates for an agent? Select all that apply. (Select all that apply.)

  • A. Routing support tickets by plan type and issue category into three queues
  • B. Validating that a form field matches an email format
  • C. Investigating why a service failed, by reading logs, events and recent changes across several systems
  • D. Finding a requested product and, if it's out of stock, searching alternatives and other stores
Show answer

Answer: C, D. It's open-ended, multi-step and the path can't be fully enumerated in advance. The goal is fixed but the path depends on what the agent finds, which is what agents are for.

4. Your agent prototype answered every test question well. What is the most important reason it isn't ready for production yet?

  • A. The model may need a bigger context window
  • B. You still can't tell whether an answer is grounded in evidence or a confident guess, and there's no runtime, permissions or cost control
  • C. Prototypes built with coding assistants never work in production
Show answer

Answer: B. Day-two issues — trust, runtime, identity, observability and cost — separate a prototype from a production system.

5. A wrong answer from your system would silently write bad data into reports. What does the decision test suggest?

  • A. Add an agent that double-checks the first agent
  • B. Use the most capable model available to reduce errors
  • C. Prefer a deterministic solution for that step
Show answer

Answer: C. When errors are silent, you want behavior that is predictable and testable.

About the author
Arnold Explorer

I build small online businesses, mostly the unglamorous kind: directories and niche sites that quietly pay the bills while bigger ideas burn cash. When someone shares a success story, I take it apart to see what they actually did and which numbers don't hold up. There are always a few.