Prompts, assistants or agents? How to tell when a custom AI agent is worth building
Not every AI problem needs an agent, and building one too early wastes money. A simple ladder for deciding how far to go, with the signs that a task has outgrown copy-and-paste prompting.
“AI agents” has become one of those phrases that means everything and nothing. Vendors attach it to chatbots, automations and products that are really just a prompt with a logo. Meanwhile, plenty of organisations are paying people to copy and paste the same prompt fifty times a week, when a simple agent would do the job in the background.
The useful question is not “should we use agents?” but “how much automation does this particular task deserve?”. Here is a way to answer it.
The automation ladder
Think of AI support for a task as a ladder with four rungs. Each step up costs more to set up and saves more time per run. Most tasks should stop on the lowest rung that works.
| Rung | What it is | Set-up effort | Best for |
|---|---|---|---|
| 1. Ad hoc prompting | A person writes a prompt each time | None | Varied, one-off tasks |
| 2. Saved prompts | A tested prompt in a shared library, reused by the team | An hour or two | Recurring tasks with a consistent shape |
| 3. Custom assistant | A configured assistant (a custom GPT, Copilot agent, Claude project) with instructions and reference files built in | A day or two | Recurring tasks that need the same context every time |
| 4. Custom agent | A system that pulls data from your tools, carries out several steps, and hands results to a person or another system | Days to weeks | High-volume, multi-step work that crosses systems |
The jump from rung 3 to rung 4 is the important one. An assistant still waits for a person to start it and paste things in. An agent goes and gets what it needs.
Signs a task has outgrown prompting
A task is a good candidate for a custom agent when most of these are true:
- It happens often. Daily or weekly, not quarterly. Volume is what pays back the build.
- It follows a recognisable pattern. The inputs and outputs look similar each time, even if the content varies.
- It involves fetching and moving information. People spend more time gathering material from email, a CRM, a shared drive or a spreadsheet than doing the thinking.
- Mistakes are catchable. There is a natural review point where a person checks the result before it goes anywhere important.
- Someone already does it with AI, manually. If a team member has built a clever prompt workflow and runs it by hand every morning, that is a prototype waiting to be automated.
Signs it has not
Hold off on an agent if:
- Nobody has tried the task with a good prompt yet. You do not know whether AI can do it well. Start on rung 2.
- The process itself is unclear. If three people do the task three different ways, automating it will encode the confusion. Agree the process first.
- The task is rare or highly variable. The set-up cost will never pay back.
- The output goes straight to a client or regulator with no review. Design in a human check first; consider whether the risk is acceptable at all.
Examples that tend to work well
A few patterns come up again and again in small and mid-sized organisations:
- Inbox triage. An agent reads incoming enquiries, classifies them, drafts a reply using your knowledge base, and queues it for a person to approve.
- Meeting to action. After a client call, an agent turns the transcript into minutes in your format, logs actions in your project tool, and drafts the follow-up email.
- Report assembly. An agent pulls figures from several spreadsheets or systems each month, writes a first-draft commentary, and flags anything unusual for the analyst.
- Document checking. An agent reviews incoming documents (contracts, applications, supplier forms) against a checklist and highlights missing or unusual items.
Notice what these share. The agent does the gathering and the first draft. A person keeps the judgement.
Questions to ask before you build
Whether you build in-house or with a partner, settle these first:
- What does success look like, in numbers? Hours saved per week, turnaround time, error rate. Measure the current state before you start.
- Where does the data come from, and who is allowed to see it? For Luxembourg and other EU organisations, check the data stays within the processing terms you have agreed, and record the agent as a processing activity where personal data is involved. UK firms should do the same under UK GDPR.
- Where is the human check? Name the step and the person.
- Who maintains it? Agents need occasional adjustment as your tools, templates and processes change. Someone in the team should understand how it works and be able to make small changes.
- How will people learn to use it? An agent nobody trusts is an agent nobody uses. Plan a short handover session, and ideally teach the team enough to build simpler assistants of their own.
Start small, then climb
The organisations that get the most from AI agents rarely start with them. They start with good prompting habits, notice which tasks keep coming back, move those into saved prompts and assistants, and then automate the two or three tasks where volume and value are highest. Each rung teaches you what the next one needs.