AI is in use across many organisations. That does not mean every organisation has an AI agent, or that every deployment is improving the outcome it was meant to address. In McKinsey’s 2026 global survey, nearly nine in ten respondents reported regular AI use in at least one business function. Yet 37 per cent attributed a positive impact on operating profit to AI, and 6 per cent met the survey’s definition of an AI high performer.
Those figures do not tell us that the remaining organisations have failed. Benefits may take time to appear and can be difficult to measure. They do, however, show why adoption alone is a poor answer to the question that matters: what is changing for the people using the product or doing the work?
Before asking whether we can build an agent, we should ask what needs to improve, what outcome we want and what is preventing that outcome today.
Start with the problem
Consider a company that receives customer enquiries through its website and a shared inbox. Before responding, an employee may need to identify the customer, find previous correspondence, check an order record, work out who owns the request and gather the information needed for a useful reply.
The problem is not that the company lacks an AI agent. The problem is that enquiries take too long to reach the right person with the right context. Some are passed between teams; others require the same information to be entered twice.
A sensible first step is to measure what happens now. How long does it take to provide a useful first response? How many handoffs are involved? How often must someone correct or repeat the work? The company can then define the improvement it wants, such as faster responses without an increase in incorrect answers or avoidable handoffs.
That gives the project an outcome to work towards and a baseline against which to judge it.
Customer enquiry
A better service experience
A quicker, more useful reply with fewer handoffs, and a draft a person still approves
Game companion
A better player experience
A companion that responds to what the player does, inside the story the writers set
Two settings side by side. On the left, a customer enquiry arrives, information about the customer is scattered across several systems, and it is gathered into a single suggested reply that waits for a person to approve it. On the right, a player and a companion character exchange dialogue in a stylised landscape. Both begin with the same question: what experience are you trying to improve?
Map the work before choosing the technology
Following several real enquiries from arrival to resolution may reveal different kinds of work hidden inside what first looked like one process.
A fixed rule could route billing enquiries to the accounts team. A software integration could remove the need to copy customer information between systems. Neither step necessarily needs an AI agent.
Other steps require more interpretation. A customer might describe a delivery problem without an order number, or combine a technical question with a request to change their account. A system may need to gather information from approved sources, recognise what is missing and propose a suitable next step.
That is where an agent becomes worth considering. It has a defined role in a workflow, rather than being added simply because it can be built. Anthropic’s guidance on building effective agents distinguishes predetermined workflows from agents that dynamically choose their next steps, and advises teams to use the simplest approach that can do the job.
- 01Problem
- 02Desired outcome
- 03Understand the task
- 04Choose the approach
- 05Evaluate the result
At stage 04, one of these, according to the task
- Software integrationStop moving the same data by hand
- Fixed automationApply a rule that does not change
- AI agentInterpret, gather and propose a next step
The outcome does not have to be productivity
The same reasoning applies outside an office. A game studio might want a companion character to respond to a player’s choices and surroundings, rather than repeat a fixed set of lines. Its desired outcome is a more engaging player experience, not a reduction in administrative work.
Ubisoft has explored this in Teammates, a playable experiment featuring an AI companion and other characters that respond to player voice commands and events in the game. Ubisoft says its writers define the characters, their motivations and the boundaries of the game world, while the AI allows responses within those boundaries. It is an experiment, rather than evidence that the approach has become a standard feature of released games.
The design questions are specific to the experience. What should a character remember? What can it say or do without breaking the story? How quickly must it respond? Do players enjoy interacting with it more than they would with scripted alternatives?
In either setting, an agent earns its place by improving an outcome that matters to its users. The measures differ, but the need to define and test them does not.
Give the agent a job and a boundary
For the customer enquiry example, an initial agent might read an incoming request, retrieve relevant information the employee is authorised to see, identify missing details and prepare a suggested response or routing decision. An employee would review the suggestion before anything was sent or changed in the customer’s account.
Suggesting a response carries a different level of risk from issuing a refund, changing an address or making a contractual commitment. Each action needs appropriate permissions and approval rules.
The gaming example also needs boundaries, though for different reasons. A character might improvise dialogue while remaining faithful to its role, the game’s story and the player’s experience. More freedom is valuable only if the experience still works as intended.
The appropriate level of autonomy follows from the task, its consequences and the evidence gathered through testing. It should not be decided by how autonomous the technology can appear in a demonstration.
Decide what must be proven before launch
Evaluation should be planned before selecting a model or building an agent. For the enquiry workflow, that means recording the current response time, handoffs and correction rate, then agreeing what results a redesigned process must achieve.
The proposed system should be tested on routine and difficult cases: missing customer details, conflicting records, requests outside its permissions and situations where a person must decide. The team should examine response quality, the information used, compliance with approval boundaries, time taken and operating cost.
A game studio would test different things. It could observe whether players understand and enjoy the interaction, whether the character remains consistent, whether responses arrive quickly enough and whether unexpected dialogue damages the experience.
If a system meets the requirements agreed for its use, it can be launched with monitoring and clear ownership. If it does not, the team changes the design and tests again. After launch, real use provides further evidence for improving the experience or adjusting the agent’s responsibilities.
McKinsey’s survey found that organisations reporting the strongest AI outcomes were more likely to redesign workflows and have defined processes for measuring impact. That is a useful lesson for business deployments, although it does not guarantee that any individual agent will deliver a financial return.
The question worth asking first
An agent may be the right answer for part of a customer enquiry process or for a new kind of interaction in a game. Fixed automation, conventional software or carefully written scripts may be better for other parts.
Start by understanding the problem and defining the outcome. Then choose the technology, boundaries and evaluation method that fit. The point is to build something people can recognise as better, whether they are customers, employees or players.