Six Tests That Reveal an AI Agent Demo’s Empty Claims

Most AI agent demos are theater with an API attached. A recent analysis estimates that roughly 130 of the thousands of vendors marketing agentic AI offer genuine agentic capability, leaving the rest to sell familiar software under a more ambitious label. The report coined the term agent washing for rebranding chatbots and RPA tools as autonomous agents while the underlying engine stays exactly the same.
The timing makes this problem expensive. Worldwide AI spending is forecast to reach $2.59 trillion in 2026, a 47% jump year over year, while over 30% of this year’s AI investment is committed to agentic AI. When that much money chases a category, every vendor suddenly discovers that its existing product was an agent all along.
Six questions for the demo
1. What does the system do beyond conversation?
A chatbot can generate an answer. An agentic system must show more than a polished exchange in a browser window, especially when a vendor claims autonomy. Ask what action the system takes after receiving a request, what tools it uses, and how the underlying engine differs from the company’s existing chatbot or RPA product.
The question targets the central definition of agent washing: a new label without a new capability. If the demonstration shows the same engine, the same workflow, and the same limits, the word “agent” is doing the work—not the software.
2. Who controls each function?
Agent demos often focus on what the system can initiate and skip who can stop it. A serious evaluation must identify the control attached to every function, along with the person or team accountable when something goes wrong.
The governance standard is clear: “Every function has an independent control sitting to the side, and there’s a clear answer for who’s accountable when things go wrong.” Without those controls, autonomy becomes a marketing feature with an invoice.
3. What happens when the system meets a trap?
A canary tools study planted diagnostic probe tools inside agents’ Model Context Protocol toolsets across eight models. Susceptibility to those traps varied roughly 36-fold, showing that agents can respond very differently to the same kind of hidden test.
Ask the vendor to explain its failure cases, not just its successful path. A demo that avoids adversarial conditions measures presentation quality. It does not measure whether the agent can handle the environment it claims to operate in.
4. How was the agent evaluated before deployment?
Evaluation is not a final checkbox. Seventy percent of large security operations centers will pilot AI agents by 2028, yet only 15% of those centers will achieve measurable improvements unless they run structured evaluation first.
That gap separates deployment from progress. Ask which tests were run, what counted as measurable improvement, and how the system performed when its tools or instructions created conflicting demands. “It just doesn’t work” is not an evaluation plan, but it is the likely result of skipping one.
5. What measurable result does the vendor promise?
About 90% of CEOs expect AI agents to deliver measurable ROI in 2026, and half of those CEOs believe their job depends on getting AI right. A demo should therefore connect the agent to a defined business result rather than rely on speed, novelty, or an impressive transcript.
The test is simple: what changes after deployment, and how will the organization measure it? A vendor that cannot answer that question is selling activity, not value. The distinction matters when spending reaches $2.59 trillion worldwide.
6. Is the market claim larger than the real capability?
The same analysis forecasts that 33% of enterprise software applications will include agentic AI by 2028, up from under 1% in 2024. That forecast describes a fast expansion in product labels and investment, but it does not prove that every application will deliver genuine agentic behavior.
Use the number as a reason to inspect claims, not accept them. Ask whether the product belongs among the roughly 130 vendors offering genuine agentic capability or among the thousands marketing the category. The answer should come from its actions, controls, testing, and measurable results—not from the demo’s vocabulary.
Governance decides whether autonomy survives contact
CEOs are taking direct ownership of the decision. A recent survey found that 72% now call themselves their company’s main AI decision-maker, while half believe their job depends on getting AI right.
That responsibility makes agent washing more than a branding nuisance. If leaders approve systems that lack structured evaluation or clear accountability, the organization inherits the failure along with the software.
The market will keep expanding from under 1% of enterprise applications with agentic AI in 2024 toward 33% by 2028. Buyers do not need more theatrical demos; they need evidence that a system can act, withstand traps, operate within independent controls, and produce a result someone can measure. That is the difference between an agent and a chatbot wearing a tiny executive badge.
Based on




