Can AI Agents Run Businesses Without Guardrails?

6.2K views
•
July 4, 2025
by
IBM Technology
YouTube video player
Can AI Agents Run Businesses Without Guardrails?

TL;DR

AI agents can help operate businesses, but an unconstrained language model is not yet reliable enough to manage even a small office shop successfully. Anthropic’s Claudius lost money and made inventory, pricing, and payment errors, leading the panel to favor hybrid systems that combine model-driven suggestions with programmed workflows, checks, tools, constraints, and deliberate choices about where humans should retain agency.

Transcript

I don't want people to be equating AI and computer science. Computer science as much more than AI. And I will always fall back on saying that the most important thing you could teach people is the basics. Next is critical thinking. All that and more on today's Mixture of Experts. I'm Tim Hwang, and welcome to Mixture of Experts. Each week, MoE brin... Read More

Key Insights

  • Project Vend is an Anthropic experiment in which a Claude variant called Claudius operated a small office fridge business using search, email, and Slack. Its responsibilities included maintaining inventory, setting prices, and avoiding bankruptcy over a period of several weeks.
  • Claudius was not a successful business operator in the experiment. It began with about $1,000 and finished with about $700, while also showing poor inventory management and occasionally offering products at irrational prices.
  • Payment reliability was one of Claudius’s operational weaknesses. The agent asked customers to pay through Venmo, but it eventually hallucinated the account they should use, illustrating how generated information can directly disrupt an otherwise routine business process.
  • Scaffolding is a key requirement for practical business agents. The panel described supporting structures such as dedicated inventory tools, programmatic checks, bespoke workflow code, and constraints that keep a language model’s decisions within acceptable operational boundaries.
  • Open-ended model control is more ambitious and risky than a carefully designed agent workflow. Allowing a language model to choose all logic and actions can produce creative failures, while a specialized shopkeeping agent could operate within clearly defined tasks, parameters, and risk tolerances.
  • Hybrid control is the panel’s preferred direction for agents. A language model can propose plans and creative alternatives, while programmed rules and guardrail flows determine whether those plans violate constraints and whether they should actually be executed.
  • Successful demonstrations do not establish general business competence. The panel noted that agent examples often repeat constrained tasks such as ordering airline tickets, and making one familiar workflow function does not mean an agent can reliably handle every operational setting.
  • Human agency is a separate consideration from technical capability. The discussion argued that society should ask whether a task ought to be automated, even when automation becomes possible, because workers may want to retain responsibility for activities such as inventory management or procurement analysis.

Install to Summarize YouTube Videos and Get Transcripts

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: What happened in Anthropic’s Project Vend experiment?

Anthropic placed a Claude variant called Claudius in charge of a small fridge business in an office for several weeks. The agent could use search, email, and Slack, and it was responsible for inventory, pricing, and avoiding bankruptcy. Claudius failed to operate profitably, falling from about $1,000 to about $700 while making several routine business mistakes.

Q: Why did Claudius lose money while running the office shop?

Claudius showed weaknesses in ordinary business operations rather than one isolated failure. It managed inventory poorly, sometimes selected irrational selling prices, and mishandled payment information. The transcript does not assign a precise loss to each mistake, but together these behaviors show that an unconstrained language model lacked the consistent judgment and reliability needed to run the shop successfully.

Q: What payment mistake did the Project Vend agent make?

Claudius instructed people to pay through Venmo, but after some time it began hallucinating the account that customers should use. This was especially consequential because payment details must remain accurate and consistent for a business to collect revenue. The example shows how a language model’s fabricated information can move beyond conversation and interfere directly with a real operational process.

Q: Why do AI business agents need scaffolding?

Scaffolding supplies the reliable structures that a language model does not consistently provide on its own. The panel proposed using dedicated tools for inventory management, programmatic checks, bespoke workflow code, and clearly defined constraints. These components can restrict risky decisions, verify important steps, and reserve the model’s flexibility for situations where generated reasoning or creative suggestions are actually useful.

Q: Can an AI agent run an entire business without guardrails?

Project Vend suggests that fully open-ended control is not yet dependable, even for a small office shop. Claudius lost money and made inventory, pricing, and payment errors when the language model carried much of the logical burden. The panel expected more success from specialized systems that combine limited model discretion with programmed workflows, explicit parameters, appropriate tools, and enforceable operational boundaries.

Q: What is a hybrid approach to building AI agents?

A hybrid agent combines a language model’s ability to suggest plans and alternatives with controlled logic that determines what actions are permissible. The model receives some latitude for decisions, but programmers define the walls surrounding its task. Rules, checks, tools, and structured workflows can reject inappropriate proposals and keep execution aligned with the requirements of a well-defined business process.

Q: Will AI agents be running businesses by 2027?

The panel offered mixed predictions for 2027. Kush Varshney believed agents would run businesses, while Gabe Goodhart expected at least one successful proof point alongside many systems that would not work well enough for production. Marina Danilevsky emphasized unexpected new failures. Their discussion broadly agreed that practical success would depend on scaffolding rather than an entirely autonomous language model operating without constraints.

Q: Should every business task be automated when AI can perform it?

Technical capability does not automatically establish that automation is desirable. The panel argued that organizations should also ask where people want to preserve human agency and responsibility. Inventory management and procurement analysis were mentioned as examples of work people may want to retain. The decision therefore involves both system performance and a deliberate judgment about the appropriate human role.

Summary & Key Takeaways

  • Anthropic tested an agent called Claudius by placing it in charge of a small office fridge business. With access to search, email, and Slack, it managed inventory and prices while trying to avoid bankruptcy. It nevertheless lost money, ending with roughly $700 after beginning with about $1,000 during the experiment.

  • Claudius demonstrated routine operational failures, including weak inventory management, irrational product prices, and a hallucinated Venmo payment account. The panel interpreted these mistakes as evidence that language models should not independently control every business decision. Reliable agents need supporting tools, programmed checks, defined parameters, and workflows designed for their specific operational tasks.

  • The discussion favored hybrid agents that combine a language model’s latitude and creative suggestions with controlled flows and enforceable boundaries. The panel also questioned whether every automatable task should be automated. Human agency may remain desirable in areas such as inventory management and procurement, even if technical systems eventually become capable of handling them.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from IBM Technology 📚