How to Build Better Agents with File Systems

135.9K views
•
September 14, 2026
by
AI Engineer
YouTube video player
How to Build Better Agents with File Systems

TL;DR

Build specialized agents around a sandboxed file system, familiar tools such as file reading and Bash, and rich company-specific context. Vercel’s internal data agent improved substantially after evolving from one large prompt, through a chain of narrow agents, to a single stateful file-system agent whose evaluation score doubled and whose recurring queries were distilled into reusable skills.

Transcript

Hey everyone, thanks for coming. I'm Andrew. I'm the chief of software at Verscell and I'm here to talk to you about how we solved agent building at Verscell. I'm the chief of software. So I work on a mix of internal engineering, external experimentation, and generally being at the frontier and building new libraries, frameworks, and technologies. ... Read More

Key Insights

  • Vercel identified its data team as the strongest initial agent use case because data scientists repeatedly had to stop planned work, write queries, analyze results, and report recommendations whenever marketing or sales asked questions about customers or products.
  • The first data-agent prototype was a large prompt containing a dump of the Snowflake schema and a user question. It tested whether available models could generate valid SQL, but Andrew Qu still copied the generated SQL and executed it manually.
  • A multi-agent workflow divided data analysis into planning, schema exploration, SQL execution, and reporting. Each agent received specialized instructions and narrowly scoped tools, allowing the system to complete the process from a question to an answer without manual SQL copying.
  • A single stateful agent improved on the chained design because it retained the broader execution history instead of receiving only a summary from a previous agent. When SQL execution or joins failed, it could revisit earlier exploration, read more context, and try again.
  • Internal evaluations did not predict the experience of trusted users. Although the system cleared about thirty percent of Vercel’s evaluations, early users called it awful because they asked questions the development team had not anticipated, exposing the limits of manually mapped scenarios.
  • The file system was the central architectural unlock observed in Claude Code and Opus 4.5. A minimal interface built around listing files, reading files, writing files, and running Bash let the model use familiar capabilities and discover effective workflows without heavily prescriptive tools.
  • Vercel’s rebuilt agent ran in a sandbox containing the company’s semantic layer and a small number of Vercel-specific tools. This file-system-based architecture doubled the evaluation score and supported end-to-end exploration, execution, and reporting within a purpose-built environment.
  • Recurring employee questions shared common structures, including aggregations, product lookups, billing information, customer metrics, sales metrics, and npm downloads. Vercel used a recurring job to distill recent queries into roughly one hundred skills, giving future runs accumulated context instead of starting from nothing.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: How did Vercel choose its first internal agent use case?

Andrew Qu asked teams across marketing, sales, finance, and legal what they hated most about their jobs. The most compelling answer came from Vercel’s lean data team. As the company and its available customer, analytics, sales, and product data grew, data scientists repeatedly had to interrupt their planned work to write queries, analyze results, and respond to other teams’ questions.

Q: How did Vercel’s first data science agent work?

The first prototype used one large prompt. Qu obtained a dump of the Snowflake schema, placed it in the system prompt with a user question, and asked the model to generate SQL. He then copied that SQL and ran it himself. The experiment provided some confidence that models could produce valid queries when given useful structure, while also showing the need for better context and guardrails.

Q: Why did Vercel initially split data analysis across multiple agents?

Vercel mapped the work of a data scientist into distinct phases, including understanding the question, exploring the semantic layer, identifying join patterns, executing SQL, correcting failed or expensive queries, and reporting the results through visualizations or written analysis. Separate agents received focused system prompts and tools for these phases, enabling an automated path from the initial question to the final answer.

Q: Why did one stateful agent replace the chained agent workflow?

The chained workflow passed only a summary and a small portion of completed work to the next agent, limiting its ability to reconsider earlier decisions. A single agent could manage its own state while switching among planning, building, execution, and reporting. If execution or a join failed, it could inspect prior work, explore more context, read additional material, and revise its approach.

Q: Why did Vercel’s data agent disappoint its first users?

The team believed the agent was performing well because it cleared about thirty percent of its evaluations. Trusted users nevertheless described it as awful because they submitted questions that the builders had not anticipated. Expanding the system by manually mapping every new scenario did not appear scalable, so the feedback exposed an architectural problem rather than merely a need for more predefined cases.

Q: How did Claude Code influence Vercel’s agent architecture?

Claude Code working with Opus 4.5 answered many of the same questions far more effectively than Vercel’s hand-built agent. Qu concluded that the important advantage was its file-system model and minimal tools, including listing files, reading files, and running Bash. These were capabilities the models handled well, and they permitted flexible exploration instead of forcing work through a highly prescriptive tool sequence.

Q: How does Vercel’s file-system data agent operate?

The rebuilt agent runs inside a sandbox where the company’s semantic layer is made available as files. It can use Bash and file-reading and file-writing capabilities to inspect context, execute work, and store intermediate results. Vercel adds only a few tools required for its specific data environment. This purpose-built file-system design doubled the agent’s evaluation score.

Q: How did Vercel turn repeated employee queries into reusable agent knowledge?

After broader internal adoption produced thousands of daily queries, Vercel noticed that many requests had recurring shapes. Common patterns included aggregations, product lookups, billing information, customer metrics, sales metrics, and npm downloads. A recurring job processes recent queries and distills them into roughly one hundred skills, allowing later runs to begin with accumulated organizational context instead of solving every request from nothing.

Summary & Key Takeaways

  • Vercel began exploring workplace agents by asking teams which tasks they disliked most. The strongest opportunity came from its lean data team, whose members repeatedly interrupted their work to answer customer and product questions. The first prototype placed the Snowflake schema and a question into one prompt, then required manual SQL execution.

  • The architecture evolved into narrowly scoped planning, SQL execution, and reporting agents, each receiving limited outputs from the previous stage. Vercel then consolidated these roles into one stateful agent that could plan, execute, inspect errors, explore additional information, and retry. Although it performed better internally, trusted users still found major weaknesses.

  • Claude Code and Opus 4.5 suggested a stronger design based on a file system and a minimal set of familiar tools. Vercel rebuilt its agent inside a sandbox containing the semantic layer, file operations, Bash, and a few company-specific capabilities. Its evaluation score doubled, while recurring queries were distilled into roughly one hundred reusable skills.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from AI Engineer 📚