How to Control AI-Generated Code Without Reviews

6.7K views
•
July 31, 2026
by
AI Engineer
YouTube video player
How to Control AI-Generated Code Without Reviews

TL;DR

Reliable AI-generated software requires stable architectural rules, mandatory design review, automated dependency checks, and continuous agent-based testing. Boundary lets engineers use different AI tools and generate code quickly, while compact documentation, enforced invariants, execution traces, and comparative tests reveal errors and prevent generated changes from crossing critical system boundaries.

Transcript

[music] >> Fighting slop with slop. My name is Vaibhav and I'm going to talk about something that is a little I would say maybe a little silly at first. I'm going to show you our team's engineering practices really quickly. We do no code reviews. We require every engineer to work on things in parallel. And we have no standardization on how people d... Read More

Key Insights

  • Slop is any code that people do not read, regardless of whether a human or an AI system produced it. Because engineers will increasingly leave generated code unread, reliability must come from enforceable constraints, observable behavior, and systematic validation instead of universal line-by-line inspection.
  • A small architecture document can standardize work across different AI tools. Boundary records only durable rules, such as the compiler's layers, so Claude, Codex, or another system can share the same architectural context without depending on a tool-specific instruction format.
  • Written design is treated more strictly than generated implementation code. Boundary built a versioned design-document system backed by Markdown and command-line scripts, then connected it to Slack so updates became visible and attracted immediate discussion from people across the company.
  • Mandatory readership protects design documentation from becoming generated noise. After one person began producing about ten design documents per day, Boundary required authors to secure actual readers, creating social accountability and improving the quality of proposals before implementation proceeded.
  • Automated dependency checks keep the codebase convergent. Boundary visualizes internal and external dependencies, semantic boundaries, and packages, while command-line checks and continuous integration identify commits that introduce forbidden dependencies or allow architectural layers to leak into one another.
  • Continuous agent testing evaluates both correctness and efficiency. Agents create BAML programs from scratch, preserve their full interaction transcripts, and inspect which tools were used, what failed, and which tasks required more tool calls than expected, giving humans structured evidence for reviewing reported issues.
  • Comparative experiments can replace intuition when evaluating language features or agent instructions. Boundary can test alternatives against measures stated in the talk, including fewer tool calls, fewer errors, and correct outcomes, then use the results to guide fixes and system design.
  • Foundational language behavior determines how much slop higher layers inherit. The talk argues that JavaScript and TypeScript contain behaviors designed partly around human productivity, while a language designed for agents should use strong types and function boundaries to make invalid or ambiguous outcomes impossible to ignore.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: How can teams control AI-generated code without code reviews?

Teams can replace universal line-by-line review with several reinforcing controls: a compact architecture document, mandatory review of design proposals, automated dependency checks, continuous program generation, transcript inspection, and tests of actual outcomes. Boundary allows implementation code to be generated quickly, but it constrains that code with durable rules and identifies failures through agents, command-line checks, continuous integration, and human assessment of reported issues.

Q: What should an architecture document for AI agents contain?

An architecture document should contain only compact, durable information that is unlikely to change for months or years. Boundary uses it to describe the layers of its compiler and to tell agents when deeper work requires communication with another person. Keeping the document small makes it understandable across different models and prevents temporary preferences from being mistaken for permanent architectural rules.

Q: Why does Boundary require people to read design documents?

Boundary requires actual readers because AI can make document production so cheap that written proposals themselves become slop. The talk describes a period when one person generated about ten design documents per day, leaving the team to fight the resulting volume. Requiring readership creates accountability, forces proposals to earn attention, and helped the company produce design documents of much higher quality.

Q: How does Boundary prevent architectural dependencies from leaking?

Boundary uses a tool that visualizes its dependency graph, including semantic boundaries, individual packages, and some external dependencies. It also builds command-line checks that guarantee selected invariants cannot be broken. When an agent creates a package or introduces a leaky dependency, continuous integration or the Git commit history can identify where the violation appeared, helping the codebase remain architecturally stable.

Q: How do agents test other agents' programming work?

Agents continuously create BAML programs and attempt to build working examples from scratch. Boundary records the complete agent transcript, including tool usage and intermediate events, then lets humans and other agents inspect it. Reviewers can identify incorrect language behavior, hallucinated findings, and inefficient sequences where an operation used three tool calls even though it should have required only one.

Q: How can teams compare AI coding approaches objectively?

Teams can run comparative tests instead of guessing which language feature, workflow, or agent skill is better. The talk proposes measuring which alternative uses fewer tool calls, produces fewer errors, and reaches the correct outcome. This turns agent development into a data-driven process, while humans still help decide which reported problems are genuine, hallucinated, or lacking appropriate judgment.

Q: Why are TypeScript and JavaScript criticized for AI coding?

The talk argues that TypeScript balances correctness with productivity, particularly human productivity, while JavaScript contains surprising behavior such as converting values to strings during sorting. Because TypeScript is layered over JavaScript, it inherits foundational behavior that later tools must patch or constrain. Systems built for agents from first principles could instead eliminate ambiguous behavior at the language and type-system level.

Q: How does BAML constrain AI agents while supporting fast development?

BAML is presented as a language that works across Python, TypeScript, or Rust while placing strong boundaries around each function. Its type system can infer types, require division by zero to be handled before code builds, and avoid silent unknowns that agents might otherwise guess about. The goal is to let agents move quickly inside explicit walls that they cannot breach.

Summary & Key Takeaways

  • Boundary replaces conventional code review with a system of stable constraints, mandatory written design review, automated architecture checks, and continuous testing. Engineers may use whichever AI tools they prefer, but their output must remain within documented compiler layers and other invariants that are designed to stay unchanged for months or years.

  • Design documents are stored as Markdown, managed through simple command-line tools, versioned, discussed through comments, and announced in Slack. Because excessive AI-generated documentation became another form of slop, the team added a requirement that people must actually read each submitted design document, which raised the quality of the resulting proposals.

  • Agents continuously create BAML programs, record their complete tool-use transcripts, inspect one another's work, identify errors and inefficient workflows, and propose fixes. The team can compare language features or agent skills by measuring tool calls, errors, and correct outcomes, while stronger type boundaries prevent entire categories of ambiguous behavior.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from AI Engineer 📚