The Prompt Is the Mind: What Benchmark Design and Lucid Dreaming Have in Common

Mark Erdmann

Hatched by Mark Erdmann

May 29, 2026

10 min read

41%

0

What if intelligence is less about what a system knows, and more about how you ask it to come alive?

A strange thing happens when you look closely at modern AI behavior: small changes in prompting can flip a model from sharp to confused, from useful to unreliable, from brilliant on objects to blind on motion. That sounds like a technical nuisance. But it may actually be a clue about something deeper. Perhaps performance is not just a property of the model itself. Perhaps it is also a property of the frame in which the model is invited to think.

That idea becomes even more provocative when you compare two very different human experiences. On one side, there is the engineering problem of building benchmarks that can generate vast numbers of tailored tasks to reveal what a multimodal model really understands. On the other, there is the attempt to induce lucid dreaming, where a chemical nudge can alter the mind’s state enough to make the dream become self-aware. In both cases, the hidden question is the same: how much of intelligence is latent, and how much is evoked by the right conditions?

The answer matters far beyond AI or sleep. It changes how we think about evaluation, teaching, creativity, and even self-knowledge.


The real problem is not whether intelligence exists, but whether it can be elicited

Most people imagine evaluation as a neutral act. You test a system, measure the result, and discover the truth about its capabilities. But that model is too simple. In practice, every evaluation is also a kind of invitation. It tells the system what counts as relevant, what kind of response is expected, and what context to treat as important.

That is why a benchmark engine that can generate hundreds of millions of image and video question-answer pairs is so interesting. It does not merely measure a fixed ability. It exposes how fragile and context dependent that ability can be. A model may excel at identifying objects and attributes, yet fail badly at spatial relations or temporal change. It may recognize a chair and a red ball, but lose track of whether the ball moved left or right, or whether the rotating object is still the same object from frame to frame.

This is not just a failure of visual perception. It is a failure of world model coherence. The system can name the pieces but struggles to maintain the dynamics that connect them.

Now compare that with lucid dreaming. The dream is already happening. The mind is generating vivid scenes, people, motion, and narrative without any external input. But lucidity is not automatic. It has to be induced, often through a subtle shift in state. A chemical like galantamine can help make the dream self-aware, as if the dream were a model that briefly realized it was generating its own world.

That parallel is not decorative. It points to a shared principle: capability can remain invisible until the environment, prompt, or state change makes it legible. A model may know more than it appears to know. A dreamer may see more than they realize. The challenge is not simply to add more information, but to create the right conditions for the latent structure to surface.

Intelligence is often not a pile of facts. It is a coordination problem between capacity, context, and state.


Why prompt sensitivity is not a nuisance, but a diagnostic

In AI, prompt sensitivity is usually treated as a flaw. One prompt yields excellent results, another disappoints. One model prefers detailed instructions, another does better when you keep it brief. That seems like inconsistency, and inconsistency is bad, right?

Not necessarily. Prompt sensitivity may be the most revealing evidence we have that many systems do not possess a fully stable internal abstraction of the task. Instead, they behave more like instruments that resonate differently depending on how they are struck. A violin does not become less musical because it responds to bowing rather than hammering. The difference is that we know a violin needs the right excitation. We are only beginning to accept that multimodal models may be similar.

This reframes prompt engineering from trickery into diagnosis. If a model performs much better with a detailed prompt, it may need explicit scaffolding to assemble the task structure. If it performs better with a succinct prompt, too much instruction may actually interfere, perhaps by cluttering the relevant signal or activating the wrong heuristics. The prompt is not merely a query. It is a test of the system’s ability to self-organize under different kinds of guidance.

That has a striking implication: some failures are not failures of knowledge, but failures of orchestration.

Imagine a pilot who can fly only when the dashboard is uncluttered, and a different pilot who needs every instrument labeled in detail. Neither is simply better in every environment. Their performance reveals different forms of dependence. The same is true for AI models. Some are robust to terse commands but brittle when overloaded. Others need rich context to unlock their best behavior. If you only measure one prompting style, you may mistake a contextual dependency for a true capability gap, or miss a real strength entirely.

This is why tailored benchmarks matter. They do not just rank models. They reveal the shape of dependence itself.


The hidden frontier is not recognition, but continuity

The most important finding hiding in these evaluations is not that models can recognize objects. It is that recognition is the easy part. The hard part is continuity.

A static image is a frozen world. You can identify a dog, a cup, a tree, a traffic light. Once the scene starts moving, the task changes. Now the model has to preserve identity across frames, understand causality, infer spatial transformations, and keep track of time. That is a much higher bar, because the system is no longer just matching labels. It must sustain a model of how the world changes.

This is exactly where many advanced systems still stumble. They can tell you what is present, but not what persists. They can detect an object, but not reliably track whether it has rotated, moved behind another object, or changed relation to the others in the scene. That is why a model can look impressive in image tasks and still be exposed by video tasks. Motion reveals the difference between description and comprehension.

A lucid dream offers a human analogy. In a non-lucid dream, scenes unfold without stable self-monitoring. You are inside the narrative, but you do not fully track the fact that you are in a narrative. Lucidity is a form of continuity: the dreamer maintains a thread of awareness through a state that normally disrupts that thread. The dream is still vivid, but now it is also observed.

This suggests a broader framework for intelligence evaluation:

  1. Naming: Can the system identify discrete elements?
  2. Binding: Can it associate the right attributes, relationships, and roles?
  3. Persistence: Can it track those entities across time and transformation?
  4. Self-monitoring: Can it recognize the limits or structure of its own state?

Many benchmarks stop at naming. The more interesting frontier is persistence and self-monitoring. In other words, not just what the system sees, but whether it can maintain a coherent model of what is happening as the world changes.

That is a more human problem than it first appears. We too are often excellent at naming and surprisingly poor at continuity. We remember fragments, not trajectories. We notice facts, but not transitions. We think we are describing reality when we are often just snapshotting it.


Benchmarks and lucid dreams are both state machines

Here is the deepest connection: both benchmark design and lucid dreaming are ways of manipulating a state machine.

A benchmark engine creates a controlled environment that can systematically probe different aspects of cognition. By varying object categories, attributes, relationships, and task formats, it reveals which internal states a model can sustain. It does not just ask, “What can you answer?” It asks, “Under what conditions can you reliably enter the state in which you answer well?”

Lucid dreaming works similarly, though inwardly. The mind is already generating a simulation, but it is usually immersed in it. A lucid dream is a special state in which the dreamer can step into meta-awareness without fully waking. A pharmacological trigger may help the system transition into that state. Again, the point is not raw content. The point is state transition.

This is a powerful mental model for thinking about capability in general. Many of the most important human and machine abilities are not binary traits. They are attractor states. You are not simply “good” or “bad” at something. You may be able to enter a mode where the relevant skills cohere, but only under the right setup.

Consider these examples:

  • A writer who needs silence to think may not lack ideas, only the right cognitive state.
  • A student who performs poorly in a traditional exam may understand the material but fail to enter retrieval mode under stress.
  • A model that shines on detailed prompts may require explicit scaffolding to stabilize its internal representation.
  • A dreamer who rarely becomes lucid may still have the latent capacity, awaiting the right trigger.

Seen this way, evaluation becomes less like grading a statue and more like tuning a radio. You are searching for the frequencies at which the signal becomes clear.

We often overestimate the importance of content and underestimate the importance of state.

That single shift in perspective changes everything. It suggests that the best systems are not necessarily those with the most knowledge, but those with the widest repertoire of usable states, and the smoothest transitions between them.


The practical lesson: stop asking only what works, and ask under what conditions it works

The most actionable insight from this synthesis is simple: if you want to understand capability, inspect the conditions that activate it.

This applies to AI teams designing evaluations. It applies to educators deciding how to test students. It applies to managers trying to get better performance from themselves and others. It even applies to anyone trying to cultivate better sleep, creativity, or insight.

Instead of treating variability as noise, treat it as data. Ask:

  • What kind of framing improves performance?
  • Which tasks collapse when context is removed?
  • Which forms of instruction help, and which obstruct?
  • What state changes reveal hidden ability?
  • What does the system do when it has to maintain continuity across time?

The goal is not to eliminate dependence on context. The goal is to understand it so well that you can shape it intentionally.

A benchmark that can generate immense numbers of tailored tasks is valuable because it lets us see the contour of a model’s dependence on prompt style, modality, and task structure. A lucid dreaming aid is valuable because it lets a person enter a different relation to their own mind. In both cases, the real prize is not the content produced in the moment. It is the discovery of the conditions under which deeper capacities become accessible.

For humans, this is liberating. It means a bad performance is not always a verdict on ability. Sometimes it is a mismatch between the person and the state. A thoughtful prompt, a better environment, a different rhythm of practice, or a shift in bodily arousal can change everything.

For AI, it is even more important. As systems become more capable, the central challenge may no longer be whether they can answer questions, but whether they can reliably enter the mode in which their abilities are stable, interpretable, and safe.

Key Takeaways

  1. Performance is state dependent. Do not assume a model or person’s visible output fully reflects latent capability.
  2. Prompt sensitivity is diagnostic. Differences between detailed and succinct instructions can reveal how much scaffolding a system needs.
  3. Continuity matters more than recognition. The hard test is not identifying objects, but tracking identity, motion, and relationship over time.
  4. Evaluate conditions, not just outcomes. Ask what setup, framing, or state transition unlocks the behavior you want.
  5. Treat variability as information. Inconsistent results often show where the real cognitive boundaries are.

Conclusion: intelligence is not just what appears in the frame

We tend to think of intelligence as a thing in the box, a property lodged inside a mind or model. But the more carefully we look, the more intelligence resembles a relationship: between system and prompt, between dream and trigger, between capability and state. What matters is not only what is inside the system, but what kinds of invitations it can answer.

That is why the deepest lesson here is unsettling and useful at the same time. A model that looks weak may be one prompt away from competence. A dream that seems ordinary may be one shift away from lucidity. A person who seems stuck may simply be in the wrong state to show what they know.

The future of evaluation, learning, and self-understanding may belong to those who stop asking, “How smart is it?” and start asking, “What conditions let its intelligence appear?” That is a subtler question. It is also the one that gets us closer to the truth.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣