Why the Next Great Interface Will Not Be an App, But a Behavior
Hatched by Darren LI
May 25, 2026
9 min read
3 views
71%
The real competition is not for features, but for habits
What if the next killer app for a new device is not really an app at all?
That sounds like a provocation, but it is closer to how platform shifts actually happen. New hardware rarely wins because it has the most capabilities on day one. It wins when it discovers a behavior people are already trying to do, then makes that behavior so natural, so immediate, and so satisfying that it starts to feel like part of life. The deeper prize is not software adoption. It is behavioral habit formation.
This is why the history of major devices looks less like a catalog of apps and more like a sequence of newly unlocked human gestures. The smartphone did not become indispensable because it had more programs than a desktop. It became indispensable because it turned cameras, maps, messages, music, and games into constant companions. The smartwatch did not need to be a miniature computer. It found its footing where the body lives, in health tracking, nudges, and quick glances. In every case, the winning interface did not just display content. It changed what users were willing to do effortlessly.
That is the deeper tension now facing spatial computing. A headset can show almost anything, but showing is not the same as becoming necessary. The real challenge is to identify which experiences can migrate from novelty into ritual.
A new device needs a new grammar of action
The most powerful technologies do not merely offer new outputs. They teach users a new grammar of action. A person does not think, “I am opening a device with a new feature.” They think, “This is the obvious way to do this task.”
That is why multimodal robot control and spatial computing may be more related than they first appear. In one case, a machine must learn how to act in the physical world from language, images, and examples. In the other, a device must learn how to fit into the physical world of the user, where vision, gesture, memory, and presence all matter at once. Both are trying to solve the same hidden problem: how to map intention into action when the world is messy, ambiguous, and context dependent.
Consider a robot receiving a prompt that combines a demonstration, language instruction, and a visual goal. The difficulty is not just moving an arm. The difficulty is interpreting the task as a coherent structure: what matters, what can vary, and what must be preserved. A human user wearing a spatial device faces a similar burden. They do not want a menu forest. They want the device to understand when they are trying to capture a memory, collaborate in 3D, annotate a space, or simply watch something beautifully.
The breakthrough is not a bigger feature set. It is a smaller gap between intention and action.
This is where the idea of a killer app becomes misleading. We tend to imagine a single application that suddenly dominates a category. But with new interfaces, the real killer often begins as a behavioral shortcut. It compresses a repeated human need into a new form factor until the form factor itself becomes the reason to use it.
Spatial video and 3D capture are promising not because they are flashy, but because they may solve a subtle problem: people do not merely want to record the world. They want to re-enter it later with context intact. That is a different behavior than photography, and perhaps a more durable one.
Why the first killer experience is usually a ritual, not a product category
Most people look for the next great app in the wrong place. They look for a software category label, like games, fitness, messaging, or productivity. But the earliest winners on a new platform often resemble rituals more than categories.
A ritual has three properties:
- It is easy to repeat.
- It gives immediate satisfaction.
- It becomes socially legible, meaning other people can recognize why it matters.
That is why some of the most successful new-device experiences feel almost embarrassingly simple at first. They are not trying to replace an entire workflow. They are finding the one moment where the device feels magical. A rhythm game on a VR headset is not just a game. It is a proof that your body can be part of the interface. Health tracking on a watch is not just a feature. It is a daily ritual of self-observation. Short-form cameras on a phone are not just cameras. They are memory rituals, social rituals, and identity rituals.
The same logic applies to spatial computing. If a headset is going to matter, it will likely start by making a few experiences feel uncannily better than the alternatives. Spatial video could be one of those experiences because it transforms memory from a flat artifact into a place you can return to. That is powerful because memory is not just storage. Memory is reconstruction. We revisit events by reassembling atmosphere, scale, and presence from fragments. A device that captures those qualities is not simply recording. It is preserving a version of lived experience that 2D media often flattens.
Think of the difference between a postcard and a room. A postcard tells you where you were. A room brings back how it felt to be there. That distinction may be the kernel of a true killer experience in spatial media.
The deepest prize: making the world legible in the same form it is lived
There is a hidden convergence between multimodal AI systems and spatial devices. Both are pursuing a world where input and output are not limited to text, icons, and taps. They are moving toward a mode of interaction in which the world itself is the interface.
This is a profound shift. Traditional software requires users to translate reality into representations: write a message, take a photo, file a ticket, enter coordinates, choose a dropdown. The translation burden sits on the human. Multimodal systems reduce that burden by accepting richer forms of intent, such as demonstrations, language, and visual goals. Spatial devices do something parallel on the output side by presenting content in the dimensions, scale, and perspective that better match human perception.
A useful mental model here is the friction ladder:
- At the top are systems that force you to abstract your intent into rigid commands.
- In the middle are systems that accept partial naturalness, like voice or drag and drop.
- At the bottom are systems that accept the same modality in which the world is experienced, like image, gesture, gaze, spatial context, and demonstration.
The lower the friction, the more likely a behavior becomes habitual. But there is a catch. Lower friction does not automatically create value. It can also create emptiness, where a device feels smooth but pointless. A technology becomes transformative only when reduced friction meets a real human desire that is repeated often enough to matter.
That is why the search for a killer app should be replaced by a search for a killer loop. A loop is a sequence of action and reward that becomes self-reinforcing. The loop does not need to be large at first. It needs to be frequent, emotionally satisfying, and easy to resume.
Examples:
- A quick spatial capture of a child’s birthday that later feels like stepping back into the room.
- A collaborative 3D annotation session that makes remote teamwork feel physical instead of abstract.
- A hands-free reference overlay that helps a technician complete a task without breaking flow.
- A robot that learns from a demonstration and then repeats the task reliably in a new environment.
These are not just features. They are loops that compress intent, action, and reward into one coherent motion.
Why general intelligence still needs a narrow wedge
There is a temptation to believe that once a system becomes more general, it can simply take over everything. The robot benchmark story pushes against that assumption. Even a highly capable agent benefits from structured prompts, systematic evaluation, and carefully chosen tasks. Generality does not remove the need for a wedge. It makes the wedge more important.
This matters for new hardware too. A headset may eventually support a vast range of experiences, but its success will not come from range alone. It will come from finding one or two experiences where the device is not just better, but categorically different. The difference must be visible to the user immediately. If a task only becomes 15 percent better, it may not overcome inertia. If it becomes 10 times more emotionally vivid, socially shareable, or physically intuitive, it can.
The hidden lesson from both robotics and spatial devices is that system-level ambition should be paired with task-level humility. Do not ask, “What can this platform do?” Ask, “Which repeated human behavior does this platform make unmistakably better?”
This is also why benchmarks matter. If you cannot specify the hard cases, you cannot know whether you are improving genuine capability or merely optimizing for demos. In consumer tech, the equivalent benchmark is not raw usage time. It is whether people return when the novelty wears off.
A good test for a candidate killer experience is simple:
- Does it solve a frequent need, not a rare one?
- Does it reduce effort while increasing emotional payoff?
- Can it be explained in one sentence to a non-expert?
- Does it become more valuable the second or third time you use it?
If the answer is yes, you may be looking at a behavior, not just a feature.
Key Takeaways
- Stop hunting for app categories. Look for recurring human behaviors that a new device can compress, amplify, or make newly natural.
- Treat the interface as a behavioral grammar. The best products teach users a new way to act, not just a new place to click.
- Search for killer loops, not killer features. The winning experience is usually a repeatable cycle of intent, action, and reward.
- Use friction as a diagnostic tool. If a workflow still feels awkward, the device has not yet found the right modality for the task.
- Ask whether the experience re-embeds the user in reality. The strongest spatial experiences preserve presence, context, and emotional texture, not just information.
The future belongs to interfaces that remember how humans actually live
The next great platform will not win by insisting that people adapt to it. It will win by noticing that humans already think, remember, coordinate, and act in multimodal ways, then making technology fluent in those modes.
That is the real connection between robots and spatial devices. Both are moving away from the assumption that intelligence is best expressed as text or commands. Both are learning that intelligence, for humans at least, is often demonstration, context, and situated action. A robot that understands a one-shot example and a headset that preserves a lived moment are solving the same civilizational problem from opposite ends: how to make machines fit into the texture of human life.
So the next time someone asks which app will define a new device, the better question is more uncomfortable and more interesting: What human ritual is this device finally making effortless enough to become a habit?
That is where killer apps come from. Not from categories. Not from feature lists. From the moment a new interface stops feeling like software and starts feeling like the world responding back.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣