The Hidden Skill Behind Great AI Work: Giving Machines a Point of View
Hatched by john ke
Jul 16, 2026
11 min read
2 views
86%
The real bottleneck is not intelligence, it is direction
What if the biggest leap in AI productivity is not getting a smarter model, but giving the model a point of view?
That sounds almost too simple, especially when the hype around AI usually centers on scale, speed, or raw capability. Yet in practice, the difference between mediocre output and astonishing output often comes down to something far less glamorous: whether the AI knows how to behave, what role to assume, and what kind of result it should optimize for. A generic model is like a brilliant camera in the hands of someone who never chooses a lens. It can capture reality, but it cannot decide what reality should feel like.
That is why two seemingly different ideas belong in the same conversation. One is the ability to create a custom mode in an AI coding tool, where you define a mode name and preferences. The other is a catalog of camera work techniques for AI video generation, where movement, framing, and perspective dramatically change the emotional and narrative effect. On the surface, one is about software workflow and the other about visual creativity. At a deeper level, both are about the same thing: constraining a generative system so it can express intention.
The deeper lesson is this: AI does not just need instructions. It needs a stance.
Why generic AI feels powerful but still disappoints
Most people start with AI as if it were an infinite utility knife. They type a prompt, get an answer, and assume the issue is quality of prompt writing. But the more serious limitation is not prompt wording. It is that a default AI has no persistent identity, no stable style, and no preferred operating mode. Each interaction begins from near zero, which means the burden of context is always on you.
That is why custom modes matter so much. A custom mode is not just a convenience feature. It is a way of turning repeated judgment calls into an enduring decision framework. Instead of telling the AI, again and again, to be concise, careful, rigorous, optimistic, or implementation focused, you encode those preferences once and let them shape the interaction every time. In effect, you are creating a working persona, a reusable cognitive environment.
This is more important than it sounds. Human experts do the same thing internally. A great editor does not approach every sentence from scratch. A great architect does not evaluate each beam as if it were an isolated object. They operate from a stable set of priors: what matters, what to ignore, what standard of quality to enforce, what tradeoff is acceptable. Custom modes externalize that expertise.
The highest leverage in AI is not better prompting. It is better defaults.
That principle becomes even clearer when you look at AI video generation. When users want better outputs, they do not merely ask for “a better video.” They specify camera movement: dolly in, pan left, handheld shake, crane shot, over the shoulder framing, low angle, slow zoom, orbiting camera, and so on. Why? Because visual meaning is not contained only in subject matter. It is encoded in how the camera chooses to see.
The same actor, same scene, same lighting, same costume can feel tender, ominous, majestic, or anxious depending on the camera language. In other words, generation is not just about content. It is about perspective design.
Camera work is the visual version of a custom mode
A camera shot is more than a technical choice. It is a narrative decision. A close up says, pay attention to feeling. A wide shot says, pay attention to context. A handheld shot says, this is unstable and human. A smooth tracking shot says, follow the motion and trust the journey. Each choice turns the same raw scene into a different interpretation.
That is exactly what a custom mode does in a coding environment. It tells the AI, in advance, what kind of intellectual camera to use. Should it zoom in on edge cases or stay wide and strategic? Should it be a cautious reviewer or an optimistic builder? Should it privilege speed, safety, maintainability, or elegance? The mode is not the answer. It is the frame through which answers become coherent.
This analogy matters because it reveals a hidden truth about generative tools: output quality is often a framing problem disguised as a capability problem.
Think about how often AI work goes wrong. The model writes code that is technically correct but architecturally messy. It generates prose that is fluent but emotionally flat. It produces video that is visually plausible but strangely lifeless. In many cases, the model is not failing to create. It is failing to choose the right lens. Without that lens, the system averages across possibilities instead of committing to a specific creative intention.
A useful mental model here is to imagine three layers of control:
- Content layer: what is being made.
- Behavior layer: how the system should act while making it.
- Perspective layer: what should be foregrounded, what should be suppressed, and what kind of meaning should emerge.
Most users only work at the content layer. Advanced users operate at the behavior layer. Experts build at the perspective layer.
The camera work list is, in effect, a vocabulary for perspective. Custom modes are a vocabulary for behavior. Together, they show that the future of AI is not simply “ask and receive.” It is design the conditions under which the model thinks and sees.
The deeper shift: from prompting outputs to shaping operating systems
The most important change here is philosophical. We are moving from treating AI as a one off answer machine to treating it as an operating system for judgment.
That shift changes what skill means. In the old paradigm, skill was mostly about producing the right artifact. In the new one, skill is about shaping the conditions that produce repeatable artifacts. That means setting modes, defining roles, selecting camera language, and building reusable patterns that preserve intention across sessions.
This is why custom modes are more than personalization. They are the beginning of institutional memory inside a tool. A team can encode product standards, code review philosophy, brand voice, or research rigor into a mode rather than relying on whoever happens to be prompting that day. Likewise, a video creator can build a library of camera patterns that consistently generate a desired emotional rhythm. Over time, these choices become part of the creative infrastructure.
The shift also explains why so many people feel that AI is simultaneously magical and unreliable. They experience the raw power of generation, but not the discipline of direction. It is like having a world class orchestra with no conductor. Every section can play beautifully, yet the result may still feel unfocused. The conductor is not there to create notes. The conductor is there to create coherence.
In AI, custom modes are conductors. Camera work is conducting for the eye. Both reduce entropy.
Great AI output is not the product of more freedom. It is the product of better constraints.
This can sound counterintuitive because constraints are usually treated as limits. But in creative systems, constraints are often what make expression legible. A sonnet is constrained, and that is why it can be powerful. A film shot is constrained by lens choice, and that is why it can communicate a feeling instantly. A custom mode is constrained by preference, and that is why it can make repeated work feel consistent and intelligent.
Freedom without structure creates noise. Structure without flexibility creates sterility. The sweet spot is a constraint that protects intention.
A practical framework: intent, lens, and motion
If you want a simple way to think about all this, use the intent, lens, motion framework.
1. Intent: What is the AI trying to achieve?
Before asking for output, define the job. Is the goal to explore, optimize, explain, persuade, critique, or generate variations? A custom mode is useful precisely because it can encode intent repeatedly. If you always want the model to think like a careful pair programmer, say so once and preserve it.
2. Lens: How should the system interpret the task?
This is the equivalent of choosing a camera angle. Should the model think at a high level or zoom into details? Should it prioritize business impact, user empathy, technical correctness, or visual drama? The lens determines what becomes salient.
For AI video generation, this might mean choosing between a wide establishing shot and a tight close up. In coding, it might mean whether the AI behaves like a strict architect, a rapid prototyper, or a defensive reviewer. Same underlying model, different interpretive posture.
3. Motion: What kind of change should the output express?
Motion is the dynamic part. In video, this is the camera move. In coding or writing, it is the progression of the argument, the stepwise reveal, the transition from confusion to clarity, or from draft to production quality. Motion gives the output momentum.
The beauty of this framework is that it applies across mediums. A good custom mode sets intent and lens. A good camera move creates motion. When those are aligned, the output starts to feel less like a machine answer and more like a directed work of craft.
Consider a concrete example. Suppose you are using AI to create a demo video for a new product feature. If you simply say, “make a product video,” you will get something generic. If instead you define a mode that emphasizes clarity, then choose camera movement that starts with a wide contextual shot, moves into a slow push toward the key interaction, and finishes with a focused close up, you have created a communicative arc. The viewer does not just see the feature. They understand why it matters.
The same applies to coding. If your custom mode says: prioritize readability, explain tradeoffs, avoid unnecessary abstraction, and propose tests, the AI stops behaving like a text generator and starts behaving like a teammate. The output becomes less random because the model is being asked to inhabit a role.
The real opportunity: teaching AI to think in genres
The most powerful use of these techniques may be neither coding nor video creation alone. It may be the ability to teach AI to operate in genres.
Genres are not just categories. They are bundled expectations about tone, structure, pacing, and purpose. A thriller, a lecture, a product demo, a code review, and a meditation guide all have different rules for what counts as good. Human creators know this instinctively. AI users often do not, which is why so much output feels technically fine but spiritually wrong.
Custom modes let you define a genre for the machine. Camera work lets you reinforce that genre visually. Together, they reduce the gap between generic generation and intentional expression. They move the system from “produce something” to “produce something in the right register.”
This is where the opportunity becomes strategic. Teams that learn to encode genres into AI workflows will stop treating each output as a one off craft problem. They will build repeatable creative systems. The result is not bland standardization. Quite the opposite. It is the ability to preserve a consistent identity while exploring many variations within it.
That is the paradox: the more carefully you define the mode, the more creative the variation becomes.
A filmmaker who understands camera grammar can improvise more boldly because the grammar gives the improvisation shape. A developer who uses a strong AI mode can move faster because fewer decisions need to be renegotiated every time. In both cases, the frame does not restrict imagination. It enables it.
Key Takeaways
-
Stop treating AI as a blank slate. The best results come from persistent settings, roles, and preferences that shape behavior before the first prompt.
-
Think in terms of perspective, not just content. In video, camera movement changes meaning. In coding and writing, custom modes change judgment.
-
Use the intent, lens, motion framework. Define what the AI should achieve, how it should interpret the task, and what kind of dynamic progression the output should have.
-
Build reusable modes for repetitive work. Encode standards for code review, brainstorming, summarization, product thinking, or creative direction so quality does not depend on memory.
-
Choose constraints that protect intention. Good constraints do not reduce creativity. They make it easier for the system to express a clear point of view.
The future belongs to people who can direct, not just ask
We are still early in the age of generative tools, which means many people are optimizing the wrong thing. They are asking how to get better outputs from AI, when the deeper question is how to shape the identity of the output process itself.
That is why custom modes and camera work belong in the same mental drawer. Both are methods for turning generative power into directed expression. Both say that the true skill is not commanding the machine to do more, but helping it see more precisely. And both reveal a larger truth about working with AI: the system becomes genuinely useful when you stop thinking like a requester and start thinking like a director.
The future of AI fluency may not be defined by who can write the cleverest prompt. It may be defined by who can build the best lens.
And once you understand that, you begin to see every custom mode, every camera choice, every preference setting for what it really is: not a small tweak, but a way of teaching intelligence what to pay attention to.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣