The Productivity Trap: Why AI Must Give Doctors More Attention, Not Just More Time

Charles DeShazer

Hatched by Charles DeShazer

Sep 02, 2026

11 min read

92%

0

What if the most important productivity gain in medicine is not seeing more patients, but noticing what would otherwise go unsaid?

That question sounds almost heretical in an era fascinated by autonomous systems, intelligent software, and the prospect that generative AI could become a new productivity frontier. The usual promise is straightforward: machines will absorb routine work, accelerate decisions, and free professionals to focus on higher value activities.

But primary care exposes a difficulty that productivity models often miss. A physician’s work is not simply a sequence of tasks. It is a chain of judgments, interpretations, emotional signals, and small acts of attention. A message generated by an algorithm may reduce one burden while creating another. A compensation model may reward visible output while neglecting the invisible work that prevents future problems. One carefully chosen word can reveal an unmet concern that an entire workflow would otherwise overlook.

The deeper issue is this: technology does not create productivity merely by reducing labor. It creates productivity when it improves the quality of human attention.

That distinction changes how we should evaluate generative AI in medicine, and perhaps in every profession where the work involves people rather than objects.

The false equation: less time equals more productivity

Productivity is often treated as a simple equation: produce more while consuming fewer resources. In manufacturing, that logic can be remarkably powerful. If a machine produces twice as many identical components with the same inputs, the gain is easy to measure.

Human services are different. The output is partly hidden inside the interaction itself. A primary care visit may result in a diagnosis, a prescription, or a referral. Yet its most valuable outcome might be that a patient finally mentions a fear, understands a treatment, or feels safe enough to return. These outcomes are difficult to count, but they shape adherence, trust, and long term health.

Consider two appointments that take the same amount of time. In the first, the physician efficiently completes every required field, orders the appropriate tests, and closes the visit. In the second, the physician notices hesitation after asking a routine question. A slight change in wording invites the patient to explain that they are not worried about the symptom being discussed. They are worried about losing their job if the symptom becomes worse.

The second visit may appear less efficient. It may even generate more work. But it has discovered the problem that actually determines whether the care plan will succeed.

This is why administrative efficiency and clinical productivity are not identical. One concerns the speed of processing. The other concerns whether the system identifies and addresses the right problem.

Generative AI can improve the first while accidentally damaging the second. If it floods a physician with polished summaries, suggested responses, and algorithmically generated messages, it may reduce typing while increasing the amount of information that must be reviewed. If those messages are treated as completed work, the system may reward throughput even as physicians become less able to distinguish urgent signals from routine noise.

A faster river is not necessarily a better water supply. It depends on whether the water is clean, whether it reaches the right place, and whether anyone can tell the difference between a flood and a resource.

The hidden bottleneck is not labor. It is attention

The modern economy has spent decades automating physical effort and routine calculation. Smartphones then placed enormous computational capacity in the hands of individuals. The next wave of AI promises something more ambitious: systems that can interpret language, generate content, coordinate actions, and operate with increasing autonomy.

Yet as systems become more capable, the scarce resource may become human attention that remains capable of judgment.

A physician does not merely need information. A physician needs to know which information matters, what is missing, and when a seemingly ordinary statement carries unusual significance. The same is true of a teacher listening to a student, a manager conducting a performance conversation, or a lawyer questioning a client.

These professions depend on what might be called diagnostic attention: the ability to detect discrepancies between the official topic and the real concern. Diagnostic attention is not equivalent to concentration. Concentration focuses on a task. Diagnostic attention remains sensitive to what the task may be hiding.

Generative systems can assist this kind of attention, but only if they are designed as instruments of inquiry rather than substitutes for noticing. An AI system might flag that a patient’s current concern differs from the concern recorded at the previous visit. It might identify a pattern of missed appointments, confusing instructions, or repeated questions. Those functions could help a physician see more clearly.

But an AI system can also become a second patient. Its generated messages require supervision. Its recommendations demand verification. Its confident language can tempt users to accept a plausible answer before asking whether it addresses the real question. The system may reduce keystrokes while increasing verification load, the mental effort required to inspect, correct, and contextualize machine output.

This creates a crucial design test:

An intelligent tool is productive only when it returns more usable attention than it consumes in supervision.

The test is more demanding than asking whether the tool saves minutes. A system that saves five minutes of documentation but creates ten minutes of uncertainty has not improved productivity. It has merely moved work from visible labor to invisible vigilance.

Why one word can matter more than a large model

The importance of small conversational choices reveals a limitation in how organizations think about automation. Large systems tend to search for large interventions: new platforms, comprehensive dashboards, autonomous agents, and sweeping workflow redesigns. But the decisive moment in human work may be tiny.

A question such as “Anything else?” can function as a polite closing ritual. A variation such as “What else is on your mind?” communicates that the physician expects there may be another concern and is willing to hear it. The difference is only a few words, but the social meaning changes. One asks whether the checklist is complete. The other opens a door.

This is not an argument for romanticizing human intuition or rejecting automation. It is an argument for understanding where value is created. In many care interactions, value does not come from adding more information to the conversation. It comes from making it easier for the other person to disclose information they have been withholding, minimizing, or struggling to formulate.

Generative AI is especially relevant here because language is both its medium and its risk. It can produce countless versions of a message, but it does not automatically know which version creates trust in a particular relationship. It can summarize a patient’s words while erasing hesitation, ambiguity, or embarrassment. It can make communication smoother while making it less revealing.

The central distinction is between semantic completion and human discovery. Semantic completion gives an answer that appears to fit the words already present. Human discovery finds the issue that has not yet been stated clearly.

Primary care is full of the second kind of work. A patient says they forgot to take medication, but the underlying issue is cost. A patient says a treatment is not helping, but the real obstacle is fear of side effects. A patient agrees with the plan, but does not understand it. If technology optimizes only for what is explicit, it can make the record look more complete while leaving the person less understood.

The best AI therefore should not only generate answers. It should help professionals ask better questions, notice uncertainty, and preserve room for surprise.

The compensation problem: what gets measured becomes what gets practiced

There is another connection between technology and care that is easy to miss. Tools do not operate in a vacuum. Their effects depend on the incentives surrounding them.

If an organization rewards the number of encounters, messages closed, notes completed, or tasks processed, an AI system will naturally be used to increase those visible outputs. Its autonomy becomes attractive because autonomy appears to scale. Yet the system may then amplify the wrong objective. It can help a practice process more interactions without helping clinicians resolve more of the concerns that make those interactions necessary.

Compensation models matter because they translate institutional values into daily behavior. What an organization rewards tells professionals what counts as work. If only measurable throughput is valued, then listening becomes a luxury, follow up becomes a cost, and emotional labor disappears from the operational picture.

This is not merely a moral concern. It is an economic one. Unresolved concerns tend to return as repeat visits, escalations, nonadherence, avoidable testing, and deteriorating trust. The work that appears to slow down one encounter may reduce the total amount of work required later.

A narrow productivity metric counts the first interaction. A broader metric counts the entire care journey.

We can express the difference with a simple model:

Net productivity = useful output minus coordination cost, correction cost, and unresolved problem cost.

Generative AI may increase useful output by drafting notes or organizing information. But it can also increase correction cost if the output must be carefully checked. It can increase coordination cost if messages proliferate across clinicians and patients. Most importantly, it can increase unresolved problem cost if it creates the appearance of completion without actual understanding.

This model suggests that organizations should evaluate AI over time, not at the moment a task is closed. Did the patient understand the plan? Did the physician’s cognitive burden decrease? Did the number of repeated contacts fall? Did important concerns surface earlier? Did clinicians report greater capacity for judgment, or merely greater pressure to process machine generated work?

The answer to those questions should influence compensation and staffing. Otherwise, an organization may pay for speed while depending on unpaid attention to repair the consequences.

From autonomous replacement to augmented inquiry

The most useful mental model for AI in primary care is not “digital employee.” It is cognitive scaffolding.

Scaffolding does not replace the building. It gives workers access to places that would otherwise be difficult to reach. In clinical practice, this might mean helping a physician prepare for a visit, identify contradictions in a record, translate instructions into accessible language, or detect that a patient’s stated concern has changed over time.

A scaffold is valuable when it supports judgment without disguising its own limits. It should expose uncertainty rather than bury it under fluent prose. It should make the next question easier to see, not make questioning feel unnecessary.

This leads to a practical hierarchy for automation:

  1. Automate retrieval. Reduce the time required to find relevant facts.
  2. Automate formatting. Reduce repetitive documentation and administrative conversion.
  3. Assist interpretation. Surface patterns, contradictions, and missing information.
  4. Support inquiry. Suggest questions that could reveal the patient’s actual concern.
  5. Reserve relational judgment for people. Let humans decide how to respond when trust, fear, dignity, or ambiguity are central.

The order matters. Many organizations rush toward the fifth step in the hope of achieving autonomy before mastering the first four. But autonomy without reliable context is not intelligence. It is accelerated overconfidence.

A well designed system might draft a response to a patient message and also display why it made that suggestion, what evidence is missing, and what question could clarify the situation. A poorly designed system might simply produce a polished answer that closes the loop. The first expands clinical attention. The second narrows it.

The difference can be measured through a concept called the attention return on investment. For every unit of clinician attention spent using an AI tool, how much additional understanding, prevention, or care quality does the tool produce? This metric would favor systems that reveal hidden problems, not merely systems that generate more text.

What leaders and practitioners can do now

The transition to generative AI should therefore begin with a more precise question than “What can we automate?” Ask instead: Which parts of this workflow prevent professionals from noticing what matters?

That question produces better experiments. It points toward reducing inbox clutter, improving preparation, highlighting unresolved concerns, and making follow up more intelligent. It also prevents organizations from confusing a larger volume of machine output with a larger capacity for care.

Key Takeaways

  1. Measure completed problems, not completed tasks. Track whether concerns are resolved, instructions understood, and unnecessary repeat contacts reduced.

  2. Treat generated messages as drafts, not decisions. Preserve clinician review wherever the patient’s intent, risk, or emotional state is uncertain.

  3. Design AI to expose gaps. Ask systems to identify missing information, contradictions, and possible underlying concerns, not only to summarize what is already known.

  4. Protect conversational space. Train teams to use small, open prompts that invite patients to disclose what a checklist may miss. Technology should create more room for this practice, not eliminate it.

  5. Align incentives with long term value. Compensation and performance measures should recognize prevention, continuity, reduced confusion, and the relational work that makes treatment effective.

The real productivity frontier

The popular image of the AI future is a machine that does more independently. That image is incomplete. In human services, the more important question is whether people can do more of what only people can do: interpret context, recognize vulnerability, earn trust, and discover the problem beneath the problem.

A smartphone made computing portable. Generative AI may make language production abundant. But abundance creates its own scarcity. When text, recommendations, alerts, and responses can be generated almost without limit, the valuable act becomes deciding what deserves attention.

Primary care makes this visible because its central task is not simply to process symptoms. It is to understand people whose symptoms are entangled with habits, fears, resources, relationships, and incomplete explanations. The technology that succeeds will not be the technology that makes the human presence unnecessary. It will be the technology that makes that presence more perceptive.

The future of productivity is therefore not a race toward the least human workflow. It is a race toward the workflow in which human attention is spent where it has the highest consequence.

The best AI assistant may not be the one that lets a doctor finish the visit fastest. It may be the one that quietly helps the doctor realize there is one more question worth asking.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣