When Machines and Patients Decide If You Are a Doctor: Rethinking Clinical Competence in an Age of Simulation and AI
Hatched by SEAN SYLVIA
Apr 16, 2026
9 min read
6 views
82%
What if a machine, watching a transcript or a simulation, could decide whether you are safe to treat patients? What if patients, not examiners in white coats, became the primary arbiters of clinical skill? We are fast moving into a world where the old ritual of an in-person clinical skills exam is gone, and the authority that once sat in the testing center is splintering into simulations, algorithms, and patient voices. The stakes are not academic. They are about trust, safety, and how we teach the next generation of clinicians.
The setup: A vanished ritual and a reassembled verdict
For decades one of the clearest public rituals in medical education was the clinical skills exam: students in neatly ironed coats, standardized patients performing the same symptoms in a small exam room, examiners checking boxes while a clock ran. That ritual signaled something simple and reassuring: someone had watched you do the thing that defines being a clinician. It was tangible. It felt real.
That ritual is no longer universal. The in-person, high-stakes clinical skills test has been dismantled. In its place are several overlapping trends: computer-based case simulations, bolstered communication tasks embedded in knowledge exams, telehealth encounters, video-recorded or workplace-based assessments, and rapidly improving AI systems that can parse words and behavior for markers of competence.
This transition is not merely logistical. It recasts the question: how do we know someone can care for a real patient? The answer used to depend largely on direct observation in a standardized context. Now the answer is diffuse, populated by synthetic scenarios, algorithmic outputs, and crucially, the patient experience itself.
The tension: Measurement, meaning, and the missing patient gaze
Measurement technologies and educational priorities are sliding in different directions. On one hand, simulation and digital cases scale easily, are easier to standardize, and produce tidy metrics. On the other hand, these formats can be impoverished proxies for real care. They often capture performance under test conditions, which can be theatrical, rather than the messy, relational work of real clinical practice.
There are three core tensions that emerge from this shift:
-
Validity versus practicality. Simulations are efficient and reproducible. But clinical competence includes elements that are not easily scripted: improvisation, managing uncertainty, integrating complex social context, and the small gestures that patients call empathy. Can a checkbox or a simulated scenario capture that?
-
Observation versus experience. Tests observe behavior at a moment in time. Patients live through encounters and form judgments about care that only accrues over time. If we value patient outcomes and experience, we need assessments that can sample those longitudinal, contextual signals.
-
Expertise versus authority. Historically, credentialing authorities set standards. Today, authority is contested by new voices: patients advocating for technologies that reflect their priorities, and AI systems that claim to measure communication, clinical reasoning, or risk. Who gets to define what matters, and how will power shift?
These tensions matter because measurement drives behavior. When assessment emphasizes checklists and simulated signals, training systems will incentivize performance that aligns with those measures. That can lead to proficiency on the test and blind spots in the clinic.
Reframing clinical competence: A new model for a distributed assessment ecosystem
To move beyond the old either or, we need a conceptual framework that integrates three distinct epistemic sources of evidence: simulated performance, workplace observation, and patient experience. I call this the Triangle of Clinical Competence.
Triangle of Clinical Competence
- Knowledge and simulated performance: This corner includes knowledge exams, computer-based cases, and high-fidelity simulations. These are strong at assessing clinical reasoning, pattern recognition, and procedural steps.
- Observed practice: This corner includes direct observation in real clinics, video review of encounters, multisource feedback from colleagues, and workplace-based assessments like entrustable professional activities. These are strong at assessing adaptation, teamwork, and safety behaviors.
- Patient experience and outcomes: This corner includes patient-reported experience measures, longitudinal outcome data, and narrative patient feedback. These capture relational competence, communication in context, and the lived effects of care.
Competence sits in the center when evidence from all three corners aligns. Critically, no single corner is sufficient. A high score in simulation paired with poor patient feedback indicates a dangerous mismatch; likewise, patients may praise a charming clinician whose diagnostic reasoning is weak.
This triangle gives us a mental model for designing assessments in which AI and patients are not intruders, but central contributors to valid judgments.
How technology and patients can be combined constructively
The anxiety around AI and simulation is real: algorithms can be brittle, reproduce bias, and reward surface-level cues. But technology also offers tools that human observers cannot scale. Patient involvement offers perspectives that no simulation replicates. The challenge is to design hybrids that capture the strengths of both while mitigating the weaknesses.
Here are practical design principles to guide that synthesis.
-
Measure processes and outcomes in parallel. Use simulations and AI to measure specific processes such as differential diagnosis, safety checks, and communication structure. Simultaneously, collect patient experience and outcome data to assess whether those processes translate into meaningful benefit. Treat discordance between the two as the most informative signal.
-
Use AI as a measurement amplifier, not a final judge. Algorithms are excellent at parsing large-scale patterns in audio, video, or textual data. They can highlight potential lapses, quantify talk-time balance, tag empathic phrases, or detect missed safety steps. But human review, especially by patients or clinicians, should adjudicate high-stakes decisions.
-
Make the patient voice explicit and structured. Patients can provide narrative feedback, but for assessment to be fair and actionable, organize patient input into structured instruments and narrative prompts that map back to competencies. For example, ask patients to rate clarity of explanation, feeling of being heard, and confidence in the plan, and couple these with short narrative answers about what mattered.
-
Sample longitudinally and contextually. Single encounters are noisy. Build portfolios that accumulate evidence across settings and time. Video clips, patient reports, AI-flagged anomalies, and supervisor notes together create a richer picture than any isolated exam.
-
Audit for equity and bias continuously. Algorithms trained on transcripts or video can encode demographic biases. Patients from marginalized groups may experience and report care differently. Regularly audit assessment outputs against demographic data, and use mixed-methods investigations when disparities arise.
-
Design for formative plus summative use. Most educational assessments are more powerful when they inform learning as well as gatekeeping. Provide learners with rich, interpretable feedback that links simulated performance, workplace behavior, and patient-reported signals to specific growth goals.
Concrete example: the hybrid assessment portfolio
Imagine a senior medical student whose portfolio contains the following items: a timed computer case sim showing diagnostic choices; a recorded telemedicine visit analyzed by AI for empathic phrases and talk time; three patient experience surveys tied to different clinical episodes; a workplace-based entrustable activity observation; and a reflective narrative linking the pieces. An assessment committee reviews this portfolio. AI highlights a pattern of low open-ended questioning and a mismatch between simulation teams and patient narratives. The committee then uses targeted coaching, a focused clinical immersion, and follow-up patient feedback to confirm improvement.
That workflow uses simulation and AI to reveal patterns at scale, while centering the patient gaze to judge meaning and impact.
Risks, ethical guardrails, and power considerations
Technological and assessment redesigns carry real risks if power is not redistributed thoughtfully.
-
Performance theater: When assessments focus on specific checklists, learners may perform for the test rather than internalize habits that matter to patients. This is a known problem in simulated exams, and it can be magnified when signals are easy to game.
-
Overreliance on proxies: Shortcuts like AI-derived empathy scores are attractive, but they are proxies. Overreliance can institutionalize measures that do not match patient priorities.
-
Privacy and consent: Recording clinical encounters, feeding them to AI, and using patient narratives in high-stakes decisions require informed consent frameworks that respect confidentiality and the power imbalance between patients and institutions.
-
Unequal voice: If patient input is nominal, tokenized, or only solicited from certain populations, assessments can reproduce inequities. Patients must be genuine partners in defining what matters.
Ethical guardrails to mitigate these risks include transparent algorithmic governance, patient co-design of assessment instruments, robust data governance plans, and iterative validity studies that examine real-world outcomes after assessment changes.
Design assessment ecosystems so that no single instrument decides destiny: technology should illuminate, patients should contextualize, and humans should adjudicate.
Key Takeaways
-
Center the patient gaze: incorporate structured patient-reported experience and outcome measures into assessment portfolios, not as adjuncts, but as core evidence of competence.
-
Use AI to surface patterns, not to finalize judgments: apply algorithms to flag concerns at scale, but require human-patient adjudication for high-stakes decisions.
-
Build triangular evidence portfolios: integrate simulation, workplace observation, and patient experience to make balanced competence decisions.
-
Sample across time and context: prioritize longitudinal, multi-source evidence rather than single high-stakes snapshots.
-
Guard against theater and bias: continuously audit measures for fairness, and design assessments to reward genuine clinical habits over performative signals.
A final reframing: assessment as a shared social contract
The elimination of the in-person clinical skills ritual is a symptom of a broader shift. Authority over what it means to be a competent clinician is moving from centralized testing centers into a distributed ecosystem populated by simulations, algorithms, workplaces, and patients. That shift is neither inherently good nor bad. It is an opportunity.
If we accept the old test as the only legitimate way to decide clinical safety, we risk clinging to a theater that hides long-term failures. If instead we allow technology and patient voices to set the agenda without careful design, we risk building brittle systems that reward style over substance.
The healthier path is deliberately pluralistic. Treat assessment as a social contract co-created by educators, clinicians, patients, and technologists. Make the terms explicit: what counts as evidence, how that evidence will be collected, how privacy will be protected, and how remedial paths will work when gaps are found.
In practice this means investing in patient-centered measurement tools, transparent AI governance, longitudinal portfolios, and coaching-oriented remediation. It means admitting uncertainty and using disagreement between measurement corners as the most productive signal for investigation and improvement.
The real test for our generation of medical educators is not whether we can build more predictive algorithms or fancier simulations. It is whether we can build an assessment ecosystem that respects complexity, centers lived patient experience, and produces clinicians who are not merely simulacra of competence but are trusted in the messy, everyday work of healing.
The question is not who sits in the exam room. The question is who gets to say you cared well enough that day. Make that answer public, distributed, and accountable, and you will have a defensible way to decide who is ready to carry a patients life in their hands.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣