COLM 2026

MedConcealA Benchmark for Clinical Hidden-Concern
Reasoning Under Partial Observability

Yikun Han1Joey Chan1Jingyuan Chen1Mengting Ai1Simo Du2Yue Guo1

1 University of Illinois Urbana-Champaign2 NYC Health + Hospitals/Jacobi

Can AI uncover what patients leave unsaid?

Clinical dialogue goes beyond medical knowledge. It requires discovering the fears,
beliefs, and practical barriers a patient has not yet shared.

01
Partial observability

Hidden

The chart tells only part of the story.

02
Confirmation

Revealed

Skillful questions surface the concern.

03
Intervention

Addressed

The response engages with the barrier.

300curated patient cases
159clinician participants
5AI clinician baselines
4hidden-concern categories

Reasoning beyond the patient chart.

Research summary

Patient-clinician communication is an asymmetric-information problem: patients often do not disclose fears, misconceptions, or practical barriers unless clinicians elicit them skillfully. MedConceal evaluates this challenge through an interactive patient simulator that separates clinician-visible context from simulator-internal hidden concerns.

Built from clinician-answered online health discussions, the benchmark comprises 300 curated cases. A reserved, stateful patient simulator tracks whether concerns have been revealed and addressed using theory-grounded, turn-level communication signals. Two tasks test complementary abilities: confirmation, surfacing hidden concerns through dialogue, and intervention, addressing the primary concern to support a target care plan.

Comparisons with 159 human clinician participants show that no single AI system leads across all metrics. Longer dialogues help some models, but effective interaction still depends on eliciting the right concern and responding to it. Read the full abstract ↗

Same visible context.
Hidden concerns to discover.

Human and AI clinicians see the same patient chart. The simulator keeps psychosocial concerns private until the interaction elicits them.

MedConceal framework: clinician-answered AskDocs threads become structured cases with private concerns; human and AI clinicians interact with a stateful patient simulator in confirmation and intervention tasks.
Figure 1 from the paper. Case construction, controlled disclosure, and matched-information human–AI evaluation. Select the figure to view it at full size.
Task 01

Confirmation

Surface hidden concerns through multi-turn dialogue, then submit structured findings. Evaluation separates what was actually revealed from what was inferred afterward.

Reveal Rate · Fine-grained F1 · Matched-but-no-reveal

Task 02

Intervention

Address the primary hidden concern and guide the patient toward a target plan. Success requires the concern to reach the simulator’s addressed state.

Success · Reveal Rate · Turn-to-Address

Four kinds of hidden concern

  • 01Misinformation or misconceptions
  • 02Emotional discomfort or fear
  • 03Communication barriers
  • 04Financial or insurance-related concern

The right plan starts
with the hidden barrier.

In a case discussed in the paper, a knee injury is only part of the story. Cost and travel constraints change what a useful next step looks like.

Paraphrased case from Section 4.3.

Case walkthrough1 / 3
What the clinician sees

A patient with a knee injury.

The clinical presentation suggests a need for further evaluation. But the visible problem does not explain what could prevent the patient from accessing care.

A recommendation alone may miss the barrier.

More dialogue helps.
What happens in it matters.

Explore the reported results by task and turn budget. Human conversations averaged 8.2 turns; the extended AI setting is a separate condition.

42.7%

Human intervention success

Compared with 29.3% for the strongest 8-turn AI baseline, Claude Sonnet 4.5.

55.8%

Human confirmation reveal rate

Higher than any 8-turn AI baseline; Claude Sonnet 4.5 reaches 52.2%.

11.164 vs 7.125

Turns to address the concern

Doctor-R1 at 20 turns ties human success, with a higher reported Turn-to-Address.

Intervention performance · AI 8-turn setting
SystemSuccess ↑Reveal Rate ↑Turn-to-Address ↓
Human clinicians Reference42.7%48.7%7.125
Claude Sonnet 4.529.3%31.3%5.239
Doctor-R115.0%34.7%6.938
GPT-5.27.7%8.7%5.043
Qwen-3.5-9B3.0%6.3%6.333
Llama3-OpenBioLLM-8B1.3%1.3%5.500

Success means the primary concern reached the simulator’s addressed state. Turn-to-Address records the first turn at which the primary concern becomes addressed. Human results are an observed reference, not an 8-turn-capped condition.

Reported values from Table 3.

Cumulative performance by dialogue turn: confirmation reveal rates and intervention success differ across AI models, compared against the final human reference and the mean human conversation length.
Figure 2 from the paper. Dashed horizontal lines show final human reference performance; vertical dotted lines mark mean human conversation length.

A controlled testbed
for a difficult interaction.

MedConceal evaluates hidden-concern reasoning under a specific simulator and interaction protocol. It does not establish effectiveness in clinical deployment.

Concerns are reconstructed from patient-authored online text, rather than prospectively confirmed by patients. Intervention uses a single source-derived target plan, and outcomes depend on simulator state transitions. Prospective validation remains necessary.

Limitations and future directions ↗

Citation

@inproceedings{han2026medconceal,
  title = {MedConceal: A Benchmark for Clinical Hidden-Concern
           Reasoning Under Partial Observability},
  author = {Han, Yikun and Chan, Joey and Chen, Jingyuan and
            Ai, Mengting and Du, Simo and Guo, Yue},
  booktitle = {Conference on Language Modeling},
  year = {2026}
}

Download .bib file ↓