MedReason
In development

Patient encounters for medical AI agents

MedReason builds interactive clinical cases for training and evaluating AI agents. Each case plays out like a real workup, and every step is scored.

Custom cases on request

One case, start to finish

34-year-old man with 3 days of fever and a new heart murmur. He injects drugs.

Any recent dental work? Rash? Joint pain?

No dental work, no rash, no joint pain.

Order blood cultures ×3low cost

Order an echocardiogramhigh cost

S. aureus in 3 of 3 bottles. Vegetation on the tricuspid valve.

Diagnose: acute S. aureus endocarditis, tricuspid valve

ScoreCorrect · moderate cost

Cases are built from peer-reviewed data. Every decision is logged with the agent's reasoning, and tests the patient didn't need count against it.

Where it's going

Most medical AI benchmarks hand a model one case and one question. The plan is to build many different scenarios. Maybe the agent sees five patients and then has to do five workups after. Maybe it's running late and scrambling through IHC. Physician-scientists bring a whole different set of tasks: grant writing, wet lab, analysis.

That's before we get to hypotheses. These are just the low-hanging fruit for writing a grader.

Status

This site is a placeholder while the first case set is built privately. If you train or evaluate medical AI agents and want a sample case, or one built for your use case, get in touch.

jacobhhastings@medreason.live