PV Intake Assistant
A change to my own system raised its coverage from 52% to 80%. I threw it away.
Reading the output showed why: it had coded a skin lesion as “Feeling tense” and a nerve toxicity as “Poisoning”, matching on one incidental shared word. I reverted to 54% — every code left standing is one I would defend in a review.
Three stages — triage (F1 0.79, recall 0.86), drug and effect extraction, then mapping to SNOMED CT across 461,000 terms. 32 unit tests, CI, Docker.
So whatA wrong code and a missing code are not the same error. A gap gets caught by a human; a wrong one enters the database looking exactly like evidence. Any metric that scores them equally will recommend the wrong system.