AI Doctors on Trial: New Framework Exposes Hidden Dangers in Medical Language Models
A new proof-of-concept framework reveals that aggregate accuracy conceals four-fold differences in dangerous-recommendation rates and stochastic failures among large language ...
