News: AI tool quicker, not more accurate than humans when diagnosing rheum disease, study finds
Prof. Valmed, a large language model (LLM) cleared for medical diagnosis, outperformed human physicians for diagnosing rheumatologic conditions in terms of speed, but not accuracy, a trial found.
The trial evaluated Prof. Valmed’s performance in the rheumatology field through clinical scenarios. The primary outcome was accuracy, defined as the percentage of instances in which a participant's most likely diagnosis matched the actual published one. This was achieved in 33.3% of intervention group cases versus 35.0% of those approached conventionally, which was not statistically significant.
Working on its own, the system was also 33.3% accurate for its top likelihood matching the true diagnosis, and 47.6% accurate in having one of its three candidates be correct.
Diagnostic accuracy was almost the same in rheumatology case scenarios for physicians who used Prof. Valmed versus those reaching their diagnoses in their usual ways, according to researchers. However, human doctors relying on their own resources required an average of 206 seconds to make their diagnoses, compared with 94 seconds among those aided by Prof. Valmed.
Prof. Valmed increased physicians’ confidence in their diagnoses. In the intervention group, even before Prof. Valmed was called in, the mean confidence rating was 39%, compared with accuracy of 22%. Their 33% accuracy when using the artificial intelligence (AI) system came with confidence ratings averaging 57%. The same pattern was seen in the control group.
Analyses suggested substantial AI over-reliance in the intervention group, and under-reliance was uncommon. This suggests that diagnostic support may be able to increase confidence more readily than correctness and highlight the need to evaluate calibration and behavioral reliance alongside accuracy.
AI system developers have targeted medicine as one of the most important potential applications for emerging technology. Whether the field has reached a point that the tools are doing tasks quicker and more accurately, however, remains uncertain. General AI systems have been tried with mixed results, sometimes finding diagnoses that had eluded human professionals. However, they are also prone to “hallucinations,” or false conclusions that appear to stem from the systems’ emphasis on pleasing their users.
Prof. Valmed is designed specifically for medical applications with guardrails to limit these hallucinations. As an LLM, it can be directed through ordinary language and can be used to assist, not replace, physicians. Participants gave Prof. Valmed a high rating for ease of use and pleasing interface, with two-thirds saying it was trustworthy and 80% saying they would use it again.
Editor’s note: To read the full study, click here. To read additional coverage from MedPage Today, click here.
