false
OasisLMS
Login
Catalog
Benchmarking AI in Radiology: Datasets, Clinical E ...
WEB05-2026
WEB05-2026
Back to course
[Please upgrade your browser to play this video content]
Video Transcription
Video Summary
This RSNA webinar focused on benchmarking AI in radiology and why it is becoming essential in the era of foundation models and large language models. George Shea argued that benchmarking may be one of the most important activities radiologists can do to ensure AI tools are evaluated fairly and clinically meaningfully. He emphasized that benchmarks should use new, public, transparent, and reproducible datasets that reflect real clinical complexity, rare cases, and changing practice conditions. He also described RSNA’s benchmarking workgroup and its early chest X-ray benchmark, with future efforts planned for neuro, abdomen, and MSK.<br /><br />Maggie Chung expanded on the need to shift from a model-centered to a community-centered benchmarking approach. She explained that current AI evaluations are hard to compare because they use different datasets, reference standards, and metrics. She outlined how the field should agree on priority findings, representative datasets, task definitions, reference standards, and metrics, while preserving uncertainty and capturing subgroup performance. She also highlighted special challenges for foundation models, including multimodal inputs, report generation, and better ways to evaluate outputs beyond word overlap.<br /><br />The webinar concluded with audience Q&A on provenance, bias, benchmark versioning, difficulty scoring, and governance.
Keywords
AI benchmarking
radiology
foundation models
large language models
clinical evaluation
public datasets
benchmarking workgroup
chest X-ray
community-centered benchmarking
×
Please select your language
1
English