
CMU Machine Learning Blog · 实时热榜
HISTORY2026年8月3日1 不同热搜
07/24—08/22 有历史数据
- 01Healthcare Benchmarks Are Only as Good as Their Assumptions
In healthcare settings where patients use LLMs as a medical assistant, LLM performance differs between evaluation and deployment. (a) Bean et al. (2025) find a 61 percentage point difference between evaluation and deployment. (b) We argue this gap arises not from poorly designed benchmarks, but from implicit assumptions embedded in evaluation protocols that fail to hold at deployment. (c) We propose a taxonomy that categorizes assumptions into two types, task and outcome, to diagnose where the g
最高第 1 名00:00 达到当日首次采集时已在榜当日结束时仍在榜累计约20小时37分

































































































