Ying, Lance, Wu, Jinzhou, Wang, Yingshan Susan, Aarya, Shivam, Buschoff, Luca M. Schulze, Chen, Harry, Collins, Katherine M., de Varda, Andrea, Fu, Shuhao, Houlihan, Sean Dae, Jagadish, Akshay K., Jiang, Guangyuan, Kiegeland, Samuel, Kurumisawa, Tetsu, Liu, Rongzhi, Liu, Ryan, Ma, Ningshan, McGregor, Kathryn, Strittmatter, Younes, Tsvilodub, Polina, Vigly, Jacob Hoover, Wu, Sarah, Xu, Enjie, Yun, Yiling, Allen, Kelsey, Brooke-Wilson, Tyler, Christian, Brian, Fedorenko, Evelina, Frank, Michael C., Franke, Michael, Gao, Tao, Gershman, Samuel J., Hawkins, Robert D., Hu, Jennifer, Jara-Ettinger, Julian, Kleiman-Weiner, Max, Levine, Sydney, Linzen, Tal, Lu, Hongjing, O'Donnell, Timothy, Ong, Desmond C., Piantadosi, Steven T., Saxe, Rebecca, Schulz, Eric, Shu, Tianmin, Sosa, Felix A., Sucholutsky, Ilia, Zhi-Xuan, Tan, Ullman, Tomer, Xu, Fei, Yildirim, Ilker, Zhu, Jian-Qiao, Griffiths, Thomas L., Gerstenberg, Tobias, Smith, Kevin, Tenenbaum, Joshua B.
Abstract
Understanding and modeling human intelligence are parallel goals shared by artificial intelligence (AI) and cognitive science. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models. We introduce CogGym, a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials. CogGym uses a semi-automated, human-in-the-loop pipeline to standardize diverse experimental paradigms into a task-agnostic Experiment Markup Language (EML), enabling reproducible and faithful comparison at scale. For initial release, we curate and standardize 258 cognitive experiments from 100 papers that focuses on human commonsense reasoning, and evaluate 50 large language models against human responses. We find a clear scaling trend where larger and more recent AI models better reproduce human judgments. Yet AI models' improvement on such common reasoning tasks is considerably slower than the gains observed on formal-reasoning benchmarks like math and coding, and model--human fit remains well below human splithalf reliability ($R^2 = 0.93$ on text, $0.95$ on image, and $0.92$ on video) with the best models achieving $R^2 = 0.59$ on text, $0.58$ on image, and $0.43$ on video experiments. We intend for CogGym to provide a living evaluation framework that continually incorporates new cognitive science experiments to characterize where model behavior resembles human behavior, where it systematically diverges, and how those patterns change as models and experiments evolve.
Chinese Translation
理解与建模人类智能是人工智能(AI)与认知科学共同追求的两个并行目标。随着AI系统能力的不断提升,模型响应在哪些方面与人类响应相似,又在哪些方面存在系统性差异?人类能够执行和思考的任务范围之广、种类之多,给人类与模型之间可扩展且严谨的比较带来了挑战。我们提出CogGym,这是一个以认知科学为基础的、可扩展的统一框架,用于在匹配的实验试次上系统性地比较模型与人类行为。CogGym采用半自动化的人机协同(human-in-the-loop)流程,将多样化的实验范式标准化为一种与具体任务无关的实验标记语言(Experiment Markup Language, EML),从而实现可复现且忠实的大规模比较。在初始发布中,我们整理并标准化了来自100篇论文的258个聚焦于人类常识推理的认知实验,并将50个大语言模型与人类响应进行对比评估。我们发现了一个清晰的规模扩展趋势:规模更大、更新的AI模型能更好地复现人类判断。然而,AI模型在这类常识推理任务上的改进速度明显慢于在数学和编程等形式化推理基准上的提升,且模型与人类的拟合度仍远低于人类的分半信度(文本为$R^2 = 0.93$,图像为$0.95$,视频为$0.92$),最佳模型在文本实验上仅达到$R^2 = 0.59$,图像上为$0.58$,视频上为$0.43$。我们希望CogGym能够作为一个持续演进的评估框架,不断纳入新的认知科学实验,以刻画模型行为在哪些方面与人类行为相似、在哪些方面存在系统性差异,以及随着模型与实验的发展这些模式如何变化。