Ye, Shun, Suja, Vinny Chandran, Li, Chenlong, Jiang, Chongming, Zamani, Reza, Li, Xiang, Bain, Christopher, Zhou, Yuqi, Peterson, Walker, Wang, Huidong, Hu, Chenglang, Park, Jongchan, Cheng, Xiao, Swedlund, Benjamin, Murillo, Sandra, Sivanandan, Anjali, Sun, Shiyu, Lanfeng, Liang, Islam, Mohammad Tariqul, Joy, Baju C., Khan, Ishaq N., Kumar, Sreedhar S., Mercado-Vásquez, Gabriel, Vizzard, James V., Matthews, Jonathan M., Huang, Helen, Guo, Xiaolu, Nicklow, Ethan, Chen, Guorui, Neff, Ryan A., Maity, Surjendu, Park, Hyeonjin, Joo, Han-ho, Dong, Katherine, Cai, Yuyan, Huang, Weihang, Zou, Yichen, Yan, Rui, Figueroa, Raphael, Goncharov, Artem, Schremmer, Bella Rose, Linton, Lian Elsa, Goda, Keisuke, Gao, Liang, Cheng, Ke, Morsut, Leonardo, Wilson, Jennifer L., Fu, Jianping, Teck, Lim Chwee, Sarkar, Deblina, Hierlemann, Andreas, Tay, Savaş, Hoffmann, Alexander, Griffin, Donald Richieri, Chen, Jun, Kelley, Shana O., Varghese, Shyni, Cheon, Jinwoo, Lam, Wilbur A., Moon, James J., Wong, Wilson W., Mitragotri, Samir, Di Carlo, Dino
Abstract
Large Language Models (LLMs) have demonstrated historic breakthroughs in general reasoning with early successes in biomedical science. However, existing LLM benchmarking emphasizes factual recall, offering limited insight into model performance on frontier and multimodal tasks. We assembled BioEVAL (BioEngineering Validation of AI and LLMs), a global, multi-institutional initiative designed to assess experimental reasoning capability across bioengineering (BE) subfields. BioEVAL spans 11 major BE subfields plus a set of uncategorized items, bringing together 22 research groups to create a PhD-level benchmark comprising 608 evaluation items: 1) 380 multiple-choice questions (MCQs, 359 retained after audit), 2) 218 literature synthesis tasks, and 3) 10 multimodal problems with experimental image interpretation. Benchmark items underwent authoring-group expert review and centralized quality control before evaluation. Following evaluation, a blinded cross-group consensus audit of the highest- and lowest-accuracy MCQ items flagged 21 questions for revision or removal; these were withheld, and all reported MCQ results are computed on the 359 retained items. We evaluated diverse cloud-scale foundation/multimodal models (e.g., ChatGPT, Gemini, and Grok) and locally deployable models suitable for inference on consumer-grade GPUs. Models achieved the highest accuracy of up to 90% on MCQs, similarity score of 0.72 on literature synthesis, and accuracy of 80% on a small sample of multimodal reasoning questions, with substantial performance variation across subfields. Leaderboard rankings characterize current capabilities, limitations, and development priorities across the evaluated BE task categories. BioEVAL is maintained as an extensible benchmark with standardized protocols for continuing expert item contribution and model evaluation.
Chinese Translation
大语言模型(LLM)在通用推理方面取得了历史性突破,并在生物医学科学领域取得了早期成功。然而,现有的LLM基准测试侧重于事实记忆,对模型在前沿任务和多模态任务上的表现缺乏深入的洞察。我们构建了BioEVAL(BioEngineering Validation of AI and LLMs,生物工程AI与大语言模型验证),这是一项全球性、多机构合作的计划,旨在评估生物工程(BE)各子领域的实验推理能力。BioEVAL涵盖11个主要的生物工程子领域以及一组未分类题目,汇聚了22个研究团队,构建了一个包含608个评估题目的博士级基准测试:1)380道多选题(MCQ,经审核后保留359道);2)218项文献综合任务;3)10道涉及实验图像解读的多模态问题。所有基准题目在评估前均经过出题组的专家审查和集中质量控制。评估完成后,对准确率最高和最低的MCQ题目进行了跨组盲审共识审计,标记出21道需修改或删除的题目;这些题目被剔除,所有报告的MCQ结果均基于保留的359道题目计算。我们评估了多种云端规模的基础模型/多模态模型(如ChatGPT、Gemini和Grok),以及适用于消费级GPU推理的本地部署模型。模型在多选题上的最高准确率达90%,文献综合任务的相似度得分为0.72,在小样本多模态推理题上的准确率为80%,且各子领域之间的性能差异显著。排行榜排名刻画了当前模型在所评估的生物工程任务类别上的能力、局限性和发展重点。BioEVAL作为一个可扩展的基准测试持续维护,并为后续专家题目贡献和模型评估提供了标准化协议。