Hi, nice to meet you!

I am Botao Yu (余博涛), a fourth-year PhD student at The Ohio State University, fortunately advised by Prof. Huan Sun. Previously, I earned my Master’s degree at Nanjing University.

My research focuses on language agents, tool-integrated reasoning, and scientific discovery. I build and evaluate agentic frameworks and systems for complex problem-solving, with expertise in agent architecture design, tool learning, and multi-level evaluation methodologies. I am particularly interested in developing LLM agents that can use/create tools to solve complex problems in both general and scientific domains.

I also have experience in natural language processing, information extraction, music understanding and generation, and computational chemistry, which provides me with diverse perspectives and transferable skills for tackling new research challenges.

🚀 Actively seeking a research internship for summer 2027 on language agents, tool-integrated reasoning, and AI for science. Feel free to reach out via email!

🌟 Featured Projects

SAGA
SAGA

SAGA

SAGA, a generalist agentic framework that automates objective planning for scientific discovery. SAGA employs a bi-level architecture where an outer loop of LLM agents proposes new objectives and analyzes optimization outcomes, while an inner loop performs solution optimization. Applied across antibiotic design, material design, DNA sequence design, and chemical process design, demonstrating how agents can systematically explore objective spaces.
ChemToolAgent
ChemToolAgent

ChemToolAgent

A systematic investigation into tool-augmented language agents. Using chemistry as a testbed, ChemToolAgent reveals fundamental insights about when and how tools help agents: tools don't always improve performance and can introduce new error modes; whether tools help depends on specific tasks. We also release ChemMCP, an MCP-compatible toolkit for easily building chemistry co-scientists.
Mind2Web 2
Mind2Web 2

Mind2Web 2

A benchmark for evaluating agents on realistic, long-horizon agentic search tasks with agent-as-a-judge methodology. Comprises 130 high-quality tasks requiring real-time web browsing and extensive information synthesis, advancing rigorous evaluation of complex agentic systems beyond simple task completion metrics.
LlaSMol
LlaSMol

LlaSMol

Investigating how to adapt LLMs to specialized domains through high-quality instruction tuning. LlaSMol demonstrates that careful data curation and task diversity matter more than scale—insights that generalize beyond chemistry to other domain adaptation challenges in building capable language agents.

🔥 News

  • 2026.09: Our paper PhenoAIR on agentic mechanism of action prediction from Cell Painting profiles is accepted to NeurIPS 2026 🎉.
  • 2026.01: Our paper Holistic Agent Leaderboard (HAL) is accepted to ICLR 2026 🎉.
  • 2025.12: Check out our new preprint SAGA, an autonomous agent framework that automates objective function design for scientific discovery through a bi-level architecture.
  • 2025.12: Check out our new preprint Scientific Discovery Evaluation (SDE), a scenario-grounded benchmark for evaluating LLMs in scientific discovery across biology, chemistry, materials, and physics.
  • 2025.10: Check out our new preprint Holistic Agent Leaderboard (HAL), addressing key challenges in AI agent evaluation through standardized evaluation harness and three-dimensional analysis across models, scaffolds, and benchmarks.
  • 2025.10: Our paper AutoSDT got the best paper award at the LLM for Scientific Discovery workshop @ COLM 2025 🎉🏆.
  • 2025.09: Our paper Mind2Web 2 is accepted to NeurIPS 2025 🎉.
  • 2025.09: Our paper LARC is accepted to AIAS 2025 and selected as the best paper award 🎉🏆.
  • 2025.09: Check out our new preprint LARC, an agentic framework for constrained retrosynthesis planning.
  • 2025.08: Our paper AutoSDT is accepted to EMNLP 2025 🎉.
  • 2025.06: Check out our new preprint Mind2Web 2, a benchmark for evaluating agentic search with agent-as-a-judge.
  • 2025.06: Check out our new preprint AutoSDT, an automated pipeline for generating high-quality scientific coding tasks.
  • 2025.06: Check out 🛠️ChemMCP, our newly released, MCP-compatible chemistry toolkit for LLMs and AI assistants. Let’s build it together!
  • 2025.05: Check out our new preprint Topic Association Analysis, where we investigated why LLMs misclassify benign comments as toxic from the topic association bias perspective.
  • 2025.05: Our paper MMMU-Pro is accepted to ACL 2025.
  • 2025.03: Our ChemAgent is now renamed to ChemToolAgent. Check out our new version with more experimental results at arXiv.
  • 2025.01: Our paper ChemAgent is accepted to NAACL 2025 Findings.
  • 2025.01: Our paper ScienceAgentBench is accepted to ICLR 2025.
  • 2024.11: Please check out our new preprint ChemAgent, an enhanced chemistry agent and its performance on various chemistry problems.
  • 2024.10: Please check out our new preprint ScienceAgentBench, a benchmark to assess language models in scientific tasks.
  • 2024.09: Check out our new preprint MMMU-Pro, an enhanced version of MMMU featuring full-vision evaluation.
  • 2024.07: Our paper LlaSMol is accepted to COLM 2024 🎉!
  • 2024.05: Our paper MMMU is selected as Oral (0.8%) and nominated for best paper (24 in total) at CVPR 2024 🎊!
  • 2024.02: Please check out our preprint LlaSMol, where we propose an awesome chemistry task instruction tuning dataset and a series of chemistry LLMs.
  • 2023.08: Arrived at Columbus. My PhD journey officially starts 😋!
  • 2023.05: Please check out our preprint MuseCoco, a text-to-music generation system.
  • 2022.09: Our paper Museformer is accepted to NeurIPS 2022 🎉!

📝 Publications

There is no such publication.
  • [Preprint 2026] FREA: A Multi-Source Expert Benchmark for Reaction Feasibility Verification

    Botao Yu, Bo Zhou*, Daniel Adu-Ampratwum*, Frazier N. Baker, Ziru Chen, Reza Averly, Ye Liu, Wenhao Gao, Xia Ning, Huan Sun (* equal contribution)
    FREA, an expert-labeled benchmark of 751 reactions that tests whether reaction feasibility verifiers, including LLMs and forward models, agree with chemists across different candidate sources.
  • [NeurIPS 2026] From Retrieval to Reasoning: Agentic Mechanism Prediction from Cell Painting Profiles

    Jiayuan Chen, Botao Yu, Tianyu Liu, Thai-Hoang Pham, Meng Wu, Ping Zhang
    A multi-agent framework that predicts mechanism of action from Cell Painting profiles by reasoning over retrieved neighbors as uncertain evidence, instead of directly trusting nearest-neighbor matches.
  • [Preprint 2026] Benchmarking AI Agents for Addressing Scientific Challenges Across Scales

    Tianyu Liu, Allen Xin Wang, Antonia Panescu, Lisa Xinyi Chen, Wenxin Long, Xinyu Wei, Yueqian Jing, Ziyao Zeng, Jihang Chen, Sihan Jiang, Ziqing Wang, Siyi Gu, Siyu Chen, Xinyang Hu, Haoran Shao, Leqi Xu, Wangjie Zheng, Zhiyuan Cao, Ada Fang, Botao Yu, Kunyang Sun, Rex Ying, Arman Cohan, Qingyu Chen, Lingzhou Xue, Kaize Ding, Yuanqi Du, Wengong Jin, Zhuoran Yang, Marinka Zitnik, James Zou, Hua Xu, Hongyu Zhao
    SciAgentArena, an interactive benchmark of ~200 real-world research tasks across scientific domains. Current agents handle well-specified data analysis but struggle with novel insights and open-ended exploration.
  • [Preprint 2026] ARMOR: An Agentic Framework for Reaction Feasibility Prediction via Adaptive Utility-aware Multi-tool Reasoning

    Ye Liu, Botao Yu, Xinyi Ling, Daniel Adu-Ampratwum, Xia Ning
    An agentic framework that combines multiple chemistry tools for reaction feasibility prediction by modeling each tool's utility and resolving conflicts between their predictions.
  • [Preprint 2026] A Versatile AI Agent for Rare Disease Diagnosis and Risk Gene Prioritization

    Tianyu Liu, Wangjie Zheng, Rui Yang, Benny Kai Guo Loo, Hui Zhang, Jeffries Lauran, Jianlei Gu, Botao Yu, Weihao Xuan, Kexin Huang, Nan Liu, James Zou, Yonghui Jiang, Hua Xu, Hongyu Zhao
    Hygieia, a multi-modal agent that integrates phenotypes, genetics, and clinical records for rare disease diagnosis and risk gene prioritization, outperforming physicians in expert-validated studies.
  • [Preprint 2026] MMORF: A Multi-agent Framework for Designing Multi-objective Retrosynthesis Planning Systems

    Frazier N. Baker, Trieu Nguyen, Reza Averly, Botao Yu, Daniel Adu-Ampratwum, Huan Sun, Xia Ning
    A modular framework for building and comparing multi-agent systems for multi-objective retrosynthesis planning that balances quality, safety, and cost.
  • [Preprint 2025] Accelerating Scientific Discovery with Autonomous Goal-evolving Agents

    Yuanqi Du*, Botao Yu*, Tianyu Liu*, Tony Shen*, Junwu Chen*, Jan G. Rittig*, Kunyang Sun*, Yikun Zhang*, Zhangde Song, Bo Zhou, Cassandra Masschelein, Yingze Wang, Haorui Wang, Haojun Jia, Chao Zhang, Hongyu Zhao, Martin Ester, Teresa Head-Gordon, Carla P. Gomes, Huan Sun, Chenru Duan, Philippe Schwaller, Wengong Jin (* equal contribution)
    SAGA, an agentic framework that automates objective design for scientific discovery: LLM agents iteratively propose and refine objectives while an inner loop optimizes solutions against them.
  • [Preprint 2025] Evaluating Large Language Models in Scientific Discovery

    Zhangde Song*, Jieyu Lu*, Yuanqi Du*, Botao Yu*, Thomas M. Pruyn*, Yue Huang*, Kehan Guo*, Xiuzhe Luo*, Yuanhao Qu*, Yi Qu, Yinkai Wang, Haorui Wang, Jeff Guo, Jingru Gan, Parshin Shojaee, Di Luo, Andres M Bran, Gen Li, Qiyuan Zhao, Shao-Xiong Lennon Luo, Yuxuan Zhang, Xiang Zou, Wanru Zhao, Yifan F. Zhang, Wucheng Zhang, Shunan Zheng, Saiyang Zhang, Sartaaj Takrim Khan, Mahyar Rajabi-Kochi, Samantha Paradi-Maropakis, Tony Baltoiu, Fengyu Xie, Tianyang Chen, Kexin Huang, Weiliang Luo, Meijing Fang, Xin Yang, Lixue Cheng, Jiajun He, Soha Hassoun, Xiangliang Zhang, Wei Wang, Chandan K. Reddy, Chao Zhang, Zhiling Zheng, Mengdi Wang, Le Cong, Carla P. Gomes, Chang-Yu Hsieh, Aditya Nandy, Philippe Schwaller, Heather J. Kulik, Haojun Jia, Huan Sun, Seyed Mohamad Moosavi, Chenru Duan (* equal contribution)
    A scenario-grounded benchmark that evaluates LLMs on scientific discovery across biology, chemistry, materials, and physics, at both the question level and the research-project level.
  • [ICLR 2026] Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation

    Sayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir, Zachary S Siegel, Boyi Wei, Tianci Xue, Ziru Chen, Felix Chen, Saiteja Utpala, Franck Ndzomga, Dheeraj Oruganty, Sophie Luskin, Kangheng Liu, Botao Yu, Amit Arora, Dongyoon Hahm, Harsh Trivedi, Huan Sun, Juyong Lee, Tengjun Jin, Yifan Mai, Yifei Zhou, Yuxuan Zhu, Rishi Bommasani, Daniel Kang, Dawn Song, Peter Henderson, Yu Su, Percy Liang, Arvind Narayanan
    A standardized infrastructure for AI agent evaluation, with a parallel evaluation harness and three-dimensional analysis across models, scaffolds, and benchmarks.
  • [AIAS 2025] LARC: Towards Human-level Constrained Retrosynthesis Planning through an Agentic Framework

    🏆 Best Paper Award at AIAS 2025
    Frazier N. Baker, Daniel Adu-Ampratwum, Reza Averly, Botao Yu, Huan Sun, Xia Ning
    The first LLM-based agentic framework for constrained retrosynthesis planning, which uses tool-grounded agentic feedback to evaluate constraints during route generation.
  • [NeurIPS 2025] Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge

    Boyu Gou*, Zanming Huang*, Yuting Ning*, Yu Gu, Michael Lin, Weijian Qi, Andrei Kopanev, Botao Yu, Bernal Jiménez Gutiérrez, Yiheng Shu, Chan Hee Song, Jiaman Wu, Shijie Chen, Hanane Nour Moussa, Tianshu Zhang, Jian Xie, Yifei Li, Tianci Xue, Zeyi Liao, Kai Zhang, Boyuan Zheng, Zhaowei Cai, Viktor Rozgic, Morteza Ziyadi, Huan Sun, Yu Su (* equal contribution)
    A benchmark of 130 realistic, long-horizon agentic search tasks that require real-time web browsing and information synthesis, evaluated with agent-as-a-judge.
  • [EMNLP 2025] AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists

    🏆 Best Paper Award at the LLM for Scientific Discovery workshop @ COLM 2025
    Yifei Li*, Hanane Nour Moussa*, Ziru Chen, Shijie Chen, Botao Yu, Mingyi Xue, Benjamin Burns, Tzu-Yao Chiu, Vishal Dey, Zitong Lu, Chen Wei, Qianheng Zhang, Tianyu Zhang, Song Gao, Xuhui Huang, Xia Ning, Nesreen K. Ahmed, Ali Payani, Huan Sun (* equal contribution)
    An automated pipeline that collects scientific coding tasks from real-world data-driven workflows, yielding AutoSDT-5K, the largest open dataset of its kind for training AI co-scientists.
  • [Preprint 2025] Probing Association Biases in LLM Moderation Over-Sensitivity

    Yuxin Wang, Botao Yu, Ivory Yang, Saeed Hassanpour, Soroush Vosoughi
    Shows that LLMs over-flag benign comments as toxic partly due to topic-level association biases, not just offensive keywords.
  • [NAACL 2025 Findings] ChemToolAgent: The Impact of Tools on Language Agents for Chemistry Problem Solving

    Botao Yu, Frazier N. Baker*, Ziru Chen*, Garrett Herb, Boyu Gou, Daniel Adu-Ampratwum, Xia Ning, Huan Sun (* equal contribution)
    A systematic study of when tools help language agents, using chemistry as a testbed: tools do not always help and can introduce new error modes. We also release ChemMCP, an MCP-compatible chemistry toolkit.
  • [ICLR 2025] ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

    Ziru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li, Zeyi Liao, Chen Wei, Zitong Lu, Vishal Dey, Mingyi Xue, Frazier N. Baker, Benjamin Burns, Daniel Adu-Ampratwum, Xuhui Huang, Xia Ning, Song Gao, Yu Su, Huan Sun
    A benchmark of 102 expert-validated tasks extracted from peer-reviewed publications for rigorously assessing language agents in data-driven scientific discovery.
  • [ACL 2025] MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

    Xiang Yue*, Tianyu Zheng*, Yuansheng Ni*, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Ming Yin, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, Graham Neubig (* equal contribution)
    A more robust version of MMMU featuring full-vision evaluation for multi-discipline multimodal understanding.
  • [COLM 2024] LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset

    Botao Yu, Frazier N. Baker*, Ziqi Chen*, Xia Ning, Huan Sun (* equal contribution)
    SMolInstruct, a large-scale, high-quality instruction tuning dataset for chemistry, and LlaSMol models trained on it that significantly outperform GPT-4 and Claude-3-Opus on chemistry tasks.
  • [CVPR 2024 Oral] MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

    Xiang Yue*, Yuansheng Ni*, Kai Zhang*, Tianyu Zheng*, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun*, Yu Su*, Wenhu Chen* (* core contributors)
    A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI.
  • [Preprint 2023] MuseCoco: Generating Symbolic Music from Text

    Peiling Lu*, Xin Xu*, Chenfei Kang*, Botao Yu*, Chengyi Xing*, Xu Tan, Jiang Bian (* equal contribution)
    A two-stage text-to-music generation system for creating symbolic music from textual descriptions.
  • [Preprint 2023] EmoGen: Eliminating Subjective Bias in Emotional Music Generation

    Chenfei Kang, Peiling Lu, Botao Yu, Xu Tan, Wei Ye, Shikun Zhang, Jiang Bian
    A method for generating emotional music while eliminating subjective bias in emotion labels.
  • [NeurIPS 2022] Museformer: Transformer with Fine- and Coarse-Grained Attention for Music Generation

    Botao Yu, Peiling Lu, Rui Wang, Wei Hu, Xu Tan, Wei Ye, Shikun Zhang, Tao Qin, Tie-Yan Liu
    A Transformer with fine- and coarse-grained attention for modeling the structures of long music sequences.
  • [ISMIR 2022] MeloForm: Generating Melody with Musical Form Based on Expert Systems and Neural Networks

    Peiling Lu, Xu Tan, Botao Yu, Tao Qin, Sheng Zhao, Tie-Yan Liu
    A melody generation system that combines expert systems and neural networks to follow given musical forms.
  • [EMNLP 2021] Knowing False Negatives: An Adversarial Training Method for Distantly Supervised Relation Extraction

    Kailong Hao, Botao Yu, Wei Hu
    An adversarial training method that improves distantly supervised relation extraction by addressing false negatives.
  • [APWeb-WAIM 2020] Joint Reasoning of Events, Participants and Locations for Plot Relation Recognition

    Shengguang Qiu, Botao Yu, Lei Qian, Qiang Guo, Wei Hu
    A method for recognizing plot relations by jointly reasoning about events, participants, and locations in narratives.

👨🏻‍💻 Internship

  • Research intern @ Microsoft Research New England

    2026.05 - 2026.08       Cambridge, Massachusetts, USA

    Worked with Yuanqi Du on AI scientist agents.

  • Research intern @ Microsoft Research Asia (微软亚洲研究院)

    2021.04 - 2022.03       Beijing, China

    Worked with Xu Tan, Peiling Lu, and Rui Wang on music generation in the Muzic project.

💻 Service

  • 2026: Student helper for AI Scientist Summer Workshop; Reviewer for COLM 2026, COLM 2026 LM4Sci Workshop, ARR 2026 (Oct.), ICLR 2027
  • 2025: Reviewer for ARR 2025 (Feb., May, July, Oct.), COLM 2025, NeurIPS 2025 SEA Workshop, ICLR 2026
  • 2024: Reviewer for ICLR 2025, ARR 2024 (Dec.), AAAI 2025 AI4Research Workshop

📖 Education

  • PhD student in Computer Science and Engineering @ The Ohio State University

    2023.08 - Now       Columbus, Ohio, USA

  • Master’s student in Computer Science @ Nanjing University (南京大学)

    2019.09 - 2023.06       Nanjing, Jiangsu, China

  • Undergraduate student in Software Engineering @ Dalian University of Technology (大连理工大学)

    2015.09 - 2019.06       Dalian, Liaoning, China

  • High school student @ The High School Attached To Hunan Normal University (湖南师大附中)

    2012.09 - 2015.06       Changsha, Hunan, China

Psst! 🔍 Kudos on your keen eye! Didn't expect anyone to notice this microscopic text. Since you've ventured this far, fancy embarking on a friendship adventure?

Last updated: October 6, 2026