Hi, nice to meet you!
I am Botao Yu (余博涛), a fourth-year PhD student at The Ohio State University, fortunately advised by Prof. Huan Sun. Previously, I earned my Master’s degree at Nanjing University.
My research focuses on language agents, tool-integrated reasoning, and scientific discovery. I build and evaluate agentic frameworks and systems for complex problem-solving, with expertise in agent architecture design, tool learning, and multi-level evaluation methodologies. I am particularly interested in developing LLM agents that can use/create tools to solve complex problems in both general and scientific domains.
I also have experience in natural language processing, information extraction, music understanding and generation, and computational chemistry, which provides me with diverse perspectives and transferable skills for tackling new research challenges.
🚀 Actively seeking a research internship for summer 2027 on language agents, tool-integrated reasoning, and AI for science. Feel free to reach out via email!
🌟 Featured Projects
SAGA
ChemToolAgent
Mind2Web 2
LlaSMol
🔥 News
- 2026.09: Our paper PhenoAIR on agentic mechanism of action prediction from Cell Painting profiles is accepted to NeurIPS 2026 🎉.
- 2026.01: Our paper Holistic Agent Leaderboard (HAL) is accepted to ICLR 2026 🎉.
- 2025.12: Check out our new preprint SAGA, an autonomous agent framework that automates objective function design for scientific discovery through a bi-level architecture.
- 2025.12: Check out our new preprint Scientific Discovery Evaluation (SDE), a scenario-grounded benchmark for evaluating LLMs in scientific discovery across biology, chemistry, materials, and physics.
- 2025.10: Check out our new preprint Holistic Agent Leaderboard (HAL), addressing key challenges in AI agent evaluation through standardized evaluation harness and three-dimensional analysis across models, scaffolds, and benchmarks.
- 2025.10: Our paper AutoSDT got the best paper award at the LLM for Scientific Discovery workshop @ COLM 2025 🎉🏆.
- 2025.09: Our paper Mind2Web 2 is accepted to NeurIPS 2025 🎉.
- 2025.09: Our paper LARC is accepted to AIAS 2025 and selected as the best paper award 🎉🏆.
- 2025.09: Check out our new preprint LARC, an agentic framework for constrained retrosynthesis planning.
- 2025.08: Our paper AutoSDT is accepted to EMNLP 2025 🎉.
- 2025.06: Check out our new preprint Mind2Web 2, a benchmark for evaluating agentic search with agent-as-a-judge.
- 2025.06: Check out our new preprint AutoSDT, an automated pipeline for generating high-quality scientific coding tasks.
- 2025.06: Check out 🛠️ChemMCP, our newly released, MCP-compatible chemistry toolkit for LLMs and AI assistants. Let’s build it together!
- 2025.05: Check out our new preprint Topic Association Analysis, where we investigated why LLMs misclassify benign comments as toxic from the topic association bias perspective.
- 2025.05: Our paper MMMU-Pro is accepted to ACL 2025.
- 2025.03: Our ChemAgent is now renamed to ChemToolAgent. Check out our new version with more experimental results at arXiv.
- 2025.01: Our paper ChemAgent is accepted to NAACL 2025 Findings.
- 2025.01: Our paper ScienceAgentBench is accepted to ICLR 2025.
- 2024.11: Please check out our new preprint ChemAgent, an enhanced chemistry agent and its performance on various chemistry problems.
- 2024.10: Please check out our new preprint ScienceAgentBench, a benchmark to assess language models in scientific tasks.
- 2024.09: Check out our new preprint MMMU-Pro, an enhanced version of MMMU featuring full-vision evaluation.
- 2024.07: Our paper LlaSMol is accepted to COLM 2024 🎉!
- 2024.05: Our paper MMMU is selected as Oral (0.8%) and nominated for best paper (24 in total) at CVPR 2024 🎊!
- 2024.02: Please check out our preprint LlaSMol, where we propose an awesome chemistry task instruction tuning dataset and a series of chemistry LLMs.
- 2023.08: Arrived at Columbus. My PhD journey officially starts 😋!
- 2023.05: Please check out our preprint MuseCoco, a text-to-music generation system.
- 2022.09: Our paper Museformer is accepted to NeurIPS 2022 🎉!
📝 Publications
-
[Preprint 2026] FREA: A Multi-Source Expert Benchmark for Reaction Feasibility Verification
FREA, an expert-labeled benchmark of 751 reactions that tests whether reaction feasibility verifiers, including LLMs and forward models, agree with chemists across different candidate sources. -
[NeurIPS 2026] From Retrieval to Reasoning: Agentic Mechanism Prediction from Cell Painting Profiles
A multi-agent framework that predicts mechanism of action from Cell Painting profiles by reasoning over retrieved neighbors as uncertain evidence, instead of directly trusting nearest-neighbor matches. -
[Preprint 2026] Benchmarking AI Agents for Addressing Scientific Challenges Across Scales
SciAgentArena, an interactive benchmark of ~200 real-world research tasks across scientific domains. Current agents handle well-specified data analysis but struggle with novel insights and open-ended exploration. -
[Preprint 2026] ARMOR: An Agentic Framework for Reaction Feasibility Prediction via Adaptive Utility-aware Multi-tool Reasoning
An agentic framework that combines multiple chemistry tools for reaction feasibility prediction by modeling each tool's utility and resolving conflicts between their predictions. -
[Preprint 2026] A Versatile AI Agent for Rare Disease Diagnosis and Risk Gene Prioritization
Hygieia, a multi-modal agent that integrates phenotypes, genetics, and clinical records for rare disease diagnosis and risk gene prioritization, outperforming physicians in expert-validated studies. -
[Preprint 2026] MMORF: A Multi-agent Framework for Designing Multi-objective Retrosynthesis Planning Systems
A modular framework for building and comparing multi-agent systems for multi-objective retrosynthesis planning that balances quality, safety, and cost. -
[Preprint 2025] Accelerating Scientific Discovery with Autonomous Goal-evolving Agents
SAGA, an agentic framework that automates objective design for scientific discovery: LLM agents iteratively propose and refine objectives while an inner loop optimizes solutions against them. -
[Preprint 2025] Evaluating Large Language Models in Scientific Discovery
A scenario-grounded benchmark that evaluates LLMs on scientific discovery across biology, chemistry, materials, and physics, at both the question level and the research-project level. -
[ICLR 2026] Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
A standardized infrastructure for AI agent evaluation, with a parallel evaluation harness and three-dimensional analysis across models, scaffolds, and benchmarks. -
[AIAS 2025] LARC: Towards Human-level Constrained Retrosynthesis Planning through an Agentic Framework
🏆 Best Paper Award at AIAS 2025The first LLM-based agentic framework for constrained retrosynthesis planning, which uses tool-grounded agentic feedback to evaluate constraints during route generation. -
[NeurIPS 2025] Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
A benchmark of 130 realistic, long-horizon agentic search tasks that require real-time web browsing and information synthesis, evaluated with agent-as-a-judge. -
[EMNLP 2025] AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists
🏆 Best Paper Award at the LLM for Scientific Discovery workshop @ COLM 2025An automated pipeline that collects scientific coding tasks from real-world data-driven workflows, yielding AutoSDT-5K, the largest open dataset of its kind for training AI co-scientists. -
[Preprint 2025] Probing Association Biases in LLM Moderation Over-Sensitivity
Shows that LLMs over-flag benign comments as toxic partly due to topic-level association biases, not just offensive keywords. -
[NAACL 2025 Findings] ChemToolAgent: The Impact of Tools on Language Agents for Chemistry Problem Solving
A systematic study of when tools help language agents, using chemistry as a testbed: tools do not always help and can introduce new error modes. We also release ChemMCP, an MCP-compatible chemistry toolkit. -
[ICLR 2025] ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
A benchmark of 102 expert-validated tasks extracted from peer-reviewed publications for rigorously assessing language agents in data-driven scientific discovery. -
[ACL 2025] MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
A more robust version of MMMU featuring full-vision evaluation for multi-discipline multimodal understanding. -
[COLM 2024] LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset
SMolInstruct, a large-scale, high-quality instruction tuning dataset for chemistry, and LlaSMol models trained on it that significantly outperform GPT-4 and Claude-3-Opus on chemistry tasks. -
[CVPR 2024 Oral] MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI. -
[Preprint 2023] MuseCoco: Generating Symbolic Music from Text
A two-stage text-to-music generation system for creating symbolic music from textual descriptions. -
[Preprint 2023] EmoGen: Eliminating Subjective Bias in Emotional Music Generation
A method for generating emotional music while eliminating subjective bias in emotion labels. -
[NeurIPS 2022] Museformer: Transformer with Fine- and Coarse-Grained Attention for Music Generation
A Transformer with fine- and coarse-grained attention for modeling the structures of long music sequences. -
[ISMIR 2022] MeloForm: Generating Melody with Musical Form Based on Expert Systems and Neural Networks
A melody generation system that combines expert systems and neural networks to follow given musical forms. -
[EMNLP 2021] Knowing False Negatives: An Adversarial Training Method for Distantly Supervised Relation Extraction
An adversarial training method that improves distantly supervised relation extraction by addressing false negatives. -
[APWeb-WAIM 2020] Joint Reasoning of Events, Participants and Locations for Plot Relation Recognition
A method for recognizing plot relations by jointly reasoning about events, participants, and locations in narratives.
👨🏻💻 Internship
-
Research intern @ Microsoft Research New England
2026.05 - 2026.08 Cambridge, Massachusetts, USA
Worked with Yuanqi Du on AI scientist agents.
-
Research intern @ Microsoft Research Asia (微软亚洲研究院)
2021.04 - 2022.03 Beijing, China
Worked with Xu Tan, Peiling Lu, and Rui Wang on music generation in the Muzic project.
💻 Service
- 2026: Student helper for AI Scientist Summer Workshop; Reviewer for COLM 2026, COLM 2026 LM4Sci Workshop, ARR 2026 (Oct.), ICLR 2027
- 2025: Reviewer for ARR 2025 (Feb., May, July, Oct.), COLM 2025, NeurIPS 2025 SEA Workshop, ICLR 2026
- 2024: Reviewer for ICLR 2025, ARR 2024 (Dec.), AAAI 2025 AI4Research Workshop
📖 Education
-
PhD student in Computer Science and Engineering @ The Ohio State University
2023.08 - Now Columbus, Ohio, USA
-
Master’s student in Computer Science @ Nanjing University (南京大学)
2019.09 - 2023.06 Nanjing, Jiangsu, China
-
Undergraduate student in Software Engineering @ Dalian University of Technology (大连理工大学)
2015.09 - 2019.06 Dalian, Liaoning, China
-
High school student @ The High School Attached To Hunan Normal University (湖南师大附中)
2012.09 - 2015.06 Changsha, Hunan, China
Last updated: October 6, 2026