MathTutorBench：AI辅导模型教学能力评估基准

cs.AI updates on arXiv.org 10月14日 12:21

MathTutorBench：AI辅导模型教学能力评估基准

本文提出MathTutorBench，一个用于综合评估AI辅导模型教学能力的开源基准。通过评估模型在数学教学中的表现，发现教学专长与学科知识之间存在权衡关系。

arXiv:2502.18940v2 Announce Type: replace-cross Abstract: Evaluating the pedagogical capabilities of AI-based tutoring models is critical for making guided progress in the field. Yet, we lack a reliable, easy-to-use, and simple-to-run evaluation that reflects the pedagogical abilities of models. To fill this gap, we present MathTutorBench, an open-source benchmark for holistic tutoring model evaluation. MathTutorBench contains a collection of datasets and metrics that broadly cover tutor abilities as defined by learning sciences research in dialog-based teaching. To score the pedagogical quality of open-ended teacher responses, we train a reward model and show it can discriminate expert from novice teacher responses with high accuracy. We evaluate a wide set of closed- and open-weight models on MathTutorBench and find that subject expertise, indicated by solving ability, does not immediately translate to good teaching. Rather, pedagogy and subject expertise appear to form a trade-off that is navigated by the degree of tutoring specialization of the model. Furthermore, tutoring appears to become more challenging in longer dialogs, where simpler questioning strategies begin to fail. We release the benchmark, code, and leaderboard openly to enable rapid benchmarking of future models.

Fish AI Reader

AI辅助创作，多种专业模板，深度分析，高质量内容生成。从观点提取到深度思考，FishAI为您提供全方位的创作支持。新版本引入自定义参数，让您的创作更加个性化和精准。

FishAI

鱼阅，AI 时代的下一个智能信息助手，助你摆脱信息焦虑

联系邮箱 441953276@qq.com

相关标签

AI辅导模型教学能力评估 MathTutorBench 学科知识教学专长

相关文章

MS MARCO Web Search: A Large-Scale Information-Rich Web Dataset Featuring Millions of Real Clicked Query-Document Labels

This Week In Machine Learning & AI - 5/27/16: The White House on AI & Aggressive Self-Driving Cars

CinePile: A Novel Dataset and Benchmark Specifically Designed for Authentic Long-Form Video Understanding

取消播放时长改革，B站为什么“怂”了？

‘RAG Me Up’: A Generic AI Framework (Server + UIs) that Enables You to Do RAG on Your Own Dataset Easily

HuggingFace Releases ? FineWeb: A New Large-Scale (15-Trillion Tokens, 44TB Disk Space) Dataset for LLM Pretraining

Unlocking the Language of Proteins: How Large Language Models Are Revolutionizing Protein Sequence Understanding

MAGPIE: A Self-Synthesis Method for Generating Large-Scale Alignment Data by Prompting Aligned LLMs with Nothing

Midjourney: ↩️ @kortizart To the best of our knowledge; you are not in our dataset. Here's a "portrait by Karla Ortiz" vs a "portrait by artist". FW...

Hugging Face: We're excited to welcome @argilla_io to the Hugging Face team! ? Time to democratise good Machine Learning, one dataset at a time!...