Skip to main content
TokenCost logoTokenCost

The Best LLMs for Tool Use on tau2-bench

Which model is most reliable at calling tools?

81 of the 135 buyable models publish a tau2-bench score. It measures multi-turn function calling against a simulated user, which is a different question from writing code.

Method: Published tau2-bench score, descending, exactly as reported. tau2-bench scores multi-turn tool use against a simulated user, so it measures whether a model calls the right function with the right arguments, not whether it writes good code. Pricing as of July 2026. 28 rows below carry a score measured at a non-default reasoning tier and are marked as such; every affected model is named on the hub. Read the full method.

RankModeltau2-benchTerminal-BenchBlended $/1MContext
#1GLM-5.2Zhipu· non-default tier99.177.9$2.151M
#2GLM-4.7-FlashXZhipu· non-default tier98.8not published$0.152200K
#3Step 3.7 FlashStepFun98.539.3$0.438262K
#4GLM-5 TurboZhipu98.5not published$1.90200K
#5Claude Fable 5Anthropic98.584.6$20.001M
#6GLM-5Zhipu· non-default tier98.2not published$1.55200K
#7Qwen3.6-PlusAlibaba97.761.4$1.131M
#8GLM-5.1Zhipu· non-default tier97.761.8$2.15200K
#9GLM-4.7Zhipu· non-default tier95.945.3$1.00200K
#10Kimi K2.5Moonshot· non-default tier95.945.7$1.20262K
#11Kimi K2.6Moonshot95.965.9$1.71262K
#12Qwen3.6-Max-PreviewAlibaba95.9not published$2.92262K
#13DeepSeek V4-FlashDeepSeek· non-default tier95.656.9$0.1751M
#14Gemini 3.5 FlashGoogle95.676.2$3.381.05M
#15MiniMax M2.5MiniMax95.3not published$0.525128K
#16MiMo-V2-ProXiaomi95.0not published$1.501.05M
#17Qwen3.7 MaxAlibaba94.774.5$3.751M
#18Claude Opus 4.8Anthropic94.484.6$10.001M
#19DeepSeek V4-ProDeepSeek· non-default tier94.264.8$0.5441M
#20MiMo-V2.5-ProXiaomi94.265.2$0.5441.05M
#21Mistral Medium 3.5Mistral94.250.6$3.00256K
#22Qwen3.5-27BAlibaba· non-default tier93.9not published$0.825262K
#23Qwen3.7 PlusAlibaba93.061.0$0.7001M
#24Grok 4.20xAI· non-default tier93.0not published$1.561M
#25Claude Opus 4.6Anthropic· non-default tier92.1not published$10.001M
#26GPT-5.5OpenAI91.880.5$11.251.05M
#27Grok 4.3xAI91.2not published$1.561M
#28Kimi K2.7 CodeMoonshot90.167.4$1.71262K
#29Claude Opus 4.5Anthropic· non-default tier89.5not published$10.00200K
#30MiniMax M3MiniMax88.965.2$1.051M
#31Claude Opus 4.7Anthropic· non-default tier88.683.1$10.001M
#32Qwen3.5-Omni PlusAlibaba88.3not published$1.50262K
#33Qwen3.5-9BAlibaba· non-default tier86.829.2$0.113262K
#34GPT-5OpenAI86.5not published$3.44400K
#35GPT-5.3 CodexOpenAI86.0not published$4.81400K
#36MiniMax M2.7MiniMax84.855.4$0.525205K
#37Qwen3.5-Omni FlashAlibaba84.5not published$0.850262K
#38GPT-5.1OpenAI· non-default tier81.952.4$3.44400K
#39Nex-N2-ProNex AGI81.667.8$1.00262K
#40GPT-5.6 SolOpenAI81.086.1$11.251.05M
#41o3OpenAI80.7not published$3.50200K
#42Command ACohere80.722.8$4.38128K
#43Command A+Cohere80.722.8$4.38128K
#44Qwen3 Coder NextAlibaba79.538.2$0.282262K
#45Nova 2.0 LiteAmazon75.7not published$0.8501M
#46Claude Sonnet 4.6Anthropic· non-default tier75.771.2$6.001M
#47GPT-5.4OpenAI· non-default tier74.6not published$5.631.05M
#48GPT-5.2OpenAI74.3not published$4.81400K
#49GPT-5.6 TerraOpenAI72.872.3$4.501.05M
#50GPT-5 MiniOpenAI71.1not published$0.688400K
#51Mercury 2Inception Labs70.827.3$0.375128K
#52o1OpenAI62.6not published$26.25200K
#53Gemma 4 31BGoogle· non-default tier59.943.4$0.180262K
#54o4 MiniOpenAI55.6not published$1.93200K
#55Claude 3.7 SonnetAnthropic· non-default tier54.7not published$6.00200K
#56Gemini 2.5 ProGoogle54.128.5$3.441.05M
#57GPT-4.1 MiniOpenAI52.910.1$0.7001.05M
#58GPT-5.4 NanoOpenAI52.6not published$0.463400K
#59GPT-OSS 20BOpenAI· non-default tier50.3not published$0.131131K
#60GPT-4.1OpenAI47.1not published$3.501.05M
#61GPT-OSS 120BOpenAI· non-default tier45.013.9$0.077131K
#62Gemma 4 26B A4BGoogle· non-default tier43.639.0$0.120262K
#63Mistral Small 4Mistral· non-default tier41.221.0$0.263256K
#64GPT-5.4 MiniOpenAI36.5not published$1.69400K
#65Pixtral LargeMistral36.5not published$3.00128K
#66Gemma 4 12BGoogle· non-default tier36.327.3$0.00262K
#67Gemini 2.5 FlashGoogle· non-default tier31.6not published$0.8501.05M
#68Gemini 3.1 Flash-LiteGoogle31.331.1$0.5631.05M
#69GPT-5 NanoOpenAI30.4not published$0.138400K
#70o3 MiniOpenAI28.7not published$1.93200K
#71Devstral SmallMistral· non-default tier28.4not published$0.150256K
#72Ministral 3 14BMistral27.29.7$0.200262K
#73Ministral 3 8BMistral26.64.1$0.150262K
#74Magistral SmallMistral26.6not published$0.75040K
#75GPT-4oOpenAI· non-default tier25.1not published$4.38128K
#76Mistral Large 3Mistral24.612.0$0.750262K
#77Magistral MediumMistral23.1not published$2.7540K
#78Gemini 2.5 Flash-LiteGoogle· non-default tier18.4not published$0.1751.05M
#79Llama 4 MaverickMeta17.87.9$0.4151.05M
#80GPT-4.1 NanoOpenAI17.33.7$0.1751.05M
#81Llama 4 ScoutMeta15.53.7$0.1351.05M

Frequently asked

Nearby questions, answered elsewhere

All rankings