Agents'Overreliance on Unreliable Tools
LLM agents use tools to access information and perform computations beyond their parametric knowledge. Existing tool-use benchmarks evaluate whether agents select and call the right tools, assuming that tool returns are reliable. However, tools can return plausible but incorrect outputs. We evaluate 14 models with thre...