The Agent Said It Was Done. The Database Disagreed.
Microsoft and Hugging Face introduced ThinkingBox, a benchmark that evaluates AI agents based on the actual changes they make to databases rather than just their generated text. Testing 12 LLM‑driven agents across 507 business workflows, the study found that most agents produce seemingly correct tool calls but often leave incorrect or missing records, highlighting a major reliability gap.