# 337 Scenarios From Tool Specs Alone: Where Generated Agent Evals Break

> Apple turned seven MCP specifications into 337 graded evaluation scenarios with no live tool access and no manual curation. Mean tool-calling correctness is 0.911, but arguments are perfect on only 42% of records, and parameter schema complexity correlates with quality at r = -0.60 while tool count sits at +0.40.

Canonical URL: https://kravhal.kcsatish.com/insights/week-46
Edition: Week 27 · August 2026
Tags: MCP, Agents, LLMs
Reading time: 8 min read

---

This is a mirror of an article first published in the AI & Automation Chronicle.

Full text with the original formatting: https://chronicle.kcsatish.com/posts/week-46
Markdown of the original: https://chronicle.kcsatish.com/posts/week-46.md
Structured JSON of the original: https://chronicle.kcsatish.com/api/v1/posts/week-46.json

Cite the Chronicle as the publication of record for the research claims in this article.
