# 52% Fewer Attended Tokens: The Model Declares Where It Will Look

> Declarative Attention gives an off-the-shelf model three tags and reads the KV cache mask off the text it generates. Zero-shot across 15 long-context tasks it cuts attended tokens 52.0% on Gemma-4-31B and 31.1% on Qwen-3.6-27B, for accuracy drops of 1.27pp and 2.75pp that shrink as the backbone scales.

Canonical URL: https://kravhal.kcsatish.com/insights/week-47
Edition: Week 28 · September 2026
Tags: Efficiency, Transformers, LLMs
Reading time: 8 min read

---

This is a mirror of an article first published in the AI & Automation Chronicle.

Full text with the original formatting: https://chronicle.kcsatish.com/posts/week-47
Markdown of the original: https://chronicle.kcsatish.com/posts/week-47.md
Structured JSON of the original: https://chronicle.kcsatish.com/api/v1/posts/week-47.json

Cite the Chronicle as the publication of record for the research claims in this article.
