Benchmark finds telegram-style prompts cut LLM output tokens up to 49%
A new open-source benchmark called the Telegraph Test shows that instructing large language models to answer in 'cablese' — the clipped, article-free style once used by telegraph operators — reduces billed output tokens by 40-49% on several models' own API meters, while preserving or even slightly improving factual recall. The effect held across four model families tested with roughly 1,300 questions over 50 passages, with one exception: gpt-5-mini's reasoning overhead made cablese prompting roughly double its cost instead of cutting it.
GoKawiil's interpretation of the reporting above, not reported fact.
If the results generalize, machine-to-machine LLM traffic — where no human needs natural prose — could see substantial API cost reductions simply by changing a single instruction, without fine-tuning or added infrastructure. The researchers' open release of code and data invites independent verification, which matters given how counterintuitive and commercially significant the claimed savings are.
- A one-sentence 'answer in cablese' instruction cut output tokens 40-49% across multiple model families on their own API meters.
- Despite the compression, recovery of factual information stayed at or above baseline (ratios 0.99-1.10) across tested models.
- gpt-5-mini was an outlier: its inability to disable reasoning meant cablese prompting roughly doubled its write costs instead of saving money.
Source: fiveminutesforward.com — Travis Smith, 2026-10-07
Published there as: “Write Like It's 1866: LLMs Relearn Telegraphese”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.