News

01/proven / not yet

What's proven. What's not.

Most tools hide this section. We think it is the product.

Proven
  • The problem is real. Token cost is a top operational pain in multi-agent systems, and the waste sits between agents, not inside them. (public market data, sources linked)
  • External evidence the problem exists. F1 0.72 on 1,575 traces from a single LangGraph stock-market application. Hybrid method, unrelated implementation (arXiv:2511.10650).
Not yet
  • Real-world semantic separation. In our real-data probes, the zero-false-positive result came from the structural layer; the semantic layer did not cleanly separate same-topic outputs. We published this finding instead of hiding it.
  • Measured savings. We have not yet measured a single dollar saved in production. No user count, no testimonials, because there are none yet.

When something moves from the right column to the left, you'll read it here first.


02/reports

Reports

Reports published under the previous product name Clew. Boxdawn is the same product, renamed 2026-08.

2026-08

Waste-rate replicates across coding and non-coding agents

Claude Code (28 sessions) · Toolathlon (6,780 trajectories · 22 models)GO

Union of 4 deterministic detectors reports 99.3% WR_char on 28 Claude Code sessions and 93.4% WR_char on 6,780 Toolathlon trajectories (22 frontier models). Same pattern, structurally distinct corpora: coding sessions and non-coding tool-use tasks.

2026-08

Toolathlon adapter unblocked via pre-registered amendment

Toolathlon (6,780 trajectories · 22 models)GO

August 10 scan excluded 100% of trajectories due to an adapter gap. Amendment prereg (docs/WASTE_RATE_TOOLATHLON_ADAPTER_AMENDMENT_PREREG.md) locked reconstruction rules and prediction bands on August 11 before any code landed. Re-scan produced 4/4 measurable prediction passes with no post-hoc adjustment.

2026-07

A pattern we implemented but never saw

Claude Code · Toolathlon · RedundancyBenchN/A

One of the three detected patterns has never fired on real data. We label it as such rather than advertising it.

2026-07

The re-read detector we killed

Claude Code sessionsKILL

Thirty pairs drawn at random and hand-labeled. Twenty-nine were an agent reading a long file in slices. The one that was left was a tool refusing the same locked PDF twice, so precision is 1/30 or 0/30 depending on whether that counts. The pre-registered threshold was 0.70.


Boxdawn diagnoses; it does not fix. No measured cost savings yet. We report what was found, not what was saved.