01/the ask

Running multi-agent in production? We want your trace.

Boxdawn is validated on 17,881 traces across four public benchmark corpora (28 Claude Code sessions, 6,780 Toolathlon trajectories, 10,056 Exgentic agent-LLM traces, and 1,017 Claude Code sessions we did not collect). The next honest step is private production traces from your team, and that's where you come in. Send us an execution trace (OTel or OpenInference JSON) and we'll run it through Boxdawn and send back exactly what it found, including false positives. No signup, nothing sold to you.

Can't share data externally? Fair. That's the point of local-first. Run Boxdawn yourself and share only the numbers. Either way, your trace directly shapes the recalibration, and we'll credit you publicly if you want.