Accuracy, honestly
A referee — human or AI — earns trust by being right, and by being accountable when it isn't. Here is exactly how good Verdict is today, how we measure it, and what it can't call yet.
Numbers current as of July 2026. This page is updated as models improve — measured results only, never projections.
Verdict is trained and evaluated on 23,515 quality-labeled sabre touches from real competition video, each labeled by experienced fencers and referees with the winner and the action. Accuracy is always reported on held-out footage the model has never seen — including a separate benchmark drawn from international-level competition, which is deliberately harder than typical club footage.
The foundation of every call is knowing where each fencer — and each blade — is in every frame. Verdict's pose model tracks 16 keypoints per fencer, including the blade guard, midpoint, and tip, at 94% keypoint accuracy on held-out competition footage. Blade tips at full lunge speed are the hardest single thing to track in fencing video; this is where most of our engineering effort has gone.
One-light touches are unambiguous — the light decides, and Verdict simply reads the box. The AI's job is the two-light touch, where right-of-way determines the winner. On those:
| Measurement | Accuracy | Notes |
|---|---|---|
| All two-light calls (club-level held-out set) | 71% | Every call, including ones the model is unsure about |
| All two-light calls (international competition benchmark) | 69% | Harder footage, faster actions |
| High-confidence calls (top 20% by model confidence) | 92% | When Verdict is sure, it's almost always right |
Confidence matters as much as accuracy: when Verdict isn't sure, it says so rather than guessing. In the app, low-confidence calls are flagged as such — the same way a good referee acknowledges a close call.
Not all touches are equally hard to call. Our current per-action accuracy on held-out two-light touches:
| Action | Accuracy | Status |
|---|---|---|
| Attack vs. counterattack | 81% | Strongest — the most common two-light situation |
| Attack in preparation | 78% | Strong |
| Remise / reprise | 74% | Solid |
| Parry-riposte | 49% | Our hardest problem — see below |
We're benchmarking end-to-end analysis time now and will publish measured numbers here — the same rule applies to speed as to accuracy: we only claim what we've measured.
Every number on this page is a snapshot, not a ceiling. The dataset grows weekly through our expert annotation team, and the two biggest known gains — blade-contact audio and 3D pose — are in active development. When the numbers change, this page changes.
Early-access members get accuracy updates as they happen.
Free tier at launch · One email when your spot opens — no spam.