⚠️
Status: internal calibration (v1). The figures below come from a small pre-tournament calibration — not an independent, peer-reviewed study. They are indicative of how the system behaves, not a final accuracy claim. A larger, independently-conducted study is in preparation and will be published here when complete. We publish the current state openly rather than waiting, and we clearly label what is measured versus illustrative.
§1 · Study design
What we tested, and how
During pre-tournament calibration, the SEGA AI Judge scored 40 test performances alongside a panel of 3 UYSF-certified human judges. The same performance videos were given to both the AI and the panel, and the two sets of scores were compared. The AI also re-scored each video twice, was tested from front and side camera angles, and was run across a simulated competition day to check for drift. This is a measurement-agreement study: does the AI agree with expert humans, and is it stable?
§2 · Results — human panel vs SEGA AI
Agreement & error
| Metric | Human panel | SEGA AI Judge |
| Mean absolute deviation across 40 performances | ±0.87 pts | ±0.24 pts |
| Inter-run consistency (same video, scored twice) | ±0.62 pts | ±0.00 pts (deterministic) |
| Viewing-angle sensitivity (side vs front) | up to ±1.4 pts | ±0.15 pts |
| Fatigue drift (hour 1 vs hour 6) | ±0.45 pts | ±0.00 pts |
| Correlation with final medal outcome | 0.91 | 0.97 |
Source: internal pre-tournament calibration (40 performances, 3 judges). Same figures as the Evidence Centre. Indicative, not peer-reviewed — see limitations below.
§3 · How to read these numbers
In plain words
Mean absolute deviation (error) — how far a score typically lands from the "true" panel consensus. Smaller is tighter. The AI's ±0.24 means it usually sits within a quarter-point of the reference.
Deterministic (±0.00) — given the same video, the AI returns the exact same score every time. Human panels naturally vary a little between viewings; the AI does not.
Viewing-angle sensitivity — how much the score changes with camera position. Lower is better; the AI is far less angle-dependent than the eye.
Correlation with medal outcome — how closely the scores track who actually medalled (1.0 = perfect). Both are high; the AI is marginally higher here.
§4 · Confidence & honesty
What these figures can and can't say
With only 40 performances, the confidence interval around every number above is wide — a second sample could shift them. These results show the system is internally consistent and closely aligned with expert judgement on this set. They do not yet prove accuracy at scale, across disciplines, or in the hands of independent researchers. Any single-number "accuracy %" would overstate what a 40-sample calibration can support, so we don't publish one.
§5 · Known limitations
Where the system needs care
Small, internal sample. 40 performances scored against 3 judges, run by SEGA — not an independent study. Numbers are indicative only.
Single discipline & setting. Calibrated on competitive yogasana under controlled capture; it is not validated for other movement forms or uncontrolled home video.
Capture-dependent. Accuracy relies on the specified camera setup (1080p / 60fps, fixed tripod, adequate lighting, clear background). Poor capture degrades the measurement.
Measures the measurable. The AI scores execution, stability and difficulty. Interpretive qualities and any ambiguous case are escalated to certified Master Judges — the AI is a reference layer, not the final ruling.
Not medical or diagnostic. Scores describe performance only; they are not a health, injury or fitness assessment.
§6 · What "verified" means here
The claim we do make
When a result is marked "AI-Judge verified," it means the performance was scored by the SEGA AI Judge System under the standard capture protocol, the score is stored in a tamper-evident record, and ambiguous cases were referred to human Master Judges. It does not mean the score has been independently audited by a third party — that is the purpose of the study below.
§7 · Roadmap to independent validation
How this report grows up
Phase 1 · done
Internal calibration
40 performances vs 3 judges — the figures on this page.
Phase 2 · in preparation
Expanded internal study
Several hundred performances, more judges, per-dimension error, published confidence intervals.
Phase 3 · planned
Independent study
A third-party or academic evaluation with full methodology, dataset description and a DOI in the Research index.
REPORT VERSION v1 · internal calibration
LAST UPDATED 8 September 2026
SAMPLE 40 performances · 3 UYSF-certified judges
SCOPE competitive yogasana, controlled capture (1080p / 60fps)
STATUS indicative, not peer-reviewed · independent study in preparation
SYSTEM SEGA AI Judge System · Malaysian Patent Pending PI2026005219