The SEGA AI Judge System · Patent Pending PI2026005219 · Validation & Error Report Evidence Centre →
Home / AI Judge System / Validation & Error Report
◈ Transparency report

AI Judge — Validation & Error Report

How accurate is the SEGA AI Judge, how do we know, and where are its limits? This report sets out the study design, the human-vs-AI agreement measured so far, the errors we track, and — just as importantly — an honest account of what has not yet been proven.

⚠️
Status: internal calibration (v1). The figures below come from a small pre-tournament calibration — not an independent, peer-reviewed study. They are indicative of how the system behaves, not a final accuracy claim. A larger, independently-conducted study is in preparation and will be published here when complete. We publish the current state openly rather than waiting, and we clearly label what is measured versus illustrative.
§1 · Study design

What we tested, and how

During pre-tournament calibration, the SEGA AI Judge scored 40 test performances alongside a panel of 3 UYSF-certified human judges. The same performance videos were given to both the AI and the panel, and the two sets of scores were compared. The AI also re-scored each video twice, was tested from front and side camera angles, and was run across a simulated competition day to check for drift. This is a measurement-agreement study: does the AI agree with expert humans, and is it stable?

§2 · Results — human panel vs SEGA AI

Agreement & error

MetricHuman panelSEGA AI Judge
Mean absolute deviation across 40 performances±0.87 pts±0.24 pts
Inter-run consistency (same video, scored twice)±0.62 pts±0.00 pts (deterministic)
Viewing-angle sensitivity (side vs front)up to ±1.4 pts±0.15 pts
Fatigue drift (hour 1 vs hour 6)±0.45 pts±0.00 pts
Correlation with final medal outcome0.910.97

Source: internal pre-tournament calibration (40 performances, 3 judges). Same figures as the Evidence Centre. Indicative, not peer-reviewed — see limitations below.

§3 · How to read these numbers

In plain words

Mean absolute deviation (error) — how far a score typically lands from the "true" panel consensus. Smaller is tighter. The AI's ±0.24 means it usually sits within a quarter-point of the reference.
Deterministic (±0.00) — given the same video, the AI returns the exact same score every time. Human panels naturally vary a little between viewings; the AI does not.
Viewing-angle sensitivity — how much the score changes with camera position. Lower is better; the AI is far less angle-dependent than the eye.
Correlation with medal outcome — how closely the scores track who actually medalled (1.0 = perfect). Both are high; the AI is marginally higher here.
§4 · Confidence & honesty

What these figures can and can't say

With only 40 performances, the confidence interval around every number above is wide — a second sample could shift them. These results show the system is internally consistent and closely aligned with expert judgement on this set. They do not yet prove accuracy at scale, across disciplines, or in the hands of independent researchers. Any single-number "accuracy %" would overstate what a 40-sample calibration can support, so we don't publish one.

§5 · Known limitations

Where the system needs care

Small, internal sample. 40 performances scored against 3 judges, run by SEGA — not an independent study. Numbers are indicative only.
Single discipline & setting. Calibrated on competitive yogasana under controlled capture; it is not validated for other movement forms or uncontrolled home video.
Capture-dependent. Accuracy relies on the specified camera setup (1080p / 60fps, fixed tripod, adequate lighting, clear background). Poor capture degrades the measurement.
Measures the measurable. The AI scores execution, stability and difficulty. Interpretive qualities and any ambiguous case are escalated to certified Master Judges — the AI is a reference layer, not the final ruling.
Not medical or diagnostic. Scores describe performance only; they are not a health, injury or fitness assessment.
§6 · What "verified" means here

The claim we do make

When a result is marked "AI-Judge verified," it means the performance was scored by the SEGA AI Judge System under the standard capture protocol, the score is stored in a tamper-evident record, and ambiguous cases were referred to human Master Judges. It does not mean the score has been independently audited by a third party — that is the purpose of the study below.

§7 · Roadmap to independent validation

How this report grows up

Phase 1 · done

Internal calibration

40 performances vs 3 judges — the figures on this page.

Phase 2 · in preparation

Expanded internal study

Several hundred performances, more judges, per-dimension error, published confidence intervals.

Phase 3 · planned

Independent study

A third-party or academic evaluation with full methodology, dataset description and a DOI in the Research index.

REPORT VERSION v1 · internal calibration
LAST UPDATED 8 September 2026
SAMPLE 40 performances · 3 UYSF-certified judges
SCOPE competitive yogasana, controlled capture (1080p / 60fps)
STATUS indicative, not peer-reviewed · independent study in preparation
SYSTEM SEGA AI Judge System · Malaysian Patent Pending PI2026005219
Full Evidence Centre →   Why SEGA AI Judge?