Tennis3D lifter
A temporal transformer that estimates the ball centre in a court-centred, Z-up coordinate system, in metres, from timestamped 2D ball observations and court landmarks from one camera. The intended input is a continuous elevated broadcast shot of a singles tennis match.
This repository contains the compact EMA inference weights and their integrity manifest. The model takes numeric observations; the Tennis3D application supplies ball detection, court calibration, reconstruction and playback.
| Property | Value |
|---|---|
| Weight file | tennis3d-lifter-stage4.pt |
| File SHA-256 | 9cfc12371dcc698bffb5e0c25a483dbc7143c08142a47bef5eaf5611201425fa |
| File size | 7,061,198 bytes |
| Architecture | stage3-v1 transformer, width 176, four blocks, eleven heads |
| Trainable parameters | 1,759,871 |
| Weights | EMA, selected at update 83,334 by minimum corrupted reserved-validation rally-mean XYZ error |
| Training | Seed 42, batch 48, 83,334 updates |
| Source checkpoint SHA-256 | 589ea17734c0643f44393692ed9a197e90f897481fe207b8387003d15e350e67 |
| Tensor SHA-256 | 2d692487f965d1e798105cd80420d70b2d14417fa18387453a6c3f23eeb07644 |
| Checkpoint schema | 1.0, world scale 10 metres, EMA weights |
The export contains model configuration and EMA tensors. It omits optimizer, sampler and random-number state. All 71 tensors matched the selected training checkpoint, and a 250-frame CPU probe produced bit-identical predictions.
Use
With the Tennis3D application installed:
tennis3d setup
tennis3d reconstruct clip.mp4 --output my-run
The application downloads the weight file and verifies its SHA-256. The compact
file can also be loaded by tennis3d.lifting.checkpoint.load_lifter(path, device).
It is a torch.save mapping with schema_version, world_scale_m, config.model
and EMA tensors named module.<state_key>. Load with weights_only=True.
Training data
The lifter was trained from random initialization using
N-il/tennis-3d-synthetic,
revision c2e782053694cfde305362e670ebbb4f4e70884c, manifest SHA-256
b32fe9cc97226c9d67468bcf7724f7d3ffc8924541798796744f9ba62bba6707.
No real footage was used to train this lifter. Checkpoint selection used reserved
validation, rather than the untouched test split.
The simulation and lifting software build on UpliftingTableTennis. Original source attribution and notices remain with the Tennis3D software.
Final synthetic test
The selected EMA was evaluated on the frozen test split at the dataset revision listed above. It contains 5,000 physical rallies, 10,000 camera views and 1,829,263 source frames per input mode. Clean and saved-corrupted inputs use the same 250-frame windows, stride 125 and centre-weighted merging. Checkpoint selection was frozen before test evaluation.
The primary error is the mean over physical rallies of Euclidean ball-centre XYZ error, pooling each rally's declared view and frame records. It includes every test frame before export filtering. Frame RMSE is weighted over all frames.
| Measure | Clean inputs | Saved-corrupted inputs |
|---|---|---|
| Primary rally-mean XYZ error, m | 0.075494 | 0.107762 |
| Frame-weighted mean XYZ error, m | 0.067567 | 0.098652 |
| Frame RMSE, m | 0.138501 | 0.187418 |
| Frame p95 Euclidean error, m | 0.229068 | 0.272969 |
| Exported frames / all source frames | 1,766,043 / 1,829,263 | 1,734,760 / 1,829,263 |
| Exported coverage | 96.54% | 94.83% |
These errors measure accuracy against simulator truth for this pinned synthetic test. Export coverage has a fixed denominator and does not change the primary error. One training seed was used. This release has no matched convergence study; the earlier compact-validation search score is not a comparable test result.
Corrupted synthetic predictions were also fitted to the tennis-flight ODE using saved simulator contact segmentation and unchanged fit defaults. There were 55,862 fitted segments, 33,611 excluded and zero failed. Mean segment RMSE was 0.036003 m. Its Spearman correlation with true segment XYZ RMSE was 0.294511, which is weak discrimination. A small ODE residual is a consistency diagnostic and does not prove a trajectory is accurate.
Current real-footage diagnostics
The current convenience gallery contains twenty requested ranges from five
professional singles-match recordings, all prepared at 1920x1080. This gallery
replaces the earlier footage cohort. Every one of its 10,749 source frames remains
in the results. Three clips have status ok, seventeen partial, and none
failed. Cameras are valid on 10,701 frames. Raw lifting covers 10,651;
quality-filtered export retains 10,134, or 94.28% of all source frames.
The requested ranges retain 26 frames after two hard cuts and five frames under a fading graphic. The cut detector missed those transitions. The inventory has one confirmed serve and no confirmed lob or ball-occlusion example; sampled category review does not prove an unlisted event is absent. These omissions and five convenience-selected recordings limit generalization claims.
Detection, tracking and camera calibration ran once and were shared by the four evaluated lifters. The other-model gallery comparisons therefore replace lifting on those cached inputs, rather than repeat the full perception pipeline. Independent real XYZ, ball image-point and camera truth are unavailable. Reprojection against the same detector observations is a dependent diagnostic. Detector training overlap is unknown, and exact saved timestamps do not establish original broadcaster exposure timing.
Paired low-resolution copies
The separate 256x144 cohort is derived frame for frame from these twenty clips.
It preserves the same 10,749 source rows and their integer presentation
timestamps. Production has four partial and sixteen failed clips, with 1,257
valid-camera rows, 1,028 raw-covered rows and 988 exported rows, or 9.19% export
coverage. The failed clips have no accepted camera. The calibration reason
insufficient_landmarks covers too few landmarks or insufficient court-wide
spatial support, so it does not mean that every rejected frame had too few points.
All inputs and failure reasons remain in the denominator.
| Production export availability on the same source row | Rows |
|---|---|
| Both resolutions | 937 |
| 1920x1080 only | 9,197 |
| 256x144 only | 51 |
| Neither | 564 |
This comparison changes device as well as resolution. Full-resolution perception used Intel GPU float32; low-resolution perception used an RTX 2080 Ti with CUDA float32 after two Intel device faults. A one-clip CUDA versus CPU probe preserved the frame and ball-track arrays, but six court-landmark values differed by up to six pixels and camera validity differed on five of 360 rows. The devices are not established as numerically equivalent. No XYZ values are ranked between resolutions. The low-resolution automatic report has six fitted, 204 excluded and zero failed ODE segments; no manual-contact score was computed for it.
Approximate manual contacts
The high-resolution gallery has 312 frozen manual event references, 154 hits and 158 bounces. A velocity-jump heuristic derives unclassified contact episodes from predicted XYZ. It is not a trained hit-versus-bounce classifier. Global one-to-one matching uses a primary tolerance of +/-100 ms and sensitivity checks at 33.333, 66.667 and 200 ms. All 312 references remain in the denominator.
| Prediction cohort | TP | FP | FN | Precision | Recall | F1 |
|---|---|---|---|---|---|---|
| Raw | 191 | 151 | 121 | 0.558480 | 0.612179 | 0.584098 |
| Exported | 198 | 157 | 114 | 0.557746 | 0.634615 | 0.593703 |
Hit-reference recall is 127/154 raw and 132/154 exported; bounce-reference recall is 64/158 raw and 66/158 exported. These are recall by reference type, not typed classification precision. Exported F1 at 33.333, 66.667 and 200 ms is 0.389805, 0.536732 and 0.662669. Selected-frame identities and timestamps are exact, while blur and exposure uncertainty in physical contact timing remain uncalibrated. The matching tolerance is an engineering choice.
Real trajectory consistency
Raw and exported trajectories are fitted separately with automatic contacts or manual event marks. Manual-event segmentation excludes a predeclared 100 ms around contacts, plus the unchanged 10 ms fit guard. This is not calibrated annotator uncertainty. Fitted, excluded and failed segments remain separate, including outliers.
| Contact-report cohort | Fitted | Excluded | Failed | Mean segment RMSE, m | Median, m | p95, m | Max, m |
|---|---|---|---|---|---|---|---|
| Raw / automatic | 215 | 134 | 0 | 0.227890 | 0.127753 | 0.620176 | 1.690066 |
| Raw / manual events | 259 | 77 | 0 | 83.102527 | 0.076213 | 2.138674 | 5851.474348 |
| Exported / automatic | 213 | 149 | 0 | 0.226206 | 0.130691 | 0.604897 | 1.690066 |
| Exported / manual events | 256 | 93 | 0 | 0.266537 | 0.074896 | 1.286740 | 2.972269 |
The raw manual-event cohort has an extreme 5,851.47 m maximum segment RMSE and 83.10 m mean, despite a 0.0762 m median. No optimizer failure was recorded in these four cohorts; successful fitting does not make an implausible curve accurate. The exported manual-event maximum is 2.9723 m. All residuals measure consistency with the predicted curve, not independent real XYZ accuracy or unique spin identification.
The pipeline's elapsed-time automatic report remains separate from these absolute-source-time contact reports. It has 215 fitted and 148 excluded segments, with 7,395 scored frames. The raw automatic contact report has 215 fitted and 134 excluded segments. Contact extraction and exact-threshold floating-point behavior on the two time axes differ. Both original reports were retained; neither was overwritten or pooled with the other.
Comparison models
The full pinned synthetic test was also scored for a transformer trained with fresh AdamW, a two-layer bidirectional LSTM trained with the same AdamW recipe, and an additional-training candidate continued from production. All use seed 42.
| Model | Updates | Clean rally-mean XYZ, m | Corrupted rally-mean XYZ, m | Current-gallery exported contact F1 at 100 ms |
|---|---|---|---|---|
| Production transformer, Muon recipe | 83,334 | 0.075494 | 0.107762 | 0.593703 |
| Transformer, AdamW recipe | 83,334 | 0.084869 | 0.123224 | 0.571827 |
| LSTM, AdamW recipe | 83,334 | 0.092103 | 0.130666 | 0.554817 |
| Longer-trained production candidate | 166,668 | 0.072187 | 0.100460 | 0.593939 |
LSTM versus the AdamW transformer changes architecture. AdamW transformer versus production changes an optimizer recipe, including learning rate, radial braking and patience, rather than the algorithm alone. The continuation has 83,334 extra updates and is not a matched-budget ablation. One run per model provides no uncertainty or seed-variance estimate. Its better synthetic error and almost equal current-gallery exported contact F1 do not establish a real XYZ improvement. The published weights remain the original 83,334-update production checkpoint; no comparison checkpoint was promoted.
Compute and execution scope
Training used one RTX 2080 Ti, seed 42, batch 48 and 83,334 updates. It accepted 4,000,032 windows from 4,002,322 sampler draws, with 2,290 rejected crops. These are exposure counters, not unique data coverage. Measured training-workload time was 17.82 hours; scheduler allocation was 19.29 hours.
Full synthetic evaluation used 85,399 allocated GPU seconds. Its first allocation held the GPU while doing serial CPU physics fitting. A continuation completed missing predictions and moved the remaining fits to a CPU-only phase. Those 7,791 CPU fits took 42,434.68 wall seconds with four workers. These allocation figures are whole-pipeline costs, not GPU inference or kernel-time benchmarks.
The selected export passed Linux CPU and NVIDIA CLI/application execution checks. The NVIDIA check used one RTX 2080 Ti, driver 575.64.03, PyTorch 2.10.0+cu126 and CUDA 12.6 on one 120-frame 1080p clip. It checked source-time preservation, cancellation/resume, stage reuse and export. Windows NVIDIA execution has not been qualified. These checks do not establish general footage accuracy or support for every GPU.
Missing or rejected estimates remain gaps with recorded reasons. Cropped footage, camera cuts, low viewpoints, unseen balls and calibration errors can reduce coverage or produce wrong estimates. The reported results describe the original 83,334-update checkpoint in this repository. Comparison results are listed separately above. Their timings differ in node, CPU allocation, cache and exposure, and are not model speed benchmarks.
License
The lifter weights are GPL-3.0-only. See LICENSE.