Visible cloth geometry fits. Template identity still moves.
A five-step temporal DDPM reconstruction of one SAM2-segmented trouser trajectory. The videos overlay each predicted surface in blue on the observed point cloud in grey.
What the run says
Most predictions lie within this cloud-to-surface RMSE band.
M1 074 diverges on frame 6 despite fitting well elsewhere.
Twenty-four reconstructions after checkpoint loading on one GH200.
Per-frame visible-surface RMSE
| Frame | Pants 047 | M1 074 | M1 101 |
|---|---|---|---|
| 0 | 8.46 mm | 7.49 mm | 11.94 mm |
| 1 | 8.54 mm | 7.28 mm | 6.36 mm |
| 2 | 8.85 mm | 8.49 mm | 6.71 mm |
| 3 | 8.21 mm | 6.82 mm | 6.48 mm |
| 4 | 9.34 mm | 7.07 mm | 6.60 mm |
| 5 | 9.50 mm | 8.07 mm | 6.67 mm |
| 6 | 8.21 mm | 14.34 mm | 6.55 mm |
| 7 | 8.02 mm | 7.01 mm | 6.16 mm |
Green marks the lowest robust fit score in each frame. RMSE itself is shown for readability; the ranking also penalizes unsupported predicted surface.
Interpretation
The estimator captures the observed trouser silhouette and coarse folds surprisingly well. The main limitation is not simply point fit: several templates explain the visible cloud, hidden geometry is unconstrained, and plausible-looking wrinkles can be hallucinated.
These numbers are not full-mesh error. The recording has no ground-truth cloth mesh. They measure how closely the prediction explains SAM2-selected points, so calibration residuals, depth noise, and segmentation boundaries all contribute.
Provenance
CheckpointCloth-splatters/dexgarmentlab-lift-20260822-state-est-gps-tf2
Estimator commit32b7c5aab86e6ed919e3ca512f9f8c3f7fcd5e1f
Recording SHA-256409964853a39c046fa4caa193c852f93c8a3740f5d04ae72508d7b92ddb6b7be
Protocol
Five inference steps; seed equals frame index; frame 0 single-frame, frames 1–7 use the preceding segmented cloud as temporal context; 15% trimmed cloud-to-surface MSE plus 0.25× surface-to-cloud MSE for ranking.