UniClothDiff · real-world evaluation · 4 September 2026

Visible cloth geometry fits. Template identity still moves.

A five-step temporal DDPM reconstruction of one SAM2-segmented trouser trajectory. The videos overlay each predicted surface in blue on the observed point cloud in grey.

8 real frames3 mesh hypothesesGH200 inferenceNo mesh ground truth

What the run says

6–9 mm
Typical visible fit

Most predictions lie within this cloud-to-surface RMSE band.

14.34 mm
Largest observed failure

M1 074 diverges on frame 6 despite fitting well elsewhere.

4.12 s
Inference loop

Twenty-four reconstructions after checkpoint loading on one GH200.

Per-frame visible-surface RMSE

FramePants 047M1 074M1 101
08.46 mm7.49 mm11.94 mm
18.54 mm7.28 mm6.36 mm
28.85 mm8.49 mm6.71 mm
38.21 mm6.82 mm6.48 mm
49.34 mm7.07 mm6.60 mm
59.50 mm8.07 mm6.67 mm
68.21 mm14.34 mm6.55 mm
78.02 mm7.01 mm6.16 mm

Green marks the lowest robust fit score in each frame. RMSE itself is shown for readability; the ranking also penalizes unsupported predicted surface.

Interpretation

The estimator captures the observed trouser silhouette and coarse folds surprisingly well. The main limitation is not simply point fit: several templates explain the visible cloud, hidden geometry is unconstrained, and plausible-looking wrinkles can be hallucinated.

These numbers are not full-mesh error. The recording has no ground-truth cloth mesh. They measure how closely the prediction explains SAM2-selected points, so calibration residuals, depth noise, and segmentation boundaries all contribute.

Provenance

Checkpoint
Cloth-splatters/dexgarmentlab-lift-20260822-state-est-gps-tf2

Estimator commit
32b7c5aab86e6ed919e3ca512f9f8c3f7fcd5e1f

Recording SHA-256
409964853a39c046fa4caa193c852f93c8a3740f5d04ae72508d7b92ddb6b7be

Protocol
Five inference steps; seed equals frame index; frame 0 single-frame, frames 1–7 use the preceding segmented cloud as temporal context; 15% trimmed cloud-to-surface MSE plus 0.25× surface-to-cloud MSE for ranking.