Method
Segmentation. SAM 3 (facebook/sam3, text prompts chair, table, potted plant) on all 72 keyframes; each mask is associated with the reviewed object whose projected centre lies closest to it. Masks for the pendants come from the earlier shade study, with the suspension cord removed by a morphological opening. One frame per object was chosen by hand from the candidates (rationale in each inputs/<id>/input.json).
Reconstruction. Official SAM 3D Objects pipeline (pinned checkout, seed 42, 25 steps per stage, vertex colours, no texture baking) on an NVIDIA A40. The raw GLB per object is canonical: y-up, longest extent 1. The pipeline's predicted rotation/translation/scale are saved next to each mesh but are not applied: no convention we tested maps them onto the exported GLB consistently with the source camera, so they are reported, not used.
Placement. Position: the reviewed room document (centre X/Z; furniture rests on the ARKit floor plane, shades hang with their bottom at the reviewed fixture's lowest point and a drawn cable to the ceiling). Yaw: the angle that maximises silhouette overlap between the projected mesh and the SAM 3 mask in the source frame (free 5° search for chairs and the plant; the table keeps its long axis on the reviewed axis; proxy chairs take the reviewed facing). Size: uniform scales the whole mesh so its height matches the reviewed height (shades: diameter 0.42 m, a prior); box scales each world axis to the reviewed width, height, depth and can distort the shape. Matching a box is not a geometric validation.