This runs live in the page. Drag the bar or pick a step. Nothing here is a mock-up: the walls, rooms and dimensions are drawn from the same reviewed files the build itself reads.
Skip to the reconstruction you can move the camera in → four models, 104 solved photographs



Both carry Galaxy Z Fold5 EXIF dated 6 March 2026. There is no CAD file, no survey, no DWG. Everything downstream has to be recovered from paper that was photographed at an angle, under a window, with a shadow across one corner.
28 named spaces. Blue cast from the window, cast shadow at the lower left.
23 named spaces. Sharper sheet, heavier paper wrinkle through the middle.
Cleanup ran as a bake-off, and the option that looked most aggressive lost. Black-hat cleared the broad cast shadow but let too much neutral wrinkle texture through, so it was rejected as the final line source. RGB consensus separated black printed ink from blue illumination and most of the paper relief, and that is what shipped.
The file sizes say the same thing. The rejected black-hat PNG is 3,117,022 bytes against the selected consensus PNG at 349,719. Roughly nine times the noise.
Shadow gone, wrinkle kept. Rejected for being dirtier, not cleaner.
Ink coverage 25.6% down to 7.3% on floor one, 22.8% to 7.7% on floor two.
Every dropped pixel drawn back over the original, so the deletion is reviewable instead of trusted.
Optical-density threshold sweep: at 0.14 the pass began attenuating low-contrast one-pixel dimension ticks and dashed construction lines, so 0.10 was selected. Neutral-ink boost was measured and then disabled at 0.0, because that path could only add ink, never take it back.
Straight classical vision on a photographed plan. Flat-field correction to kill the window gradient, Hough transform for lines, max-pooled wall grid, a room classifier, then an SVG. Each stage was reasonable. The result was not the house.
These are kept because they are the reason the current pipeline separates what a label says from what geometry claims. That rule came out of watching this fail.
Illumination gradient removed. This part actually worked and survives in spirit today.
Finds every straight edge including hatching, dimension ticks and the sheet border. No idea which ones are walls.
Max-pooling made the walls thick enough to connect and thick enough to swallow the door openings.
Rooms invented where the grid closed a loop. Confident, coloured, and wrong in several places.
The wrong footprint, lifted into 3D. Extruding a bad plan gets you a bad plan you can orbit.
Nothing downstream is allowed to guess any more. A room exists because a person read the label on the sheet and signed off on it. A polygon exists only if its edges land on ink that survived cleanup. The two are stored separately, and the difference is written into the files rather than into a caption.
The reviewed schema carries 51 named spaces across 2 floors with 19 doors and openings. A separate wall graph carries 203 segments, 117 on the first floor and 86 on the second.
Only 3 polygons were ever promoted to geometry. The rule is written in the file itself: the label inventory is not promoted to geometry by implication.
Drawn in the photograph's own pixel coordinates, which is what source-aligned means here. Nothing was straightened to make them fit.
Every other draft polygon behaved like a bounding box or crossed a source wall under independent review. Those stayed analysis evidence and were never used.
Acceptance threshold: 75% of edge samples within uncertainty plus 2 px of preserved source linework. Kitchen and Great room overlap by 0 px². Shipped with this page: wall-graph.json, floor-01-wall-graph.svg, floor-02-wall-graph.svg, verified-schema.json, accepted-traces.json, accepted-traces-validation.json. Every figure above is counted from those files when the page is built, not typed in.
The full photo-to-geometry sequence as a standalone instrument, with its own camera and per-trace uncertainty readout. It is driven from this page rather than opened separately.
These are the trained Gaussian splats themselves, not renders of them. The camera orbits the house, so you cannot get lost — and it stops at the edge of the arc the photographs actually cover. Four models are here so the differences are yours to check rather than mine to describe: the photographic baseline, the first reconstructable generated clip, four clips fused into one model, and the orbit hop that pushed coverage past the corner of the house.
Nothing below this point is a picture standing in for a room. Every number attached to these models was measured on views the training never saw.
Step 1 — pick a model
Step 2 — load it
Drag to orbit, scroll to zoom, shift-drag to pan, R to reset. The viewer has preset viewpoints and a reset button of its own. Picking a different model above reloads it in place.
Every clip gets fed to a structure-from-motion solve under pinned intrinsics. If the cameras cannot be recovered from the frames, the clip is rejected no matter how it looks. The scoreboard is how many of 210 image pairs come back geometrically calibrated.
The strongest number on the board belongs to a clip that was thrown away, and the cheapest win belongs to a clip that was never generated.
The first guide that produced a solvable generation at all.
The middle arc of the r7 family, and the one that shipped.
Gated higher than the arc that shipped. Rejected anyway.
The best gate score in the project, guide and generated clip both.
Radial-only parallax the mapper cannot bootstrap. Caught before spending anything.
Six candidates were rendered and solved locally at no cost. Five passed at 21/21 registered, 0.73 to 0.82 px, zero watermark-configuration pairs. One failed outright. Only the survivors were allowed to become paid generations.
Looked fine. Solved as a single degenerate configuration on every pair.
The first generated clip that reconstructed at all.
Sharpest single arc. Narrow envelope: step off the path and it falls apart.
The strongest parallax signature in the project, and the first look past the corner.
There is no photograph of this house as built anywhere in the project. The baseline is the strongest prior generated clip, Kling 3.0 Pro image-to-video at $0.56, re-derived from scratch and reconstructing at 30.34 dB.
The donor repo had claimed 32.45 dB on the same data. That did not survive re-derivation and is recorded as unverified. The honest claim is the one that holds: generated video now reconstructs within 2.56 dB of the best prior generated clip, on a gate that rejected earlier attempts outright.
All seven were driven from the identical generated still with the same prompt and negative prompt, then screened by a local proxy: feature continuity, camera motion, source retention, cut safety. Thirteen runs across two scenes, $8.44 actual against a $9.00 budget.
The proxy ranked Veo first at 90.9 and Kling fifth at 85.3. Kling is the one that actually reconstructed. That is why the screening file says in its own limitations that the ranking is triage only and must not be presented as a reconstruction result.
Screening formula: 35% feature continuity, 25% camera motion, 25% source retention, 15% cut safety, over SIFT correspondences with RANSAC fundamental geometry. It never ran a structure-from-motion solve, so it could not see whether the generated 3D was right. Persistent 2D texture matches features even when the geometry underneath is wrong, which is exactly how Veo won a race it then lost.
Eight schema rooms and the exterior shell, each carrying the dimension string and review confidence straight off the reviewed plan, so what you are looking at and what the file says stay attached to each other. That link is the part worth keeping.
The images themselves are presentation plates, not measured space. They are useful for finding a room and checking its label against the plan, and they are the reason the work moved to real reconstruction: a plate cannot be walked around, and the section above can.
The Victorian pass is the clearest case of looks and geometry disagreeing. It is the best-looking exterior in the project and it was rejected, because the restyle moved the roofline and the reconstruction could not solve it.
Nothing here is cut together. A single recording drives the room change, fireplace intensity and flicker, window spill, the day to night slider, the counter practicals, an audio-reactive party mix, rain with two lightning strikes, then the exterior with a heading pan and a horizon lift.
Jump to any beat below. The values shown are the ones the script actually sends.
Lighting took three passes. At practical intensity 3.6 with room wash 2.2 the party's audio-reactive drive clipped the whole room to white-cyan for about six seconds. Dropping to saturation 0.7 and wash 1.6 stopped the clipping and washed everything pastel. The shipped values are saturation 0.9, practicals 2.6, wash 1.2, party mix 0.9. High saturation carries the colour while low intensity keeps the headroom.
The idea was sound on its face. Reproject a photograph a few degrees using its own depth, fill the sliver that opens behind the objects, repeat. Each step invents almost nothing, so a chain of them should walk a long way and stay mostly real. Eight chains were run at step sizes from 0.5° to a single 30° jump.
Every one of the six preregistered step sizes failed. Reaching 30° in ten steps leaves 27.0% of the frame traceable to the photograph. Reaching the same 30° in one jump leaves 66.5%. Small steps do invent less each time — and they invent it on top of what the last step invented.
One frame of the fused solve. Nothing in it was warped, filled or guessed.
Two thirds of it is still the photograph, and it still reads as the building.
Every individual step opened only 7% of the frame. There is no house left.
The obvious suspect was depth drifting as it is carried forward, so the 3° chain was run again with depth re-rendered from the splat at every step — a perfect oracle the real pipeline can afford. It moves the result from 27.0% to 33.4% and the frame is just as gone. The failure is not depth. It is that the same 7% hole opens every step and lands on ground the previous step already invented.
Two things decide that, and they are not the same thing. How much of the frame has to be invented is the labour. How big the largest single hole is decides whether a generative model gets one bounded region it can see the shape of, or a scatter of speckle it cannot. On r4_0031, at the far positive end of the arc, turning one way costs — of the frame and turning the other costs — — a gap of — points, which invites the tidy explanation that the occluders sit to one side of the house. Two more frames were measured and they do not support it: the cheaper direction is not the same direction for all three. It is a property of the frame, so it has to be measured per warp rather than assumed. Sliding the camera to the same place without rotating — the control — is worse than either, everywhere, on every frame, because rotation keeps the subject in the picture and translation slides it out.
The number under all of this: the photographs cover — of azimuth, —, and that is the entire capture. A full orbit of the property is 360°. Chaining was the mechanism that was supposed to cross that gap and it does not work; one direct warp reaches perhaps 30° past the last photograph while staying two-thirds real. Everything beyond that has no photograph to be warped from at any distance, which is not an inpainting problem — it is the generation problem branch A was opened on, and the arbiter has already measured what generated frames do to a reconstruction.
What this does not say: the warp itself is sound — pixels outside the hole survive every inpaint bit-identical, asserted at every step, and the run reproduces byte-for-byte. It does not say a better inpainter would look worse; it says a better inpainter would be inventing the same fraction of the frame, and the fraction is what was being measured. No reconstruction was trained on these frames. Costs $0.00: everything here ran locally. Numbers read live from dwi-chains.json and dwi-reach.json.
Fusing the R4 arc with three new clips put 104 of 104 images into a single model on the first try. 33,572 points at 0.889 px. Lateral camera coverage went from 6.5 to about 13.8 units.
It costs 1.34 dB on the R4 arc for roughly three times the view volume, and R4's worst endpoint got 1.19 dB better, because every new clip observes that region.
The walk-up view a person actually uses.
Both edges gained real edge content instead of invention.
Not a floater problem. No capture yet sees the near ground at a grazing angle.
A clip whose camera is low and pitched down while translating, so a driveway approach dolly at −20 to −25°. That is a new control family and a paid decision, so it is not in this build.
Also on the record, because leaving it out would be dishonest. One r5 clip completed and billed at $0.50625 with 175.6 s of inference, then the result store returned HTTP 500 on all 23 retrieval attempts over 24 minutes. It was excluded from the fusion rather than worked around. Request 019fd250-e52d-7713-b0f1-63e242925bc2 is still retryable if that store recovers.
Each one is a manifest with a per-file sha256 index, an authority boundary and a spend record.
Across two floors, 33 of them carrying a legible dimension string off the sheet.
Counted from the shipped wall graph, 117 plus 86, rather than asserted in prose.