This runs live in the page. Drag the bar or pick a step. Nothing here is a mock-up: the walls, rooms and dimensions are drawn from the same reviewed files the build itself reads.
Skip to the reconstruction you can move the camera in → four models, 104 solved photographs



Both carry Galaxy Z Fold5 EXIF dated 6 March 2026. There is no CAD file, no survey, no DWG. Everything downstream has to be recovered from paper that was photographed at an angle, under a window, with a shadow across one corner.
28 named spaces. Blue cast from the window, cast shadow at the lower left.
23 named spaces. Sharper sheet, heavier paper wrinkle through the middle.
Cleanup ran as a bake-off, and the option that looked most aggressive lost. Black-hat cleared the broad cast shadow but let too much neutral wrinkle texture through, so it was rejected as the final line source. RGB consensus separated black printed ink from blue illumination and most of the paper relief, and that is what shipped.
The file sizes say the same thing. The rejected black-hat PNG is 3,117,022 bytes against the selected consensus PNG at 349,719. Roughly nine times the noise.
Shadow gone, wrinkle kept. Rejected for being dirtier, not cleaner.
Ink coverage 25.6% down to 7.3% on floor one, 22.8% to 7.7% on floor two.
Every dropped pixel drawn back over the original, so the deletion is reviewable instead of trusted.
Optical-density threshold sweep: at 0.14 the pass began attenuating low-contrast one-pixel dimension ticks and dashed construction lines, so 0.10 was selected. Neutral-ink boost was measured and then disabled at 0.0, because that path could only add ink, never take it back.
Straight classical vision on a photographed plan. Flat-field correction to kill the window gradient, Hough transform for lines, max-pooled wall grid, a room classifier, then an SVG. Each stage was reasonable. The result was not the house.
These are kept because they are the reason the current pipeline separates what a label says from what geometry claims. That rule came out of watching this fail.
Illumination gradient removed. This part actually worked and survives in spirit today.
Finds every straight edge including hatching, dimension ticks and the sheet border. No idea which ones are walls.
Max-pooling made the walls thick enough to connect and thick enough to swallow the door openings.
Rooms invented where the grid closed a loop. Confident, coloured, and wrong in several places.
The wrong footprint, lifted into 3D. Extruding a bad plan gets you a bad plan you can orbit.
Nothing downstream is allowed to guess any more. A room exists because a person read the label on the sheet and signed off on it. A polygon exists only if its edges land on ink that survived cleanup. The two are stored separately, and the difference is written into the files rather than into a caption.
The reviewed schema carries 51 named spaces across 2 floors with 19 doors and openings. A separate wall graph carries 203 segments, 117 on the first floor and 86 on the second.
Only 3 polygons were ever promoted to geometry. The rule is written in the file itself: the label inventory is not promoted to geometry by implication.
Drawn in the photograph's own pixel coordinates, which is what source-aligned means here. Nothing was straightened to make them fit.
Every other draft polygon behaved like a bounding box or crossed a source wall under independent review. Those stayed analysis evidence and were never used.
Acceptance threshold: 75% of edge samples within uncertainty plus 2 px of preserved source linework. Kitchen and Great room overlap by 0 px². Shipped with this page: wall-graph.json, floor-01-wall-graph.svg, floor-02-wall-graph.svg, verified-schema.json, accepted-traces.json, accepted-traces-validation.json. Every figure above is counted from those files when the page is built, not typed in.
The full photo-to-geometry sequence as a standalone instrument, with its own camera and per-trace uncertainty readout. It is driven from this page rather than opened separately.
These are the trained Gaussian splats themselves, not renders of them. The camera orbits the house, so you cannot get lost — and it stops at the edge of the arc the photographs actually cover. Four models are here so the differences are yours to check rather than mine to describe: the photographic baseline, the first reconstructable generated clip, four clips fused into one model, and the orbit hop that pushed coverage past the corner of the house.
Nothing below this point is a picture standing in for a room. Every number attached to these models was measured on views the training never saw.
Step 1 — pick a model
Step 2 — load it
Drag to orbit, scroll to zoom, shift-drag to pan, R to reset. The viewer has preset viewpoints and a reset button of its own. Picking a different model above reloads it in place.
A number picked out of the air gave a −55° preset of pure smear. A sharpness metric said three of the four models hold out to ±80°; at +80° there is no building in the frame at all. A collapsing splat does not go soft — the gaussians stretch into long streaks and the foliage shatters into confetti, and both of those raise sharpness, so the metric scored the collapse as healthy. It overstated every model and understated none, which is what measuring the failure mode instead of the subject looks like. Solved camera positions replaced it and are honest, but they answer a different question — where a photograph stands, not how far the solve carries — and they are symmetric by construction when this property is not.
So the limit was set by looking. 52 frames were rendered through this viewer with the panel hidden, four models across thirteen azimuths at 10° steps, and each was marked clear, degraded or gone. The marks are recorded per frame, so any one of them can be argued with by pointing at the frame.

Line the four models up by how much camera arc each solve was given and the return collapses monotonically: 5.06× readable degrees per camera degree, then 1.75×, then 1.01×, then 0.83×. A 10× spread in capture width bought 20° of extra readable view, and the two widest captures read no wider than the arcs their own cameras stand on. Adding generated frames further round the house is buying arc on the wrong side of that knee — the next gain has to come from making the pixels we already have count for more, not from standing in more places.
Feeding generated frames into a reconstruction costs −1.43 dB against the photograph-only baseline, measured on held-out views. We also know why: a warped frame is two things stitched together — pixels resampled from a real photograph, and pixels a model invented where nothing was visible — and the warp knows which is which, per pixel. So the obvious repair is to hand the trainer the resampled half and mark the rest unknown.
That repair has one load-bearing premise: that unknown is a thing the trainer can be told. The only candidate is a transparent input pixel, and the manual sentence for it reads either way. Three instruments on this project have already measured something real and had the answer used for a question it wasn’t answering, so this one got tested instead of assumed — a rectangle punched to transparent in half the training views, the other half left intact, both arms otherwise identical.

This is strictly worse than feeding the invented pixel through. A wrong colour is one bad observation competing with good ones; a wrong emptiness deletes geometry that other views got right. Masking a warped frame down to its honest pixels would carve the disoccluded regions out of the house. Two short training runs and one rectangle cost about four minutes and stopped a full experiment that would have come back negative and been written up as confidence weighting doesn’t help — which would have been false. Sealed as refuted in alpha-semantics-20260811-r1; the same idea still has a coarser form left, selecting whole frames by how much of them was invented.
Every clip gets fed to a structure-from-motion solve under pinned intrinsics. If the cameras cannot be recovered from the frames, the clip is rejected no matter how it looks. The scoreboard is how many of 210 image pairs come back geometrically calibrated.
The strongest number on the board belongs to a clip that was thrown away, and the cheapest win belongs to a clip that was never generated.
The first guide that produced a solvable generation at all.
The middle arc of the r7 family, and the one that shipped.
Gated higher than the arc that shipped. Rejected anyway.
The best gate score in the project, guide and generated clip both.
Radial-only parallax the mapper cannot bootstrap. Caught before spending anything.
Six candidates were rendered and solved locally at no cost. Five passed at 21/21 registered, 0.73 to 0.82 px, zero watermark-configuration pairs. One failed outright. Only the survivors were allowed to become paid generations.
Looked fine. Solved as a single degenerate configuration on every pair.
The first generated clip that reconstructed at all.
Sharpest single arc. Narrow envelope: step off the path and it falls apart.
The strongest parallax signature in the project, and the first look past the corner.
There is no photograph of this house as built anywhere in the project. The baseline is the strongest prior generated clip, Kling 3.0 Pro image-to-video at $0.56, re-derived from scratch and reconstructing at 30.34 dB.
The donor repo had claimed 32.45 dB on the same data. That did not survive re-derivation and is recorded as unverified. The honest claim is the one that holds: generated video now reconstructs within 2.56 dB of the best prior generated clip, on a gate that rejected earlier attempts outright.
All seven were driven from the identical generated still with the same prompt and negative prompt, then screened by a local proxy: feature continuity, camera motion, source retention, cut safety. Thirteen runs across two scenes, $8.44 actual against a $9.00 budget.
The proxy ranked Veo first at 90.9 and Kling fifth at 85.3. Kling is the one that actually reconstructed. That is why the screening file says in its own limitations that the ranking is triage only and must not be presented as a reconstruction result.
Screening formula: 35% feature continuity, 25% camera motion, 25% source retention, 15% cut safety, over SIFT correspondences with RANSAC fundamental geometry. It never ran a structure-from-motion solve, so it could not see whether the generated 3D was right. Persistent 2D texture matches features even when the geometry underneath is wrong, which is exactly how Veo won a race it then lost.
Eight schema rooms and the exterior shell, each carrying the dimension string and review confidence straight off the reviewed plan, so what you are looking at and what the file says stay attached to each other. That link is the part worth keeping.
The images themselves are presentation plates, not measured space. They are useful for finding a room and checking its label against the plan, and they are the reason the work moved to real reconstruction: a plate cannot be walked around, and the section above can.
The Victorian pass is the clearest case of looks and geometry disagreeing. It is the best-looking exterior in the project and it was rejected, because the restyle moved the roofline and the reconstruction could not solve it.
Nothing here is cut together. A single recording drives the room change, fireplace intensity and flicker, window spill, the day to night slider, the counter practicals, an audio-reactive party mix, rain with two lightning strikes, then the exterior with a heading pan and a horizon lift.
Jump to any beat below. The values shown are the ones the script actually sends.
Lighting took three passes. At practical intensity 3.6 with room wash 2.2 the party's audio-reactive drive clipped the whole room to white-cyan for about six seconds. Dropping to saturation 0.7 and wash 1.6 stopped the clipping and washed everything pastel. The shipped values are saturation 0.9, practicals 2.6, wash 1.2, party mix 0.9. High saturation carries the colour while low intensity keeps the headroom.
The idea was sound on its face. Reproject a photograph a few degrees using its own depth, fill the sliver that opens behind the objects, repeat. Each step invents almost nothing, so a chain of them should walk a long way and stay mostly real. Eight chains were run at step sizes from 0.5° to a single 30° jump.
Every one of the six preregistered step sizes failed. Reaching 30° in ten steps leaves 27.0% of the frame traceable to the photograph. Reaching the same 30° in one jump leaves 66.5%. Small steps do invent less each time — and they invent it on top of what the last step invented.
One frame of the fused solve. Nothing in it was warped, filled or guessed.
Two thirds of it is still the photograph, and it still reads as the building.
Every individual step opened only 7% of the frame. There is no house left.
The obvious suspect was depth drifting as it is carried forward, so the 3° chain was run again with depth re-rendered from the splat at every step — a perfect oracle the real pipeline can afford. It moves the result from 27.0% to 33.4% and the frame is just as gone. The failure is not depth. It is that the same 7% hole opens every step and lands on ground the previous step already invented.
Two things decide that, and they are not the same thing. How much of the frame has to be invented is the labour. How big the largest single hole is decides whether a generative model gets one bounded region it can see the shape of, or a scatter of speckle it cannot. On r4_0031, at the far positive end of the arc, turning one way costs — of the frame and turning the other costs — — a gap of — points, which invites the tidy explanation that the occluders sit to one side of the house. Two more frames were measured and they do not support it: the cheaper direction is not the same direction for all three. It is a property of the frame, so it has to be measured per warp rather than assumed. Sliding the camera to the same place without rotating — the control — is worse than either, everywhere, on every frame, because rotation keeps the subject in the picture and translation slides it out.
The number under all of this: the photographs cover — of azimuth, —, and that is the entire capture. A full orbit of the property is 360°. Chaining was the mechanism that was supposed to cross that gap and it does not work, so every pose has to be one direct warp from a real photograph — and how far one of those actually reaches is a question this curve could not answer, because it stopped at 30° and because the measurement under it was wrong. 09c corrects it. Everything past the reach it establishes has no photograph to be warped from at any distance, which is not an inpainting problem — it is the generation problem branch A was opened on, and the arbiter has already measured what generated frames do to a reconstruction.
What this does not say: the warp itself is sound — pixels outside the hole survive every inpaint bit-identical, asserted at every step, and the run reproduces byte-for-byte. It does not say a better inpainter would look worse; it says a better inpainter would be inventing the same fraction of the frame, and the fraction is what was being measured. No reconstruction was trained on these frames. Costs $0.00: everything here ran locally. Numbers read live from dwi-chains.json and dwi-reach.json.
The curve gave itself away before anything else did. On r4_0031 a 120° orbit scored — points cheaper than a 90° one. Cost cannot fall as the camera moves further from the only photograph of the subject, so the measurement was supplying pixels it had no right to. It was: a front-only capture has no back of the house in its depth map, so nothing occludes the facade when it is reprojected into a camera on the far side. The facade arrives again, inverted and see-through, and the metric counts it as real. Three other explanations were tested first and all three failed — sky counted as supplied (it is 0.0% of every frame), depth noise (median-filtering the source makes it slightly worse), and sampling cracks (a fatter splat closes them by smearing the house away). The fix is one test: a source pixel may only supply a destination pixel if its surface still faces the destination camera.
Backfacing material is being supplied here, and nothing in the metric objects to it.
This is why it survived a round. At 30° the defect is not visible — it is a few points of cost.
The camera is behind the house. This scores cheaper than the 90° view of the same building.
What a front-only capture actually knows about the far side, which is almost nothing.
Which raises the obvious question about everything above this section, so it was re-measured rather than argued: does the chaining verdict survive its own correction? Both decisive arms were run again with the facing test on. One 30° warp keeps — of the frame traceable to the photograph, against — before. Ten 3° warps to the same place keep —, against —. Every number in 09 and 09b moves down — and the gap between them widens, from — to — points. The correction makes the falsification stronger. The plate figures above still read as they were sealed, which is why they are labelled as measured rather than quietly restated.
What this does not say: it does not say the reach numbers are precise. Invented-pixel fraction is a labour estimate, not a usability bar, and the section above already showed it moving ten points on a free parameter of its own measurement. It says something narrower and harder to argue with — 60% of the orbit has nothing to warp from, and that is the part no better inpainter, warper or threshold touches. Costs $0.00: everything here ran locally against the sealed r2 geometry, read-only. Numbers read live from dwi-far-reach.json, generated from the sealed r3 manifest; the four plates above carry their payload hashes in dwi-facing-plates.json.
Fusing the R4 arc with three new clips put 104 of 104 images into a single model on the first try. 33,572 points at 0.889 px. Lateral camera coverage went from 6.5 to about 13.8 units.
It costs 1.34 dB on the R4 arc for roughly three times the view volume, and R4's worst endpoint got 1.19 dB better, because every new clip observes that region.
The walk-up view a person actually uses.
Both edges gained real edge content instead of invention.
Not a floater problem. No capture yet sees the near ground at a grazing angle.
A clip whose camera is low and pitched down while translating, so a driveway approach dolly at −20 to −25°. That is a new control family and a paid decision, so it is not in this build.
Also on the record, because leaving it out would be dishonest. One r5 clip completed and billed at $0.50625 with 175.6 s of inference, then the result store returned HTTP 500 on all 23 retrieval attempts over 24 minutes. It was excluded from the fusion rather than worked around. Request 019fd250-e52d-7713-b0f1-63e242925bc2 is still retryable if that store recovers.
Each one is a manifest with a per-file sha256 index, an authority boundary and a spend record.
Across two floors, 33 of them carrying a legible dimension string off the sheet.
Counted from the shipped wall graph, 117 plus 86, rather than asserted in prose.