There is no photograph of this building anywhere in this project and there cannot be one. It has not been built. It exists as two sheets of paper. The traced plan gives room outlines and those give massing, but massing has no siding, no brick, no roof pitch and no light in the windows. So the plan and a written style specification went to an image model, and the front elevation below is what came back.
Everything below descends from that one image: every video clip, every solved frame, every reconstruction. They inherit what it got right and what it got wrong.
What the model was told it could not change, written down before any frame was generated: 2 stories, 4 porch columns, 3 balcony bays, a five-sided brick bay tower on the right and a single-story glazed sun-room on the left. The building is deliberately not left-right symmetric, which is what tells a frame apart from its own mirror image. Every later frame is checked against this list.
All seven were driven from the identical generated still with the same prompt and negative prompt, then screened by a local proxy: feature continuity, camera motion, source retention, cut safety. Thirteen runs across two scenes, $8.44 actual against a $9.00 budget.
The proxy ranked Veo first at 90.9 and Kling fifth at 85.3. Kling is the one that actually reconstructed. That is why the screening file says in its own limitations that the ranking is triage only and must not be presented as a reconstruction result.
Screening formula: 35% feature continuity, 25% camera motion, 25% source retention, 15% cut safety, over SIFT correspondences with RANSAC fundamental geometry. It never ran a structure-from-motion solve, so it could not see whether the generated 3D was right. Persistent 2D texture matches features even when the geometry underneath is wrong, which is how Veo scored first and then failed the solve.
Each clip is fed to a structure-from-motion solve under pinned intrinsics. If the cameras cannot be recovered from the frames, the clip is rejected however good it looks. The score is how many of 210 image pairs come back geometrically calibrated.
The best score on the board belongs to a clip that was thrown away. The cheapest result belongs to a clip that was never generated at all.
The first guide that produced a solvable generation at all.
The middle arc of the r7 family, and the one that shipped.
Gated higher than the arc that shipped. Rejected anyway.
The best gate score in the project, guide and generated clip both.
Radial-only parallax the mapper cannot bootstrap. Caught before spending anything.
Six candidates were rendered and solved locally at no cost. Five passed at 21/21 registered, 0.73 to 0.82 px, zero watermark-configuration pairs. One failed outright. Only the survivors were allowed to become paid generations.
Looked fine. Solved as a single degenerate configuration on every pair.
The first generated clip that reconstructed at all.
Sharpest single arc. Step off the path it was solved on and it falls apart.
Strongest parallax in the project, and the first view past the corner.
There is no photograph of this house as built anywhere in the project. The baseline is the strongest prior generated clip, Kling 3.0 Pro image-to-video at $0.56, re-derived from scratch and reconstructing at 30.34 dB.
The donor repo had claimed 32.45 dB on the same data. That did not survive re-derivation and is recorded as unverified. The claim that survives is narrower: generated video reconstructs within 2.56 dB of the best prior generated clip, on a gate that rejected earlier attempts outright.
These are the trained Gaussian splats themselves, not renders of them. The camera stops where the solved frames stop, so you cannot get lost. On the four models that only ever saw the front, that edge is a wedge and the drag hits it quickly. On the four that went all the way round there is no edge, and the drag keeps going.
Eight models are here to compare directly. The front four are the re-derived donor clip that set the bar, the first clip generated to a camera guide of our own, four clips fused into one model, and the orbit hop that pushed coverage past the corner. The four with a back all come from section 09 and differ in where the camera stood and how it got there: solved from the frames, read straight out of the render manifest, flown close enough to the wall to pick up brick, and both laps trained together.
Nothing below is a still image standing in for a model. Every number attached to these was measured on views the training never saw.
Step 1: pick a model
Step 2: load it
Drag to orbit, scroll to zoom, shift-drag to pan, R to reset. The viewer has preset viewpoints and a reset button of its own. Picking a different model above reloads it in place.
This limit took four attempts. The first was a guess, and its −55° preset showed pure smear. The second came from a sharpness metric, which said three of the four models hold out to ±80°; at +80° there is no building in the frame at all. A collapsing splat does not go soft. The gaussians stretch into streaks and the foliage breaks into confetti, and both of those raise sharpness, so the metric scored the collapse as healthy. The third used solved camera positions, which are accurate but answer a different question: where a solved frame stands, not how far the solve carries. They are symmetric when this property is not.
The fourth was set by looking. 52 frames were rendered through this viewer with the panel hidden, four models across thirteen azimuths at 10° steps, and each was marked clear, degraded or gone. The marks are recorded per frame, so any one of them can be argued with by pointing at the frame.

Line the four models up by how much camera arc each solve was given and the return collapses monotonically: 5.06× readable degrees per camera degree, then 1.75×, then 1.01×, then 0.83×. A 10× spread in capture width bought 20° of extra readable view, and the two widest captures read no wider than the arcs their own cameras stand on. Past that knee, adding generated frames further round the house buys very little. The next gain has to come from the pixels already captured rather than from more camera positions.
Feeding generated frames into a reconstruction costs −1.43 dB against a baseline built from solved frames alone, measured on held-out views. The cause is known too. A warped frame is two things stitched together: pixels resampled from a solved frame, and pixels a model invented where nothing was visible. The warp records which is which, per pixel. So the obvious repair is to give the trainer the resampled half and mark the rest unknown.
That repair assumes unknown is something the trainer can be told. The only candidate is a transparent input pixel, and the manual is ambiguous about what one means. So it was tested rather than assumed: a rectangle punched to transparent in half the training views, the other half left intact, everything else identical.

This is strictly worse than feeding the invented pixel through. A wrong color is one bad observation competing with good ones; a wrong emptiness deletes geometry that other views got right. Masking a warped frame down to its resampled pixels would carve the disoccluded regions out of the house. Two short training runs and one rectangle cost about four minutes and stopped a larger experiment that would have returned a negative result for the wrong reason. Sealed as refuted in alpha-semantics-20260811-r1; the same idea still has a coarser form left, selecting whole frames by how much of them was invented.