The plan calls for front, sides and back before any camera moves. Only the front exists. Four video clips were generated from it and fused into one reconstruction, and all four begin at the same anchor frame, so the entire solved dataset, 104 frames of it, stands within about 60° of the front door. That is the ceiling section 07 measured four ways.
The reconstruction is not failing here. It is declining to invent three elevations nobody showed it. The rest of this section is the work of getting those elevations made.
The first plan for the other elevations was hierarchical: get the whole orbit roughly, then refine it. Priced at $5.29, authorised, and stopped after $0.44 on its own pre-registered kill gate. The opening probe asked an angle-controlled image endpoint for 0°, 10° and 20°, and got the same picture three times. At 30° it crossed some internal bucket boundary and returned a different house altogether: different light, different massing, different roof. A second model from a different family did the same thing.

The correction is to stop asking for angles in words. A camera moving continuously inside a single video clip never names an angle at all. So the next round runs one wide pass the whole way around the building, elevation drifting as it goes, held together by keeping the subject small enough in frame to stay consistent. Then successive passes zoom into segments of the pass before them, each conditioned on the one it came from, until every face is covered at full detail. Sections 05 to 08 are the constraints that plan has to survive.
A pass around the building needs references for the angles it travels through, and buying those from an image model is what 09b just failed at. The drawings already describe those angles. The reviewed footprints were extruded into a massing on 16 July (first-floor envelope, garages, the angled suite wing, the double-height great room and foyer, the gazebo, the covered entry), and the script that built it rendered one front-entry camera, wrote the PNG and threw the scene away. The geometry that answers “what is behind this house” was sitting there the whole time the measurements were calling the rear unrecoverable. Nobody had pointed a camera at it.
Below is that same massing from four azimuths. The geometry was not re-authored: the run reads the original script’s module source and executes it verbatim, minus its single fixed-camera render, and contributes the camera only. Three of these four views had never existed in any form. $0.00 billed, no network calls.

What stays hypothesis stays labeled, in the massing’s own words: wall heights, roof pitch and intersections, facade window rhythm, material placement, terrain form. What is plan-grounded: the footprint components, the gazebo outline, the covered entry, the south balcony relationship, the angled suite wing, the garage recession, and the double-height great room and foyer.
Everything was held even across the four: one prompt, one five-image reference set (the generated front elevation for appearance, the four plan cardinals for shape) and one clip from each vendor, kept whether it passed or failed. Every request body was checked against that endpoint’s own published schema before anything was billed, which caught one vendor being sent fields it does not have. Unknown fields get dropped without complaint, so that clip would have rendered, charged, and applied none of the constraints the test is about. $3.31 billed across four clips.

All four failed the same way, which turned out to be worth more than a pass. Every clip is recognizably the generated front elevation (cream siding, red brick corner tower, standing-seam hipped roof, wet driveway) and none of them is the 47.9 × 8.5 × 30.3 building in the drawings. The gray cardinals were read as hints about how the house should look, not as its shape. Reference images carry appearance. Handing a model four views of a shape does not make it build that shape.
That points at the endpoint left out of the test. Every clip this project has ever reconstructed came from a depth-controlled model, one that takes a video and follows its depth, rather than taking pictures and guessing. Earlier runs had to render that control from a splat that was itself solved from generated frames. The orbit in 09c is a better control than any of those, because it comes off the plans. So a fifth arm sends the same prompt with the reference set cut down to one image, and hands the shape over as 241 frames of control video instead.
The only change is how the shape got there. The four gray cardinals came out of the reference set, and the orbit went in as control video, frame for frame at the endpoint’s own 241-frame cap, so one output frame is generated per control frame and the 1.5° per frame holds end to end. $1.51, one request.


A measurement that always passes proves nothing, so this one was run first against three answers already known (the control against itself, against itself reversed, and against itself shuffled) before it was allowed to report anything, then against a clip that never received a control at all.

The other three gates come from the same audit that produced the strip above: no cuts, and the camera travels. Frame-to-frame difference peaks at 1.11× the clip’s own median, the steadiest of the five by a distance, against 8.59× and nine disturbances for the best-looking uncontrolled clip. The first and last frames, a full circumnavigation apart, are the same house down to the planting beds.
What it did not buy: the brick corner tower the prompt asks for is absent, because the massing has no tower and the control beats the reference on anything structural. The haze is heavy, and detail softens at the top of the elevation ramp, where the control is furthest from anything the reference image showed. Both channels did exactly what 09d said they would: the reference supplies material and light, the control supplies form.
$6.32 spent to this point in the section, of which $1.51 appeared to buy nothing: an earlier submission of this same request was accepted by the queue and then lost by the runner, which gave up polling at fifteen minutes and had not yet written the request id to disk, so the job could be neither collected nor canceled. Compute that ran is compute that bills, so it was counted as spent rather than quietly dropped. The runner now writes the id before the first poll and can be told to collect a job it already owns instead of ordering a second one. That clip has since been recovered and it turned out to be worth more than the one that arrived on time. 09g covers it.
The technique has a second half: go round again at higher detail, building on the pass before rather than starting over. The obvious way to do that is to shrink the orbit, and here that fails on arithmetic. The footprint is 47.9 × 30.3 m, so half its length is 23.9 m, and any circle small enough to raise detail meaningfully puts the camera inside the dining room.
That is an argument, so it was tested. Four-angle probes at 0.86, 0.78 and 0.70 of the fitted radius cut the building at 3 of 4 angles, then at all four, twice over. Even the mildest of them, a 14% reduction, is already losing the ends of the house.

A third attempt re-solved the radius per frame to hold a constant screen fill. That stopped the cutting, and it is the right way to frame a wide pass, but as a second rung it goes nowhere: the solved radius averages 44.4 m against the fitted 47.8, so the building comes out about 7% larger on average, and on the frames where the orbit swings out to 48.3 m, slightly smaller. It is kept under its own name rather than presented as a rung.
What works is to stop trying to fit the whole building in. Hold the vertical fill at 0.92 and let the frontage run past both edges on purpose. The house is 8.54 m tall against a 16:9 frame, so the full height still fits while the width overflows, and the standoff then follows the silhouette on its own: close at the corners, further back across the long face. Every frame gets the full height of the building and a sliding slice of its wall.

The other half of “using preceding segments” is the reference set. Rung 1 was handed exactly one image, the generated front elevation, which was the only view of this house that existed. That is why its sides and back are the model’s invention. Rung 2 is handed rung 1’s own delivered frames at the four cardinals, so it starts from a house that already exists at every angle. Which output frame sits at which angle is read out of the frame-by-frame match from 09e rather than assumed, landing within 0.75° at worst.

It was submitted, and it came back: $1.51, one request, 241 frames at 1280×720. Put through the same instrument as the wide pass (every output frame matched against every control frame on edge structure, self-tested against three known answers before it was allowed to report an unknown one), it follows the close control across 357° at slope 0.998, with 99.6% of frames on the diagonal. It also stalls less than rung 1 did: 210 of the control’s 241 positions are held by a distinct output frame, against 187 for the wide pass, at a mean offset of 0.41 frames rather than 0.49. No cuts: the loudest frame-to-frame difference is 1.58× the clip’s own median.

What did not carry over is the surface. Rung 1 delivered a smooth cream-rendered house; rung 2 came back in cream weatherboard with a brick entry. The palette held (cream walls, pale metal roof, the same overcast wet light) and the material did not. Four reference images tell the model roughly what the house is made of, and nothing in the request holds any one of them at any particular frame. The field that would do that is the one 09i found empty.

That takes the run to $7.82 of the $7.83 authorized: six requests that came back with a clip, plus the seventh charge 09e describes. Two laps of the same circumference, the second at 2.43× the wall detail of the first, both driven by cameras that came off the plans and both measured against them.
Measuring the rung turned up something about the one before it. The wide control in 09e, the one that was paid for, cuts the building on 16 of its 241 frames, worst case 1.04× the half-frame. Mild, and nobody knew, because the framing check was written after the money was spent. It was measured afterwards by replaying each frame’s own camera position out of that frame’s filename and re-rendering nothing, so the number describes the exact bytes that were submitted. The check now names that pass on every run instead of quietly passing it.
The lost job was never gone. fal keeps a request history, and 1.9 hours after that job ended the history still held the record: status 200, 833 s of compute, and a video URL that was still live. The clip was pulled down and put beside the one that had been collected. So the run ends up holding something it never meant to buy: two independent executions of a byte-identical request, submitted 17.7 minutes apart, run on whatever workers the queue happened to pick.
The two files are the same size to the byte and their hashes differ, which sounds like a difference and is not one. Hashing each part of the container separately shows where it is: 318 bytes differ out of 14,737,917, all of them inside the metadata box, and the 14,720,606-byte payload, the encoded video itself, is identical bit for bit. Decoded and compared as pixels, all 241 frames come back at zero mean squared error on every plane. It is the same clip twice.

This changes what the next dollar can buy. The authorization for this arm was written as one, and if it is bad, two, which assumes a second run gives a second sample the way re-rolling an image model does. On this endpoint it does not. An identical request returns the identical clip, so resubmitting cannot rescue a pass you dislike; only changing what gets sent buys anything. The close control in 09f is that changed request, and it is where the last $1.51 went instead of into a second lap of the wide ring.
The clip went through the same solve as every other reconstruction on this page. How far apart the sampled frames are turned out to be the whole variable. At 11.2° between frames, only 6 of 32 images survived the solve at all. At 5.6°, the frames broke into two models that share no coordinate frame, 44 images in one and 21 in the other. At the clip's own 1.5°, all 241 registered into a single model at 0.897 px mean reprojection error. Every earlier splat here came from a front-arc clip whose rear frames registered zero times (83 consecutive frames and 224° of orbit), measured twice.
Trained at the standard 30k steps with every eighth frame held out, it returns 31.35 dB across 31 held-out views, against the 30.34 dB baseline of section 06b. Split by where the camera was standing, the front arc is the worst of the four at 29.92 dB and the rear is 31.74 dB. SSIM does not follow: 0.860 against the baseline's 0.910, on a harder split (31 holdouts rather than 4), so the two numbers come from different tests, and only the PSNR comparison is like for like. Loadable as Circumnavigation in section 07. Local GPU, $0.00.
Every clip on this page fails the same way: the house is right where the camera starts, and has quietly become a different house by the time the camera comes back round. The standard fix is to pin a known image to a specific position in the output, so the appearance cannot drift away from it. Nothing above does that. So the schemas were read rather than recalled: 48 video endpoints, live OpenAPI documents, sorted on the two capabilities this technique needs.
The lit box is where the answer was. 11 endpoints take a pinned frame and a control video whose geometry the output must follow. Every one of them is a VACE variant, and one of them is wan-22-vace-fun-a14b/depth, the endpoint 09e and 09f bought their two clips from. It accepts first_frame_url and last_frame_url. Both requests sent to it left both unset. So did the other 4 of the 6 submissions this section logged, but those went to endpoints that have no such field to leave unset, which is the distinction the survey exists to draw. On the two that could have been pinned, this was a capability sitting there unused, and only the schema could tell that apart from a capability that was never on offer.
The 23 endpoints that pin frames and nothing else are the ones with the familiar names: Veo 3.1, Kling, Wan FLF2V, Flux 3 keyframes. They do not substitute here. They interpolate between the images they are given, inventing the path in between, and the path is the one thing this project can already compute exactly: the control video in 09e is a 241-frame orbit whose every camera position comes off the blueprint. Handing that to a model with no geometry channel throws away the reason the orbit was built. Each family gives you half the request.
Pinning at an arbitrary index (a fixed frame at each of the four cardinals, not only at the two ends) is rarer still. Of the 48 it exists only on the inpainting task of the two VACE families, carried by a mask video rather than a field. That is the request this section has been circling since 09: the four elevations already exist as stills, and the orbit already knows which frame each one belongs at.
None of this changed the rung that was in flight while the schemas were being read. Control standoff and reference set were the only two things it was meant to vary from 09e, and a third difference would have made the comparison unreadable. There was also nothing to pin it to: no styled still exists at the facade framing, and the only candidate is a crop of a wide frame, which would have to be enlarged by the same 2.43× the closer standoff buys, anchoring the clip to the softness the rung was bought to remove. That rung has since landed, and what it failed to inherit was its materials, which is the failure this field addresses. The pinned run is the next request, priced and authorized on its own. Reading the schemas cost $0.00.
Both rungs were checked against their control and against each other, and both passed. Neither check asks whether the house in the frame is still made of anything. This one does. Every frame of both clips was scored inside the building’s own outline for three things: color, how much of the wall is opening or shadow, and edge detail.

Color saturation on the building runs 0.62 across the front arc and 0.46 from 90° to 270°. The fraction of it in shadow or opening falls from 0.68 to 0.42, which is the windows and doors going away. Edge detail barely moves: the worst frame of 241 still holds 62% of the front-arc figure, and no run of eight frames anywhere on the orbit drops below half. Cladding lines and planting keep the edge count up while every opening in the wall closes over, so a sharpness measure on its own would have passed this clip.
Both rungs fail across the same arc, which follows from how they were ordered. Rung 1 was given one reference image showing the front, so its sides and back were invention. Rung 2 was given rung 1’s own delivered frames at the four cardinals as its reference set, so three of its four references were already showing the blank walls, and it inherited them.
This is what the holdout figures in 09h cannot see. PSNR compares a render against the frame it was trained on. If the frame stopped showing a house, a reconstruction that faithfully reproduces a blank cream box scores well for it. The rear arc coming back 1.8 dB above the front is what that looks like from inside the metric: less on the wall to get wrong.
So the reconstruction is doing its job and the ceiling is upstream of it. The earlier splats on this page cover a restricted arc of a house that has brick, glazing and interior light in every frame they were trained on. This one covers the whole circumference of a house that has those things on about a quarter of it. The circumference was paid for with the materials, and nothing in the run reported that trade until it was measured here.
Two things have to change before the next clip. The empty field in 09i pins a known image at a fixed position so the appearance cannot drift away from it. But there is nothing to pin at 180° yet, because the plan massing carries no material anywhere. The four elevations have to be styled first, at the framing the orbit will meet them at, and then pinned to the frames the orbit already knows they belong to. 09m writes that framing down, one camera per elevation. Measuring this cost $0.00.
Every splat here has been scored against its own training frames, and those scores cannot be read against each other. A front-arc model is graded on front-arc frames and does well for reproducing them. The one that covers all four sides scores 1.8 dB higher on its rear arc than on its front, which sounds like the rear came out better and means the rear has less on it to get wrong. So this plate carries no metric. Eight models, eight bearings 45° apart, camera limits released so each one is asked for views its data cannot support. Where the data runs out the renderer still draws whatever gaussians happen to face that way.

The front frame of the first four rows is the best picture of this house anything on this page has produced. Brick piers, glazed bays, lit rooms behind them, planting at the base, a roofline that reads as built. Turn 45° and it is already going. By 90° there is nothing left to read. The model with the highest holdout scores on this page is the one that falls apart hardest: row one averages 112 out of 255 in brightness at the front and 71 across the other seven bearings, which is the renderer running out of gaussians and showing the empty scene behind them.
Row five holds a house at every bearing. The proportions are right, the ridge line stays level all the way round, the porch reads from the side and the rear gable reads from behind. It has no windows, no doors and no brick, for the reason 09j measures: the clip it was trained on stopped showing those things once the camera left the front.
Four models have the house’s appearance across a quarter of a turn. One has its shape across the whole turn. The goal was one model with both, and every attempt so far has traded one for the other.
Rows six to eight change how the camera positions were obtained. Every model above them was reconstructed by COLMAP, which recovers where each camera stood by matching features between frames. None of these frames are photographs. Each was rendered by a camera this project placed, and the manifest that drove the render already lists that camera for all 241 frames. Rows six to eight skip the solve and read the manifest.
That puts the reconstruction in the plans’ own world, in meters, which can be checked. Rows six to eight come back 0.6°, 1.6° and 1.7° off vertical, with extents within 10% of the house the renderer measured off the walls. The COLMAP solve of the same 241 frames comes back 23.6° off, at about a tenth of the size. Because these three are in meters, their camera distance is computed: the measured house box against the viewer’s own focal length puts it at 66.8 m. It lands inside the apparent-size band the five fitted rows occupy.
Row seven is a second lap flown 6.3 m off the wall. Viewed from 32 m it shows brick coursing, window frames, veranda posts and lit rooms behind the glass, which no full-circumference model here has had. Above the veranda it has nothing. That camera never stood far enough back to see the roof, so there is haze where the gables belong, and at the 66.8 m the plate uses it is mostly haze.
Row eight trains on both laps at once, 482 images. It has the roof and the back, and at the front it is softer than row six, which used half the data. Averaging two standoffs into one reconstruction gives up the close lap’s detail and gains no reach. The two laps are better kept as separate models, each viewed at the range it was flown for.
Two things had to be fixed before the plate could be trusted. Each solve carries its own arbitrary world scale, the largest here 50× the smallest, so a shared camera distance had to be fitted: stand back 1.594× the radius holding 40% of the gaussians the renderer meaningfully draws. That reproduces the four distances already curated for the model picker within 6.5%. Dropping the near-transparent gaussians first is what makes it work, because 52–86% of every cloud here is floaters, so a percentile taken over all of them measures the halo.
The rule still does not reach the fifth model, because it was fitted on four clouds that all trail floaters the same way and that one does not. In the front-arc models the radius holding 90% of the gaussians is 2.1 to 2.9× the radius holding 40%; in the circumnavigation it is 1.5×, because a camera that went all the way round leaves nowhere for a halo to stream off to. The multiplier had absorbed the tail shape of the other four, so applied here it stands the camera at 2.20 in a cloud whose last 1% only reaches 2.68, and the house overfills the frame. That row is placed by eye at 3.4 against a recorded ladder, and the shot index says so.
The second was the vertical. Nothing in a splat file says which axis is up; the viewer takes it from the start view each model was given, and the circumnavigation model was added to the picker with another model’s start view pasted on. Nobody noticed while nobody orbited it. The scene is a house on a large flat lawn, so the cloud is overwhelmingly planar and its direction of least variance is the ground normal, which puts that solve’s real vertical 23.4° away from the one it was handed. Deriving it is what keeps row five upright through the whole orbit. Shooting the plate cost $0.00.
The target is a specific file: a splat of a different house, headed 106,331 splats and 64 of 64 poses, photographic up close, and restricted to one viewing angle. The job was to keep that quality and remove the restriction. Counting splats answers none of that, so the panels below are cropped at 1:1 out of rendered frames. A pixel in one panel is the same size as a pixel in the next.

The first panel is still the finest texture on this page and nothing here matches it. Glazing bars stay separate, the furniture holds its edges, the planting outside the window reads as leaves. That gap comes from what was fed in. The benchmark reconstructed photographic frames; the circumnavigation laps reconstructed frames generated from a massing render, which never had brick or leaves in them to recover.
The second panel is the same model looking outward, and it is the reason the restriction existed. Large smeared gaussians sit between the camera and the building, which is what a reconstruction does with the directions its 64 poses never covered. Panels three to five carry 241, 241 and 482 poses and do not have that halo, because the camera went all the way round and left the halo nowhere to stream off to.
The count is met and the coverage is beaten. The appearance still falls short, and the fix for it sits upstream of any reconstruction: the circumnavigation clips need the styling the front-arc clips already got, applied per elevation, before they are trained again. Panel three shows what that looks like when the source has it, at one story and one lap.
09j ends on a prerequisite and no numbers: the four elevations have to be styled first, at the framing the orbit will meet them at. There are already four cardinal stills in the run and they are the obvious thing to style. They were rendered to fit the whole 47.9 m building in frame, at radius 47.8 and a fixed 12° elevation. The facade orbit reaches those same four bearings at radius 25.1 to 32.8 and elevations from 2° to 10°, so a styled version of a cardinal still would have to be enlarged by between 2.16× and 2.70× to reach the orbit’s scale. 09i rejected a crop at 2.43× for exactly that, since it anchors the clip to the softness the closer standoff was bought to remove.
The still to style at each bearing is already rendered. The facade pass’s own control frame there is the framing the orbit meets that elevation at, by construction, so the four frames below are the ones the styling goes on and the delivered frames they pin to are the ones at the same index. The bottom row is what came back at those four frames: brick and lit windows at the front, where the one reference photo was taken, and blank walls at the other three, because there is nothing at those bearings to inherit material from.

The resolution figures come out of the renderer’s own per-frame schedule rather than off the pictures. Two relationships hold across all 241 rows of it to within 0.007 px/m: detail scales as (height / 2) / (nearFace × tan(fov / 2)), and the near-face distance is the camera’s horizontal standoff minus a support distance belonging to the massing. That second term is a function of bearing alone for a building with vertical walls, which is what lets the same decomposition be evaluated at the cardinal stills’ own radius and elevation and compared in the same units. Writing this down cost $0.00. The request it describes is priced and authorized on its own.
Before buying the styling, it is worth checking what is on the four frames 09m names. The sealed massing script carries 8 glass elements. 6 of them sit on the house and all 6 are on the front. The other two are not on the house: their coordinates are two edges of the detached gazebo, which has no footprint of its own. That leaves 0 openings across right, back and left together, which is why those three render as flat slabs.
None of the eight came from the plan. Every one is a literal pair of coordinates typed into the render script, under a comment there calling the opening rhythm a hypothesis. The verified schema holds 19 openings and all 19 are interior, doors between two named rooms, and the entry config’s own authority record says no exterior elevations were ever supplied. The front has an appearance because somebody drew eight rectangles for the one camera that script was written to render.
This section was first written to say the front is the only elevation with modeled openings on the strength of a field in my own schedule script whose value was the literal test side == "front". It was typed in, not measured. Replacing it with a parse of the massing source is where everything above came from, including the two windows on the gazebo, which nobody had noticed were there.
What the plan does carry is which rooms reach the outside and which way each one faces. A room’s exterior wall is the part of its own polygon lying on the building outline, so that is computable per room: of 51 rooms, 24 have one. The three blank elevations have 134 m of wall between them and 12 rooms behind it whose names imply glazing. The back carries 11 rooms and 75.4 m on its own.

Which way a wall faces is settled by stepping off it and asking whether that step leaves the footprint. The polygon’s winding will not do it: winding points into the house at a recess, and this building has a recessed entry and two covered porches, which put 282 plan units of entry wall on the back elevation until the test was changed. The lengths are in meters because the massing script declares its own scale constant and this reads it out of the file, so a change to the massing moves these numbers with it. The schema forbids inferring measurements from polygon scale by itself, and it is right to. Producing the schedule cost $0.00, and it narrows the request 09m priced: one elevation can be styled from the generated seed image in 04, and the other three need their openings placed from the plan first.