Reference video → structured observations
PySceneDetect, OpenCV, MediaPipe and librosa separate cuts, camera motion, pose availability, framing and timing before generation.
A purpose-built research system that decomposes reference structure and regenerates controlled short-form outputs using separate identity, look and scene packages.
Not a general AI video generator. This online version is an evidence viewer; generation and state-changing actions are disabled.
The project integrates video analysis, visual generation, image-to-video synthesis, diagnostics and repair logging into one traceable dissertation prototype.
This is not a single prompt-to-video demo. Deterministic analysis, separated conditioning, per-shot keyframes, prompt-level I2V, post-hoc metrics and decision records form one inspectable workflow.
PySceneDetect, OpenCV, MediaPipe and librosa separate cuts, camera motion, pose availability, framing and timing before generation.
gpt-image-1 combines inspectable generation units with separate identity, look and scene packages.
Kling I2V clips are shown as existing evidence. Paid generation remains disabled in this public demo.
ArcFace / InsightFace, optical flow and the pose-mask audit expose identity, motion and control-signal limitations. Executed acceptance remained human-gated.
DWPose, SAM3 and Wan2GP / Wan2.2 support the documented Mode C feasibility branch; full production was not validated.
A readable map of the four operating modes, the shared evidence core, separated conditioning inputs and the dissertation claim boundary.
Four routes define different sources of structure and different evidence maturity.
Reference-derived shot logic enters the keyframe-first evidence core.
Completed creative extension on a separate ComfyUI path.
Person replacement follows a control video through a separate Wan backend.
Post-submission: 7 s segments completed locally (4 Aug 2026); 27 s still pending.
Planned extension · hosted-service demos only.
The Mode A/B evidence path is inspectable from structured input to retained output.
Evidence or authored intent becomes an inspectable shot-level structure.
Identity, look, scene and reference-structure inputs are combined into reviewed visual anchors.
Approved keyframes become I2V clips; human-gated review records decisions before final assembly.
Executed acceptance remains human-gated and supported by post-hoc diagnostics. Automatic metric-driven Stage-4 scoring was designed and partially instrumented, but not fully executed.
Control sources stay separable so provenance and failure diagnosis remain clear.
Shot order · framing · action intent
Subject identity · view references
Outfit · hair · styling
Setting · viewpoint · lighting
Three existing artefacts show how source structure becomes a reviewed keyframe and an assembled evidence video.

Beach setting, shoulder framing and action intent.

Selected identity and look inside the retained beach scene.
Open the 20-second retained video from the corresponding opening-shot evidence.
Claim boundary. This is a qualitative evidence sequence. It demonstrates reference-scene preservation and reviewed assembly, not exact frame-level motion transfer.
This project investigates whether reference-driven video generation becomes more controllable and auditable when structural, identity, look, scene and repair decisions are externalised as inspectable records, rather than embedded in a single opaque prompt-to-video process.
The study evaluates the extent to which decomposed conditioning packages, shot-level generation units and explicitly logged repair decisions improve workflow traceability, while explicitly not claiming exact frame-level motion transfer or fully automatic metric-driven acceptance.
Reference analysis is decomposed into per-shot generation units, which are then combined with separately sourced identity, look and scene packages before keyframe-first I2V generation. Generation is therefore planned at shot level rather than treated as one opaque end-to-end process.
Completed runs, human-gated repairs, post-hoc diagnostics, feasibility extensions and archived negative results are reported under distinct evidence labels, so that the strength of each claim matches the maturity of its evidence.
Timing, framing, camera motion, character motion, scene intent and prompt blocks are exposed as inspectable per-shot records rather than hidden prompt state.
Evidence anchorReference analysis page + Mode A shot records.Reference structure, identity, look and scene conditioning retain distinct provenance, enabling failure diagnosis and separation of claims.
Evidence anchorMode A settings + looks/scenes + street scene-package run.Attempts, diagnoses, interventions and acceptance decisions are recorded in the decision log; human review remains the acceptance authority, and fully automatic Stage-4 execution is not claimed.
Evidence anchorDecision log + repair evidence page.The public site exposes inspectable implementation evidence while keeping generation and project mutation unavailable.
GET /Portfolio interfaceCurated academic explanation, modes, methods, evidence and limitations.
GET /media/…Evidence routesExisting videos, images, reports and audit packages remain directly inspectable.
READ-ONLY RECORDSDecision and evaluation recordsAttempts, diagnoses, repair actions, metrics and acceptance context remain browseable.
POST /api/… → 403Server-side safety boundaryGeneration, upload, reset and state-mutation routes are disabled in the online demo.
DISABLED IN ONLINE DEMODirect links to the current dissertation PDFs, software repositories and cited papers.
Mode A is the dissertation’s primary reference-video pathway. It analyses a reference clip into shot structure, framing logic and action intent, then combines those records with selected identity, look and scene packages.
The implementation separates two settings: A-1 preserves the reference scene; A-2 tests cross-scene recomposition.
Mode A separates reference-video structure from identity, look and scene conditioning. A-1 is completed as reference-scene preservation; A-2 is reported conservatively as cross-scene recomposition / scene-package evidence.
The beach reference scene is preserved while the selected character/look is replaced. This is the completed main Mode A evidence.
4 analysed shots · 4 approved keyframes · 4 Kling I2V clips · final assemblyThe intended setting combines reference-derived shot/action logic with a new scene package. Current evidence is reported conservatively through the completed street scene-package run and the archived pose-guidance branch.
Street scene-package final · 4 approved keyframes · 4 clips · motion auditTerminology note. Mode A-1 is the dissertation’s “reference-scene preservation” setting (§4.3–4.7); the dissertation itself does not use the A-1 label.
Mode A-2 asks whether reference-derived structure can be disentangled from reference-scene appearance and recombined with a different scene package. This is not simply a background swap: it tests whether shot logic, identity, styling and environment remain separable enough to be recomposed and diagnosed independently.
Can shot order, framing and action intent remain useful when the original environment is replaced?
Reference analysis supplies structural records; the scene package supplies environment, viewpoint and lighting; identity and look remain separately traceable.
The completed street output reaches approved keyframes, Kling I2V clips and final assembly as a scene-package replacement / identity-and-scene stress test.
The separate start/end pose-guidance branch tested direct reference-structure transfer and was archived because skeleton-image guidance was too weak for reliable hard control.
Research value. The combined positive and negative evidence identifies the boundary between prompt-level cross-scene recomposition and reliable reference-structure transfer, motivating stronger future pose-conditioning mechanisms.
Reporting boundary. Mode A-2 is supported as an intended design setting and partially evidenced. The completed street final is reported conservatively as scene-package evidence, not as a completed reference-video transfer result; exact reference-structure transfer remains future work.
The detailed record below follows the completed A-1 beach setting from four analysed shots through approved keyframes, I2V clips, human review and final assembly.
The main Mode A reference-video pipeline and primary dissertation case study.
It tests whether reference structure can be decomposed and recomposed with separated identity and look controls.
Four analysed shots, four approved keyframes, four I2V clips, final assembly and decision records.
Shot structure and action intent are preserved; exact frame-level motion transfer is not claimed.
Mode A starts from an analysed reference clip rather than a free-form prompt. The reference video supplies shot order, framing, timing cues, camera and character motion observations, and scene continuity. These observations are converted into per-shot generation units, then combined with the selected identity and look package to produce approved keyframes and I2V clips.
In the Mode A-1 setting, the reference scene is preserved while the character/look is replaced. The beach case therefore tests whether reference structure, scene continuity and character styling can remain separable instead of being collapsed into one opaque prompt.
This case demonstrates the implemented Mode A workflow: local reference analysis, shot-level structure records, identity/look conditioning, keyframe-first generation, prompt-level I2V, human-gated repair and final assembly. Each retained clip remains traceable to a shot, approved keyframe, prompt decision and review record.
Contribution boundary. The result demonstrates controlled recomposition: reference-scene structure and action intent are retained while the selected character/look is replaced. It does not claim exact frame-level motion transfer or fully automatic metric-driven acceptance; identity and pose metrics remain diagnostic.
The prompts are assembled from structured shot observations and selected control packages rather than written from scratch.
Detect four edited shots.
Timing, framing, camera motion, character motion and scene.
Convert each shot record into a generation unit.
Add the selected identity and look package per shot.
Structure, scene, identity, look, camera and action blocks.
Select one reviewed keyframe for each shot.
Convert approved keyframes into Kling I2V clips.
Record human-gated diagnosis, repair and acceptance.
Four approved keyframes became four prompt-level Kling I2V clips and were assembled into the retained Beach Mode A-1 output.
Acceptance boundary. Human review remained the acceptance authority. Exact frame-level motion transfer and fully automatic metric-driven repair are not claimed.
The selected keyframe for each shot became the visual anchor for its corresponding I2V clip.
The record below follows the completed street scene-package run — the current evidence for the Mode A-2 setting — from four static scene references through approved keyframes, I2V clips and final assembly.
Scene-package replacement completed using static street references, not a reference video.
Tests whether identity/look conditioning remains coherent when a scene package, rather than a reference video, supplies the environment.
Four approved keyframes, four I2V clips, final assembly, motion audit and decision records.
reference_video = None; this is not a completed reference-video-driven Mode A-2 result.
Four approved street keyframes became four prompt-level Kling I2V clips and were assembled into the retained street scene-package output.
Acceptance boundary. The completed street final is not claimed as a fully completed reference-video-driven Mode A-2 output; it used reference_video = None.
The selected keyframe for each shot became the visual anchor for its corresponding I2V clip.
Why there is no reference video here. Mode A-1's completed evidence preserves an analysed reference scene; Mode A-2's completed evidence instead used static scene references with reference_video = None — that is the setting boundary itself, not a missing asset. Direct reference-structure transfer was attempted separately in an archived pose-guidance branch.
A completed scene-package replacement and identity/scene stress test using static street references and reference_video = None.
This case tests whether the selected Look 3 character can remain visually coherent when regenerated and assembled in a new environment, without using a reference video as the structural input.
The street run was built to test scene-package conditioning separately from reference-video structure. Unlike Mode A-1, which preserves the beach reference scene, this case replaces the environment with static street scene references. It therefore acts as an identity-and-scene stress test: the system must keep the selected character/look coherent while changing location, viewpoint and scene texture.
This distinction is important for the dissertation claim boundary. The completed street final demonstrates scene-package replacement and final assembly, but it is not reported as a fully completed reference-video-driven Mode A-2 result.
Completed using static street references.
Tests identity/look coherence in a changed environment.
4 keyframes, 4 clips, final, motion audit and prompt repair.
reference_video = None; exact transfer is not claimed.
The page follows the street run as an evidence chain rather than presenting a single output video.
Each approved keyframe acts as the visual anchor for one generated I2V clip. The set tests whether Look 3 remains coherent across close-up, back-view, feet-detail and profile-walk shots.
Motion audit asks a different question from visual quality. It measures visible optical-flow energy relative to the reference-motion target. The result is diagnostic, not proof of semantic correctness or exact structure transfer.
Headline finding (dissertation §4.13). Across all four shots the generated clips reproduce only 14–64% of the reference's per-frame motion energy (shot_001 14%, shot_003 19%, shot_004 64%); both beach and street generations under-animate relative to the reference. The prompt-only repair below acts on the two weakest shots.
Targeted action. Prompt-only repair addressed the two weakest shots rather than regenerating the full sequence. Values are retained from the run decision log.
Archived-log terminology. “Real footage” in the archived audit log is shorthand for the motion-target role. The reference clip is the author-created AI-generated beach performance clip (dissertation §5.3); the logs are preserved unedited.
These are development findings within the engineering dissertation. They explain why pose and mask signals remained diagnostic/candidate controls rather than being promoted as the final Mode A generation path.
Documented pose/mask candidates plus an archived start/end skeleton-guidance branch.
Negative and partial findings explain why these signals were not selected as final hard controls.
Twelve samples across four shots; nine pose/mask candidates; shot 003 unavailable in all samples.
Diagnostic/candidate evidence only; these controls did not generate the final beach or street outputs.
Attempted beach-to-street reference-structure transfer. Skeleton-image guidance was too weak, so the branch was archived as a negative result and did not produce the final street video.
Four shots × start/mid/end produced 12 inspectable samples. Pose and pose-seeded mask candidates were available for 9 samples; the three feet-detail samples remained unavailable.
Displaying the individual controls avoids the unreadable black margins in the original archival contact sheet.




Finding. Start/end images did not provide sufficiently strong temporal conditioning for reliable structure transfer. This is a backend/control limitation, not the final street result.
Each card shows reference frame, pose diagnostic and mask candidate. Shot 003 is intentionally retained as unavailable evidence.
Partial-body coverage and weaker boundaries make this unsuitable as production-ready hard control.
The strongest case: full-body visibility supported stable pose and a plausible foreground candidate.
All three samples lacked reliable pose. Mask output was left unavailable rather than invented.
Available for analysis, but not evidence that pose/mask control drove final Mode A outputs.
Mode C tests a separate control paradigm from the Mode A/B keyframe-first pipeline. Instead of converting a reference video into prompt-level generation units, it uses a driving/control video, person mask, reference images and pose/body conditioning to perform replacement inside the person region.
The newest evidence is a 7.0-second Wan2.2 Animate replacement window. It extends the earlier short Phase-0 evidence retained below, but it does not replace the main Mode A evidence and does not validate full 27-second production generation.
Research boundary. This branch distinguishes driving-video replacement from prompt-level reference-guided regeneration. It is not part of the Mode A keyframe/Kling path and did not produce the final beach or street outputs.
A separate driving-video person-replacement backend outside the Mode A/B keyframe-first core.
It tests whether a masked driving-performance video can provide stronger body-motion control than prompt-level I2V, while exposing the compute and identity-consistency limits of this backend.
Privacy-redacted control display, person mask, generated 7-second replacement, complete run metadata and earlier Phase-0 evidence.
Quality remains WIP. This is a 7-second window, not a full 27-second production run; identity, body proportions, temporal consistency and compute cost remain unresolved.
This is a longer 7.0-second Mode C replacement window generated with Wan2.2 Animate 14B. The retained local run used a consented driving/control recording that is not served publicly, together with a person mask video, reference images and OpenPose/BBox/face-movement conditioning inside the mask. The retained run produced a 544×960 portrait output of 209 frames at 30 fps using seed 630814980.
The result extends the earlier short Phase-0 evidence, but it is not presented as a full production pipeline. Generation took 2h 55m 14s, showing that this backend is computationally heavy compared with the Mode A/B keyframe-first workflow. It is therefore reported as extended feasibility evidence, not as a completed or interactive production mode.
Privacy-safe display copy: the participant's full head is covered by an expanded coarse mosaic. The original driving performance and audio evidence remains preserved locally.
Person-region mask used with face movement, BBox and OpenPose processing.
Retained 544×960 portrait replacement output.
Three concrete failure/limitation cases are part of the Mode C record and are surfaced here rather than left only in the raw run folders. (1) Extreme-pose identity and body-proportion distortion, observed in both the Phase-0 test and the newer 7 s window — the character's proportions and facial identity become unreliable on fast or extreme poses. (2) A session-ending GPU driver fault on the remote RTX 4080 host (nvidia-smi: No devices were found) occurred immediately after the 7 s run completed, halting further Mode C generation; this is recorded as an operational risk of running a 14B model on a 16GB GPU. (3) The 7 s run's three sliding-window stages did not each independently reach full length — only the final 209-frame stage matches the full driving clip; the two earlier-timestamped files are partial intermediate outputs of the same run, not three separately completed replacements.
Across both runs, the backend executed end to end without crashing mid-generation, driving motion was followed, and background outside the person mask was preserved. These failure cases bound the branch's production readiness; they do not indicate that the pipeline failed to run.
The earlier short Phase-0 segment established that the backend could run end to end on the documented remote RTX 4080 setup. The new 7-second window extends this evidence, but both remain feasibility tests rather than full production validation.
Privacy-safe display copy derived from the documented Phase-0 control; the collaborator's full head is covered by an expanded coarse mosaic and audio is removed. The original evidence remains unchanged and local.
The earlier short Phase-0 segment ran end to end on the documented remote RTX 4080 setup, followed the driving performance and established backend feasibility. The newer 7-second window extends duration evidence without turning either test into full production validation.
Identity and proportions distort on extreme poses; full 27-second generation and production quality were not validated.
The wan.video and Viggle material is exploratory hosted-service demo content, not Mode C evidence. Mode C retains the documented local Wan2GP / Wan2.2 Phase-0 and 7-second feasibility runs.
Mode D is the planned fourth pathway: delegating driving-video replacement to hosted online backends instead of the local Wan2GP setup used in Mode C. It is designed but not implemented. The material on this page consists of exploratory online-service demos only; none of it is dissertation evidence.
Reference input · 4 Aug 2026
AI-generated reference image created 4 Aug 2026. It defines the intended identity, black velvet gown look and vintage dressing-room scene for the planned hosted-service replacement tests.
Mode D asks whether the Mode C control paradigm — driving video plus reference identity — can be delegated to hosted services when local GPU capacity is the bottleneck. It is therefore a capacity and deployment question, not a new control paradigm.
Capacity context. A completed 7-second local Wan2.2-Animate window required approximately 2 h 55 m on the 16 GB RTX 4080 host, while GPU driver faults interrupted further local runs. Local execution was possible, but marginal and unreliable.
The four demos form a two-service × two-condition grid mirroring the Mode A-1 / A-2 distinction: same driving motion, same AI-generated reference image, with the scene either preserved or replaced.
Hosted character replacement on viggle.ai; the original room background from the driving clip is retained. 7.0 s, 718×1280, 25 fps. Provenance: viggle.ai watermark embedded in frame.
Service parameters not fully recorded; treated as demo per archiving rules.Hosted character replacement on viggle.ai; the environment is replaced with the dressing-room scene from the reference image. 7.0 s, 720×1280, 25 fps. Provenance: viggle.ai watermark embedded in frame.
Service parameters not fully recorded; treated as demo per archiving rules.Hosted video-to-video transfer (driving video + reference image); the character is replaced while the original room background is retained. 4.0 s, 720×1280, 30 fps.
Service parameters not fully recorded; treated as demo per archiving rules.Hosted edit transferring only pose and body motion from the driving video into the reference image's dressing-room scene. 4.8 s, 718×1280, 30 fps. Provenance: Wan watermark embedded in frame.
Service parameters not fully recorded; treated as demo per archiving rules.Comparison boundary. These demos informally replicate the dissertation's scene-preservation versus scene-replacement distinction (Mode A-1 vs A-2) on hosted backends. They are exploratory demos only; no parameter records were retained, so they carry no evidential weight.
The hosted replacement demos were reviewed qualitatively against the completed Mode A evidence. They are useful for fast exploration, but they expose control limitations that the Mode A pipeline is designed to make inspectable and repairable.
The hosted services accept a single reference image and offer no identity-anchor control. Across the four demos the generated character drifts from the reference identity, and the two services produce visibly different interpretations of the same reference.
Mode A uses a locked multi-view anchor set and per-shot human-gated repair, making identity drift inspectable and allowing it to be addressed at shot level. This does not imply that Mode A solves every identity issue.
In the scene-preserved demos the replaced character is not consistently relit to match the driving environment: the subject can read as composited, with lighting direction and colour temperature inconsistent with the room.
Mode A generates each shot as a whole frame from an approved keyframe, so subject and environment are reviewed within one lighting solution. The comparison remains qualitative.
Hosted services provide fast visual feedback, but the internal model version, preprocessing, identity handling and repair logic are usually not fully inspectable. This makes the output harder to reproduce or diagnose compared with the Mode A pipeline, where prompts, keyframes, decision logs and retained media are preserved as evidence.
When a hosted result fails, the main available action is usually to rerun or change the prompt. In Mode A, failure can be localised to a shot, keyframe, identity anchor, scene package or motion prompt. This makes repair more targeted and the reason for each revision more inspectable.
The comparison supports the dissertation’s core argument: a slower, inspectable Mode A pipeline provides stronger provenance, identity control and repair visibility than faster one-shot hosted replacement demos.
Mode D remains a planned extension. Online services were only exercised as undocumented demos; prompts, seeds and parameters were not fully recorded.
Four short online-service demo clips and one reference image, retained without full parameter records. They illustrate feasibility of hosted replacement but carry no evidential weight.
Re-run selected demos with recorded settings (screenshot of settings page, prompt, visible parameters, site/date/mode) so future runs can be archived under the standard outputs/runs/ rules and reported honestly.
Public-display rules for synthetic references, consented driving material and read-only evidence access.
The beach reference clip and identity/look/scene images used in the public evidence are author-created AI-generated materials. They are used as structural or visual research assets and are not presented as real people.
The real-person driving/control recordings used for Mode C and related planned extension tests were recorded with informed consent for research use. Public display uses identity-replaced, redacted or processed views rather than exposing the raw private recording.
The online dashboard is a read-only evidence viewer. It serves curated outputs, redacted display copies, generated replacements, masks, reports and audit packages. Raw control recordings and private working files are not publicly routed.
No real public figure likeness is intentionally used in the system inputs or outputs. Synthetic identity references are treated as generated research assets.
Only the minimum material needed to explain the method and evidence is shown publicly. Raw control recordings and private working files remain local and are not exposed as downloadable public evidence.
This page separates executed local diagnostics, post-hoc metrics and designed-but-not-fully-executed automatic scoring. Evaluation evidence supports review and repair decisions, but the final dissertation evidence remains human-gated rather than fully automatic.
Shot detection, camera motion, timing extraction and reference-side pose availability were produced locally from existing inputs.
ArcFace, optical flow, the pose/mask audit and visual framing review inspect retained outputs after generation.
Generated-side pose similarity, formal beat alignment and fully automatic Stage-4 acceptance were not fully executed.
The values below are retained diagnostics. A low or unavailable metric does not silently become a failed or passed shot; human review remains explicit.
Beach approved keyframes compared with the Look 3 front anchor. The metric is diagnostic only and was not used as a hard automatic acceptance threshold.
Visible motion energy relative to the reference. It does not prove semantic correctness, pose correspondence or exact frame-level transfer.
12 samples; 9 pose detections and mask candidates. Shot 003 had no reliable pose seed and is marked unavailable_no_pose_seed. Useful as diagnostic evidence and candidate control signals, not production-ready hard control and not used to generate final beach/street outputs. Manifest ↗
Framing and shot composition were reviewed visually against reference shot intent. The retained records contain framing labels, but formal fully automatic framing accuracy remains partial/pending.
Reference durations compared with the nominal fixed 5.00-second I2V target. Shot order and action intent are retained; exact beat timing and shot duration are not.
Formal beat-boundary alignment and generated-side pose similarity are not available as complete metrics. Reference pose availability is not equivalent to a generated-versus-reference pose score.
The metric-driven Stage-4 loop was designed and partially instrumented. A safe local dry-run harness reads existing outputs, assigns pending/not-applicable states and proposes repairs without executing regeneration. Completed dissertation evidence instead uses human-gated review and decision logs.
| Metric | Status | Evidence | Limitation |
|---|---|---|---|
| Reference / shot analysis | Executed local analysis | 3 retained test clips; 1/1/4 detected shots | Structured observations, not motion copying |
| ArcFace identity | Post-hoc diagnostic | 0.374 / 0.227; mean 0.301 | Not an in-loop threshold; no-face shots n/a |
| Optical-flow motion | Post-hoc diagnostic | Street v1 and promoted v2 values | Energy only; not semantics or alignment |
| Pose / mask | Documented control trial | 9 of 12 candidates | Shot-type dependent; did not drive finals |
| Framing | Human-coded / partial | Visual review and framing labels | No complete automatic accuracy score |
| Timing | Local duration audit | Reference durations vs nominal 5s clips | Not a formal beat-lock scorer |
| Automatic Stage-4 | Designed · not fully executed | Dry-run harness and repair plans | No automatic paid retry or acceptance loop |
Status is separated from interpretation so feasibility, completed evidence and future work are not conflated.
A status map separating completed evidence, human review, diagnostics, pending metrics and future work.
It prevents feasibility tests and partial instrumentation from being read as fully executed evaluation.
Beach, street, repair, motion, archived control, Mode C and Stage-4 status records.
Missing metrics remain pending or not applicable; no value is invented to complete the table.
Reference-scene preservation with four reviewed shots and an assembled final.
Exact frame-level motion transfer is not claimed.Identity-and-scene stress test with static street references.
reference_video = None; not structure transfer.Failures, diagnoses, retries and promoted results are logged per shot.
Human review remained the acceptance authority.Prompt repair raised visible motion energy for selected street shots.
Optical flow does not prove motion transfer.Start/end skeleton guidance was tested and retained as a documented failure.
Guidance was too weak for reliable transfer.A short driving-performance replacement ran via Wan2GP.
Full 27-second production was not validated.A dry-run scoring/repair structure was designed and partially instrumented.
No claim of fully automatic production acceptance.Identity diagnostic — executed post-hoc (dissertation §4.8.1). ArcFace (buffalo_l) cosine similarity of each approved beach keyframe to the Look 3 front anchor: shot_001 = 0.374, shot_004 = 0.227 (mean 0.301); shot_002 (back view) and shot_003 (feet detail) = not applicable, no face. All four keyframes were accepted as the same character at human review. The low similarities are evidence that a single face-embedding metric is mis-calibrated for a synthetic, stylised character — this is why human review, not ArcFace, is the acceptance authority.
This matrix separates completed system evidence from post-hoc diagnostics, feasibility extensions and planned architecture.
| Component | Status | Evidence | Limitation |
|---|---|---|---|
| Reference analysis | EXECUTED LOCAL | test_01, test_02 and test_03 | Produces structured observations; no exact motion-copy claim |
| Keyframe generation | EXECUTED | Four approved beach and four approved street keyframes | Human-reviewed visual anchors |
| Kling I2V | EXECUTED | Four beach clips and four street clips | Prompt-level fixed-duration I2V |
| Human-gated repair | EXECUTED | Decision log, retained attempts and promoted results | Human review is the acceptance authority |
| ArcFace identity | POST-HOC | shot_001 = 0.374; shot_004 = 0.227 | Cosine similarity diagnostic, not an in-loop threshold |
| Optical-flow motion | POST-HOC | Street audit; repair values 14→51% and 19→47% | Visible motion energy, not semantic or frame alignment |
| Pose / mask audit | DOCUMENTED TRIAL | 12 samples; 9 pose/mask candidates | Did not drive final beach or street outputs |
| Automatic Stage-4 | DESIGNED · NOT FULLY EXECUTED | Safe dry-run harness and repair plans | No fully automatic acceptance or paid retry loop |
| Mode C | PHASE 0 ONLY | Short driving replacement, mask and output | Full 27-second production not validated |
| Mode D | PLANNED | Architecture definition only | Not implemented |
This page shows how the system turns reference videos into structured generation evidence before any generation step. The analysis stage separates shot structure, camera motion, character motion, framing and timing so downstream prompts and keyframes are grounded in inspectable observations rather than free-form guessing.
Purpose: validate multi-shot cut detection and generation-unit creation. The edited source produced four detected shots, and those four records became the structural basis of the completed Mode A-1 beach evidence.
| Shot | Timing | Camera motion | Framing |
|---|---|---|---|
| shot_001 | 0.00–2.40s | push_in | medium |
| shot_002 | 2.40–4.67s | static | wide |
| shot_003 | 4.67–6.80s | handheld | unknown |
| shot_004 | 6.80–7.90s | handheld | wide |
Validates single-take fallback and subject-motion analysis while the camera remains static.
Validates camera-motion analysis and the separation between camera motion and character motion.
The output of analysis is not a copied video. It is an inspectable set of observations passed into downstream generation-unit and prompt construction.
Authorised single-take or edited structural input.
Cut boundaries create one record per detected shot.
Static, push, pan and other camera cues are recorded separately.
Subject action and pose availability are recorded where locally available.
Shot size, duration and beat-related timing remain explicit.
Each shot becomes a structured record rather than a free-form guess.
Analysis fields combine with selected identity, look and scene packages.
A reviewed visual anchor is produced per shot before prompt-level I2V.
Each shot record stores timing, framing, camera motion, character motion and scene description. The generation unit then assembles these observations with selected control packages.
Selected subject identity and view-specific reference material.
Outfit, hair and styling constraints kept separate from identity.
Preserved reference scene or selected static scene-package evidence.
Shot size, viewpoint and camera-motion observations from analysis.
Character action intent and prompt-level motion direction.
Shot-specific exclusions and consistency constraints.
A generation unit is the hand-off between analysis and generation. It keeps the source observation, selected packages, prompt blocks and review status attached to one shot. This simplified example shows the schema shape; the run record remains the evidence authority.
{
"shot_id": "shot_001",
"timing": "0.00–2.40s",
"framing": "over-shoulder close-up",
"camera_motion": "static / slight motion",
"character_motion": "turn / look",
"scene": "beach sunset",
"look_package": "Look 3",
"prompt_blocks": {
"identity": "…",
"look": "…",
"scene": "…",
"camera_framing": "…",
"action_motion": "…",
"negative": "…"
},
"evaluation_status":
"accepted_by_human_review"
}| Clip | Purpose | Expected structure | Detected result | Downstream use |
|---|---|---|---|---|
| test_01 | Subject-motion analysis | Single shot | 1 shot | Analysis smoke test |
| test_02 | Camera-motion analysis | Single shot | 1 shot; push/pan profile | Camera-motion smoke test |
| test_03 | Edited reference structure | 4 shots | 4 shots | Mode A-1 beach run |
The executed pipeline is keyframe-first and human-gated. It preserves shot structure and action intent without claiming exact frame-level motion transfer.
The nine-stage method from structural input to final assembly.
It shows where control, human approval, evaluation and provenance enter the implementation.
Analysis fields, generation units, packages, keyframes, clips, review, repair log and assembly.
The executed path is keyframe-first and prompt-level, not exact frame-by-frame motion copying.
Architecture mapping. These nine inspectable stages are a finer-grained view of the dissertation's four-stage architecture (§3.2): stages 01–03 = Reference Analysis + Template Builder, 04–06 = IP-Conditioned Generation, 07–09 = Evaluation & Repair + assembly.
An authorised clip or scene package defines the structural starting point.
Local analysis records cuts, timing, framing and pose evidence.
Each shot becomes a structured unit with explicit fields.
Reference packages remain separated by function.
One human-reviewed still anchors each generated clip.
Existing approved keyframes were animated per shot.
Human review is primary; metrics are diagnostic or post-hoc.
Observed issues lead to a targeted retry and recorded decision.
Accepted clips are ordered into the final evidence video.
These comparisons help a reviewer inspect scene, framing and identity changes. They are not measurements of exact motion transfer.
A qualitative comparison of source/reference material, approved anchors and completed outputs.
It lets a supervisor inspect visual continuity and deliberate scene replacement directly.
Beach source/keyframe/final and street scene package/keyframe/final using existing media.
This is visual evidence, not a quantitative proof of exact motion transfer.




reference_video = None.This page shows the agentic part of the workflow. The system does not treat generation as a one-shot output. Instead, each shot is reviewed, diagnosed and revised through a recorded decision process. The evidence is therefore not only the final video, but the trace from observed issue → diagnosis → repair action → retained result.
A linear workflow would generate all clips once and accept or reject the final assembly as a whole. This system records per-shot failures and allows the failed unit to be revised without rebuilding the entire sequence. That makes the workflow traceable, modular and auditable.
Human review identifies whether a shot is acceptable, while the decision log records the observed failure, repair action and retained output. The automatic Stage-4 loop is not presented as fully executed metric-driven acceptance.
Identity-anchor repair (dissertation §4.12, Figure 4.3). Left: front and profile identity anchors. Middle: earlier street keyframes that kept outfit and scene but drifted to a generic face. Right: after the reference-ordering policy promoted the identity anchor for face-visible shots. Qualitative human-review evidence; identity is improved, not solved.
Reconstructed from eight logged attempts across four shots (1/2/2/3 attempts per shot).
Reconstructed from the decision log: first-attempt acceptance 25% (1 of 4 shots) → 100% after human-gated repair, across 8 logged attempts (1/2/2/3 per shot). This is a log reconstruction, not a controlled ablation (dissertation §4.9).
Each repair attempt is recorded as a structured decision event. The log stores the shot identifier, attempt number, observed issue, diagnosis, repair action, prompt or keyframe change, review result and whether the output was promoted or archived. This makes the repair process explicit rather than implicit.
shot_idattempt_idobserved_issuediagnosisrepair_actionreview_statusretained_output
Four bounded intervention types organise the completed and archived evidence.
When the generated character drifted, the repair prioritised identity anchors and reduced competing references.
When visible movement was too weak, the repair changed motion wording while retaining the approved keyframe.
When skeleton/start-end guidance was unreliable, the branch was archived rather than forced into the final output.
When automatic scoring was not fully executed, it was labelled as designed future work rather than completed automation.
Reference priorities conflicted.
Identity anchor priority revised.
Keyframe consistency improvedMotion wording damped movement.
Prompt-only motion repair.
Shots 001 and 003 improvedStart/end control was unreliable.
Branch archived.
Negative result retainedInstrumentation did not equal execution.
Labelled future work.
No overclaimThe approved keyframe was retained; only the motion prompt was revised for the two weakest street clips.
Visible motion-energy ratioRetained v1 → promoted v2
Visible motion-energy ratioRetained v1 → promoted v2
This page demonstrates the system’s repair logic at workflow level: failed shots are identified, diagnosed, revised and either promoted or archived. It supports the dissertation claim that the workflow is agentic through traceable decision-making and targeted repair.
It does not demonstrate fully automatic Stage-4 acceptance, exact frame-level motion transfer or an independent equal-budget ablation. Human review remains the final acceptance authority in the completed evidence.
Each limitation is connected to the evidence that exposes it and a concrete next engineering step.
Unresolved research and engineering constraints.
Prototype feasibility is kept separate from production claims.
Executed cases, diagnostics and archived trials.
Each finding leads to a specific future action.
These limitations define the next engineering steps rather than invalidating the completed evidence. The current contribution is a traceable reference-driven workflow with clear evidence boundaries.
No exact frame-level motion transfer, fully automatic Stage-4 execution, full 27-second Mode C production or pose/mask-driven final output is claimed.
Use this page to locate artefacts; use the case studies to understand their meaning.
A direct inventory of the website's evidence stories and retained artefacts.
It provides a fast route from a research claim to the files that support or limit it.
Completed runs, repair records, audits, archived branches, Phase 0 and Stage-4 design status.
The index locates evidence; interpretation remains on the associated case-study page.
SVG architecture map
Shows mode logic, shared core and claim boundaries; explanatory, not an additional experiment.ArcFace, optical flow, pose/mask audit, framing review and timing audit
Shows how outputs were inspected and diagnosed; full automatic Stage-4 scoring was not fully executed.test_01, test_02, test_03 and shot/camera/pose/framing records
Shows how reference videos become generation units; exact frame-level motion transfer is not claimed.Authored prompt, storyboard/visual assets, separate ComfyUI generation and 71.6 s edited final video
Completed creative extension; not routed through shared generation units and has no per-shot decision-log / repair instrumentation.Source frames, 4 keyframes, 4 clips, final video, logs
Proves reference-scene preservation; exact motion transfer not claimed.Street scene-package run + archived pose-guidance branch
Explains the intended scene-replacement setting; the completed street final is not reference-video-driven.Scene refs, 4 keyframes, 4 clips, final, motion audit
Identity/scene stress test;reference_video = None.Decision logs, retained iterations, v1/v2 clips
Documents diagnosis and retry decisions, not full automation.12 samples, contact sheet, manifest, ZIP
9/12 candidates; diagnostic, not production hard control.Start/end preview and limitation report
Guidance strength was insufficient.Privacy-redacted control display, mask video, generated replacement and metadata
Shows longer-window driving-video replacement feasibility; full 27-second production not validated.Driving input, mask, short replacement output
Full 27 seconds not validated.Dry-run design and partial instrumentation
No claim of fully automatic acceptance or regeneration.ChatGPT-generated dataset sample, ComfyUI-trained face LoRA, 4 generated output samples
Independent of the Mode A/B/C pipeline; qualitative stability comparison only, no metric computed.Storyboard-driven ComfyUI branch and completed edited final
EXECUTED · SEPARATE COMFYUI PATHMode B was executed end-to-end as a storyboard-driven branch: authored story prompt, storyboard panels and visual planning assets were taken through ComfyUI generation and assembled into a 71.6-second edited final video (3840×2160, 30 fps, 4 Aug 2026). This demonstrates the second structure-source pathway in practice. However, this branch ran on a separate ComfyUI path rather than through the shared generation-unit pipeline, and it was not instrumented with per-shot decision logging or human-gated repair records. It is therefore reported as a completed creative extension, not at the same evidence maturity as the Mode A beach run.
Research role. Mode B constructs sequence structure from an authored prompt, storyboard panels and visual planning assets rather than extracting it from a reference video. Its completed final demonstrates that alternate structure-source path in practice without claiming validation of the shared Mode A generation-unit workflow.
Character, environment, storyboard and VFX material are grouped by their planning role. Every item below is an existing file copied without regeneration.
Character, expression and pose references for the Mode B animated storyboard branch.






Scene and environment references used for visual planning.
Storyboard panel references for the Mode B sequence.
Visual-effects and demo material for the Mode B branch.
Evidence boundary. Mode B completed an end-to-end creative path on ComfyUI, but it did not run through the shared generation-unit, evaluation and repair path. It has no per-shot decision-log or human-gated repair evidence, and is therefore not reported at the same evidence maturity as the completed Mode A beach run.
A locally trained face LoRA, tested against the ChatGPT-generated dataset it was trained on
SIDE EXPERIMENT · QUALITATIVE ONLYThis page documents a separate, author-run experiment: training a face LoRA in ComfyUI on a synthetic identity, then comparing the LoRA's own generations back against the dataset that trained it. The training images were not photographs of a real person — they were generated by the author using ChatGPT image generation, then curated down to a training set. This experiment sits outside the dissertation's reference-driven pipeline; it does not feed Mode A, Mode B or Mode C, and no claim is made that it was used to produce any of the beach, street or Mode C evidence elsewhere on this site.
Research role. The question this experiment asks is narrow: once a synthetic identity exists only as a set of ChatGPT-generated images, can a locally trained LoRA reproduce that identity as reliably as the model that generated the training set in the first place? The comparison below is a qualitative, human-reviewed one — no identity-similarity metric (ArcFace or otherwise) was computed across the LoRA outputs, so no quantitative stability score is claimed.
The dataset row shows the ChatGPT-generated training images the LoRA was built from. The generated row shows the LoRA's own output when run in ComfyUI after training.
All 40 curated training images, shown at thumbnail size to give the full picture of the target identity's look and consistency.








































Four samples generated by running the trained LoRA in ComfyUI, retained as-is without curation of a larger batch.
What this shows. Reviewed side by side, the LoRA's own generations track the trained identity's general look (hair, colouring, styling) but are visibly less consistent than the ChatGPT dataset that trained them — facial structure, lighting character and likeness drift more from image to image. This is reported as a human-observed qualitative comparison on a small, author-selected sample, not a measured stability score.
What this does not show. This experiment does not validate a production-ready identity LoRA, was not used anywhere in the Mode A, Mode B or Mode C evidence chain, and is not reported as dissertation evidence — it is retained here only as dashboard-level context on identity-consistency work attempted outside the main pipeline.
The dissertation contribution is a traceable orchestration system. The architecture is separated into an implemented evidence path, a logical modes map and four conditioning inputs.
The implemented orchestration architecture and bounded A–D mode taxonomy.
Separating structure, identity, look, scene, review and logs makes claims and failures traceable.
Code-native system diagram, five-track schema, model/tool stack and evaluation boundary.
Logical modes do not imply equal maturity; Mode D is planned and automatic Stage-4 is incomplete.
The reorganised figure separates mode maturity from the shared evidence workflow, conditioning inputs and claim boundaries. It is an explanatory system map, not an additional experiment.
The completed cases use a readable, shot-level path from a structural or scene input to reviewed final assembly.
The system converts a reference structure or scene package into inspectable shot-level records.
Each shot is regenerated from separated identity, styling and scene evidence, with one reviewed visual anchor per shot. Keyframe-first follows the pose-to-pose principle of classical animation (dissertation §2.1): a short-form clip is a few readable key poses joined by beat-timed cuts, so the pipeline fixes the pose first and delegates in-betweens to I2V.
Approved keyframes become prompt-level I2V clips. Human-gated review and repair records are logged before final assembly.
The implemented evidence preserves shot structure and action intent, not exact frame-level motion transfer.
Mode describes how the system operates. It does not imply that every branch has the same evidence maturity.
Primary dissertation path. Mode A-1 is completed; Mode A-2 is designed and partially tested.
A-1 completed · A-2 partialExecuted end-to-end on a separate ComfyUI path and assembled into a 71.6 s edited final. No shared generation-unit or per-shot repair instrumentation.
EXECUTED · SEPARATE COMFYUI PATHSeparate backend. Not part of the Mode A/B keyframe-first core.
Phase 0 + 7s window · quality WIPA 7-second local window completed on 4 Aug 2026; full 27-second production remains unvalidated.
Person + background replacement concept. Hosted-service demos were exercised on 4 Aug, but carry no evidential weight; no local implementation exists.
PLANNED · HOSTED DEMOS ONLYEach source has one clear responsibility, so a failure can be diagnosed against the correct control signal.
Shot order · framing · action intent
Subject identity and view-specific reference material
Outfit · hair · styling
Environment · viewpoint · lighting
Keeping these inputs separate makes provenance clearer and allows each failure to be diagnosed against the correct control signal. The street scene-package case used reference_video = None.
The anchor design separates structural, identity, styling and scene responsibilities so each shot can promote the evidence most relevant to its view.
The Look 3 identity is not one image but a purpose-built anchor set: a front-facing facial anchor, a side/profile anchor and outfit references — each authored to control a specific axis of appearance. Identity is anchored per view, so a profile shot is conditioned by the profile anchor rather than a stretched frontal reference.
Keyframe generation supplies four reference images in a locked slot order: slot [0] the source frame (primary structural anchor for framing, body scale and camera distance), then identity anchor, look reference and scene reference. Slot priority is shot-type-specific: face-visible shots promote the identity anchor; back-view and feet-detail shots promote the scene/source references. This ordering is a tested heuristic, not a model guarantee.
A shot is recorded as structured evidence rather than treated as one opaque prompt. Each field supports a different diagnostic question.
Pan, tilt, push, pull, tracking direction, intensity and stability.
Action intent, body displacement, direction and visible motion requirement.
Shot size, subject placement, orientation and camera relationship.
Source start/end, duration, beat position and assembly relationship.
Preserve/replace policy, environment lineage, lighting and viewpoint.
The five-track record is the hand-off between analysis, keyframe creation, I2V, review and repair. It does not claim exact motion copying.
The stack is grouped by research responsibility and evidence status, avoiding a dense implementation table.
Shot structure, timing, pose availability and post-hoc optical-flow diagnostics.
Executed in Mode ACreates a visual anchor from separated identity, look, scene and structural references.
Completed evidenceAnimates approved keyframes with prompt-level motion instructions and fixed clip duration.
Completed evidenceHuman acceptance supported by ArcFace and optical-flow post-hoc evidence.
Human-gatedDWPose and SAM3 support the separate short driving-video feasibility branch.
Phase 0 onlyThe public page shows what was executed without using the old automatic-loop wiring diagram as completed evidence.
Human review remained the acceptance authority for promoted keyframes, repaired clips and final assembly.
Human-gated repair evidenceAttempts, observed failures, diagnoses, selected repair actions and promoted results remain inspectable.
Browseable evidenceMetric-driven scoring and paid retry orchestration were designed and partially instrumented, but not fully executed.
Designed · not fully executedArcFace and optical flow are post-hoc diagnostics; framing is human-coded, generated-side pose may be pending, and formal beat alignment remains pending.
The original engineering diagrams remain preserved in the project’s paper-export artefacts; they are no longer used as the primary public explanation.
The prototype separates identity-adjacent styling references from environment references. This makes it possible to preserve a reference scene, or deliberately replace it with a static scene package, without confusing the two experiments.
Reusable identity-adjacent look packages and static multi-view scene packages.
Separated provenance lets the system distinguish source preservation from deliberate scene replacement.
Three look sheets, two scene packages and the preserved beach source scene.
These references condition approved keyframes; they are not proof of hard pose or motion control.
Shot order · framing · action intent
Subject identity and view-specific reference material
Outfit · hair · styling
Environment · viewpoint · lighting
Keeping these inputs separate makes provenance clearer and allows each failure to be diagnosed against the correct control signal.
These sheets are existing reference packages in the project. Look 3 is the styling package used in the completed beach and street evidence shown on this website.

Fitted activewear and a movement-oriented silhouette intended for dynamic full-body and studio shots.

Black velvet, organza structure and polished editorial styling designed for controlled dramatic scenes.

Charcoal tailoring, wide denim and natural styling. This is the selected look for both completed case studies.
Select an existing viewpoint to inspect each package. The street run uses the Parisian Street package with no reference video; the beach run preserves the scene from its source clip.
Warm, low-key editorial environment with spatial, lighting and reverse-angle references.

Cobblestone, limestone façades and soft daylight; used by the completed street stress test.

This distinction is central to the dissertation claim: Mode A-1 preserves the source environment, whereas the street stress test replaces the environment using a static package.

The source frame supplies beach setting, framing logic and action intent. The approved Look 3 keyframes replace the selected character/look while maintaining that scene context.
Beach keeps source-scene structure while replacing character/look. Street uses a static scene package with reference_video = None.
Look and scene references guide approved keyframes. Pose/mask extraction was diagnostic and did not act as production-ready hard control for the final outputs.
Current project_state.json — the pipeline's central source of truth.
Per-shot attempt records (scores, verdict, diagnosis, repair) and human-in-the-loop confirmation events.
Choose a character look for this run. Identity stays fixed. Look controls outfit, hair, makeup and visual mood.
Two built-in 3D animated character looks. Select the one you want for this run — it becomes the IP identity anchor for all generated panels.
Set the six run parameters before starting the pipeline: identity (IP character), look (outfit & styling), scene (target environment), mood / scene prompt, replacement scope, and output settings. Then run pre-flight to confirm readiness.
Confirm usage rights before analysis. Reference is analysed for formal structure, not reproduced.
Look package controls hair, makeup, outfit, shoes, silhouette and styling. Select or change look in Look Library.
Choose the target environment. The reference video controls pose/timing/framing — the scene package controls location, lighting and background.
Controls whether the source video's environment is preserved or replaced by the selected scene package.
This prototype supports one primary IP character per run. If the reference contains multiple people, the system selects one primary subject for structure extraction. Background people are ignored and will not be regenerated.
The reference video provides pose, timing, framing and camera structure. The selected identity, look and scene packages control what is regenerated.
Verify all prerequisites before running the pipeline. Blocked items must be resolved. Warnings allow proceeding with reduced reliability.
Decompose the reference video into a structured source frame storyboard. Shows per shot: source frame · pose overlay · camera motion (optical flow) · character motion · framing · timing · 5-track status. Tools: PySceneDetect (cuts) · librosa (beats) · MediaPipe (pose) · optical flow (camera motion) · GPT-4o Vision (semantic enrichment).
test_03_multishot_edit_4shots.mp4| Reference video | Type | Resolution | Duration | Detected shots |
|---|---|---|---|---|
| ▶ test_01_static_camera_subject_move.mp4 | single-take · static camera | 1080×1920 | 9.3 s | 1 shot |
| ▶ test_02_camera_push_pan.mp4 | single-take · push / pan | 1920×1080 | 7.8 s | 1 shot |
| ▶ test_03_multishot_edit_4shots.mp4 main run | multi-shot edit | 1916×1080 | 7.9 s | 4 shots |
Build the source frame storyboard from analysis results — one extracted frame per detected shot, with framing, pose, timing, and beat labels. Human reviews this before exporting the structured JSON.
Redraw each reference panel as an IP-consistent keyframe (gpt-image-1). Preserves framing, camera angle, body orientation.
Animate each IP keyframe with Kling image-to-video. Generated per-shot, never as one continuous video.
Scores each shot at the keyframe / image level — identity (ArcFace cosine), pose (keypoint distance), framing, beat alignment on the still keyframe. This is not full video-level temporal evaluation. For reference-vs-generated video motion, see Evaluation Report → Reference vs Generated Video Audit.
Diagnose failure reason per shot → apply targeted repair → re-evaluate. Max 3 attempts. Terminal state → needs_human.
| Failure Type | Diagnosis Signal | Repair Action | Observed Change |
|---|---|---|---|
| Identity drift | ArcFace score < 0.80 | Re-draw keyframe with raised LoRA weight | Keyframe regenerated; video re-run |
| Pose mismatch | Pose score < 0.85 | Strengthen pose conditioning (increase ControlNet weight) | Keyframe reused; video re-run with stronger pose |
| Framing error | Framing score < 0.75 | Rewrite framing prompt (add explicit composition instructions) | Prompt updated; keyframe + video regenerated |
| Beat drift | Beat error > 250 ms | Adjust cut timing to nearest beat within snap distance | Assembly timing adjusted; no regeneration |
| Multiple failures | ≥2 metrics fail | Prioritise identity fix first, then pose | Sequential repair attempts |
Assemble approved shot clips into a 9:16 vertical short video. Add audio and export.
Approved shots ready for assembly. Shots with status needs_human are excluded.
Completed evidence uses human-gated diagnose-and-retry with a decision log. Automatic Stage-4 comparison is designed and partially instrumented, but not fully executed.
final_look3_reference_driven_demo.mp4 · Look 3 · beach scene · 4 shots · 9:16.
The completed street run used the static street scene package with reference_video = None. It produced four approved keyframes, four Kling clips and final_look3_street_demo.mp4, with human review and a decision log. shot_001 & shot_003 received a motion_v2 prompt-level repair and were promoted into the final. This is positive identity-and-scene stress-test evidence, not a completed reference-video-driven structure-transfer result.
| shot | v1 | v2 | status |
|---|---|---|---|
| shot_001 | 14% | 51% | repaired · energetic head turn |
| shot_002 | 37% | — | kept · reference itself calm |
| shot_003 | 19% | 47% | repaired · clearer step cycle |
| shot_004 | 64% | — | kept |

final/motion_audit/reference_vs_street_motion_report.md · street_motion_metrics.jsonThe Evaluation ↔ Repair loop (05→06→03) is what makes this system agentic.
| Aspect | Mode A | Mode B |
|---|---|---|
| Starting point | Reference video | Natural language story prompt |
| Character source | Fixed IP identity + selectable look package | Built-in 3D animated character look |
| Storyboard source | Extracted from video | Generated from panel plan |
| Structure source | Detected cuts and beats | 4-panel animated storyboard template |
| Source panels | Extracted keyframes | Generated storyboard panels |
| Output | Same Storyboard JSON | Same Storyboard JSON |
Both modes are converted into the same storyboard JSON. From the keyframe stage onward, they use the same generation, evaluation and repair process.
| Step | Tool | Used for |
|---|---|---|
| Story planner | GPT-4o | Converts story prompt into panel_plan.json |
| Storyboard sheet | gpt-image-1 | Generates 4-panel animated storyboard contact sheet |
| Panel review | Human review | Approves or edits panels before JSON export |
| Shot cuts | PySceneDetect | Finding shot boundaries |
| Beats | librosa | Reading rhythm and beat timing |
| Pose | MediaPipe | Extracting body keypoints |
| Framing | Bounding boxes | Checking face and body placement |
| Keyframes | OpenAI image backend | Whole-person IP regeneration (face, hair, outfit, shoes, body) |
| Video | Kling API | Turning keyframes into short clips |
| Identity | ArcFace / InsightFace | Face embedding cosine similarity vs IP references |
| Look eval | GPT-4o vision | Hair, outfit, shoes and palette consistency check |
| Repair | Rule-based planner | Targeted fix per failure type (identity, pose, framing, look) |
needs human.Post-hoc comparison only: generated street clips are compared with the beach clip using Farnebäck optical flow at a common height, native aspect preserved. Because the completed street run used reference_video = None, this audit does not establish reference-video structure transfer. It reports motion-energy / amplitude differences, not exact frame alignment or pixel-level correspondence.
| shot | reference avg (%h/frame) | generated (street) | motion recovered | reference motion character |
|---|---|---|---|---|
| shot_001 | 0.336 (pk 1.48) | 0.047 | 14% | energetic head-turn toward camera |
| shot_002 | 0.119 (pk 0.49) | 0.044 | 37% | calm slow walk away (calmest reference shot) |
| shot_003 | 0.753 (pk 2.57) | 0.141 | 19% | most active — stepping feet + turn |
| shot_004 | 0.273 (pk 0.83) | 0.175 | 64% | lateral walk with camera movement |
shot_001 and shot_003 were re-prompted (damping words removed, motion cues named) and have been promoted into the official street run — the final now uses the v2 clips for these two shots (14→51% / 19→47% motion-energy recovery). The v1 clips are retained as backups.
Archived-log terminology. “Real footage” in the archived audit log is shorthand for the motion-target role. The reference clip is the author-created AI-generated beach performance clip (dissertation §5.3); the logs are preserved unedited.
Future-work ablation design: compare Loop OFF and Loop ON under identical conditions. It has not been fully executed and is not presented as dissertation result evidence here.
Designed one-shot baseline with repair disabled. DISPLAY ONLY · FUTURE WORK.
Proposed metric-driven diagnosis and retry framework. DESIGNED · NOT FULLY EXECUTED.
| Metric | Loop OFF | Loop ON | Δ | Better |
|---|---|---|---|---|
| Formal automatic ablation not fully executed. Simulated records are excluded from dissertation evidence. | ||||
Person replacement driven by a performance video (Condition A: replace the performer, preserve the original background, camera, duration and timing). Phase 0 passed — a short low-res single segment ran end-to-end on the 16 GB GPU (via Wan2GP, int8 + CPU-offload), producing a background-preserving character replacement that follows the driving motion. Quality (proportions/identity on extreme poses) and the full 27 s loop are still WIP.
Phase 0 result: PASSED. A ~4 s low-res (480×832, 81 frames, int8) single segment ran end-to-end on the 16 GB RTX 4080 via Wan2GP. Outcome: it runs on 16 GB ✓, the generated character follows the driving motion ✓ (squat sequence transferred), the original background is preserved ✓ (room/floor unchanged outside the person mask), and identity/proportions distort on extreme poses (the deep-squat frame) — a quality limitation, not a feasibility failure. Full 27 s generation and quality tuning (more steps/res, relighting) remain WIP.
Phase 0 已通过:约 4 秒低分辨率单段在 16GB RTX 4080 上(经 Wan2GP,int8+offload)端到端跑通 —— 能替换人物、保留原背景、跟随动作;极端姿势下比例/身份会漂移(质量问题,非可行性问题)。完整 27 秒与画质优化仍在进行中。
Driving performance (control) → SAM3 person mask (white = replace, black = keep) → Wan-Animate replacement with the Look reference. This preserves action intent + background, not exact frame-level transfer.
Wan2GP on w33108. Phase 0 (short low-res segment) passed — see evidence above. Quality tuning + the full 27 s loop are next; the in-dashboard full-generation button stays disabled (generation runs on the GPU host, not here).Driving video mode_c.MP4 (27 s, ~30 fps) is on file; the remote GPU host was verified and Phase 0 was executed there. No generation runs from this dashboard host. Full repository audit + phase plan: docs/MODE_C_CONDITION_A_AUDIT.md.
Designed one-shot baseline with repair disabled. This page does not run or create dissertation evidence.
Designed experiment for metric-driven evaluation and repair. It is partially instrumented, but it is not the executed evaluation of record for the completed evidence.
Planned Loop OFF vs Loop ON comparison schema. Simulated records are not shown as dissertation evidence.
| Metric | Loop OFF | Loop ON | Improvement |
|---|---|---|---|
| No data yet | |||
Choose your 3D character, describe your animated idea, then generate a storyboard contact sheet.
One sentence is enough — click 🎬 Expand to let the director AI build the full arc. Or write it yourself.
Generate an 8-panel director's plan with GPT-4o, then draw the 4×2 contact sheet with gpt-image-1.
Uses panel_plan.json to draw all 8 panels as a single previs contact sheet (1536×1024).
Review each panel: beat role, action, camera, expression, and start / middle / end video prompts.
Compile annotated panels into structured JSON. Each panel becomes an animated_panel generation_unit — same schema as Mode A, shared downstream from here.
Experimental storyboard JSON → paid keyframe generation → paid Kling I2V → partially instrumented evaluation / repair design.
→ View results in Overview