Qwen-Image-2.1-Multiple-Angles-LoRA - v2 1500

Qwen-Image-2.1-Multiple-Angles-LoRA

LORA
Reprint


Updated:

 Qwen-Image-2.1-Multiple-Angles-LoRA  by MacrossManiac on Tensor.Art

Qwen-Image-2.1 · Multiple-Angles camera-control LoRA

Reshoot any object from a controlled camera position: feed a photo, ask for a different azimuth/elevation, get the same subject from that angle. First camera-control LoRA for Qwen-Image-2.1 — the newest Qwen image base (unified generate + edit, native RGBA), where this format did not exist yet.

  • Trigger: <mva> — every caption/prompt starts with it

  • 36 poses: 12 azimuths (30° steps) × 4 elevations (0° eye-level / 30° elevated / 60° high-angle / 90° top-down), plus a close-up variant of each → 72 addressable framings

  • Strength: 0.8–1.0 in ComfyUI / diffusers

Checkpoint evidence (stated plainly): two independent reads agree on step 1,500. (a) The run's per-250-step sample grids (samples/qwen21_multiple_angles_v2/samples/) — angle-following solid from step 1,000, peak ~1,000–2,000, slight orientation drift at 2,500. (b) An out-of-distribution real-photo curve (benchmarks/v2_curve/ — cat, dog, and a third held-out subject, v1@1000 vs v2@500→2,500 at the same prompt and seeds): v2 holds identity and background detail at least as well as v1@1000 on all three subjects, with no drift visible at step 2,500 (the drift seen in the training-sample grid does not reproduce on real photos). Step 1,500 stays the recommendation as the balanced point; 2,000–2,500 are usable for stronger angle-baking.

Character test (benchmarks/v2_verify/character_sheet.jpg): three rigged characters from the Linzhan training data — a subject class v1 never saw — rendered at back view and top-down with v1@1000 vs v2@1500. v1 attempts the poses (rendered characters resemble dome objects) but softens and drifts on character-specific structure; v2 holds identity and sculpted detail (hair spikes, creature forms) more faithfully at the hard poses. Honest scope note: these are in-distribution subjects for v2, so this verifies the character data took; generalization to arbitrary real-world characters rides on the OOD curve above.

Read this honestly — the metric, not the LoRA, is the story. CLIP image–image cosine against a same-object ground truth is nearly viewpoint-invariant: any two views of the same object on a similar background score 0.8–0.95 whether or not the camera actually moved. So this table measures global image similarity (background, composition, polish) — on which the un-LoRA'd base produces more globally-similar images. It cannot distinguish "moved the camera correctly" from "stayed put", which is the capability this LoRA adds (see the visual A/B above: base's "back view" still shows front-facing features the requested view cannot contain).

The right quantitative protocol — and the reason these ground-truth renders are published — is pose-retrieval accuracy: score each generation against all six ground-truth poses of its object and check that the requested pose wins. That v2 scoring needs one more render pass; the dataset, the prompt grammar, and the held-out protocol here make it fully reproducible by anyone.

For the capability itself, the controlled visual evidence above is the primary demonstration: matched prompts, same control images, base vs LoRA.

Prompt format

<mva> {azimuth}, {elevation}[ close-up]

Azimuth (12 steps of 30°)

Label Azimuth front view 0° front-right view 30° front-right quarter view 60° right side view 90° back-right quarter view 120° back-right view 150° back view 180° back-left view 210° back-left quarter view 240° left side view 270° front-left quarter view 300° front-left view 330°

Elevation

Label Elevation eye-level shot 0° elevated shot 30° high-angle shot 60° top-down shot 90°

Append close-up for the 62%-crop framing. Example:

<mva> right side view, high-angle shot close-up

Usage

Gradio Workflow Space — akhaliq/qwen21-multiple-angles-workflow: canvas UI, photo in → camera knobs out.

ComfyUI — load the LoRA with any Qwen-Image-2.1 workflow (LoRA strength 0.8–1.0), pass your photo as the edit/reference image, prompt with the format above.

diffusers (QwenImage21Pipeline, edit mode = pass the image):

image = pipe( prompt="<mva> back view, eye-level shot", image=input_image, # front view of your subject num_inference_steps=40, true_cfg_scale=1.0, ).images[0]

Training

  • Data: v1 — 5,028 angle-labeled edit pairs from zeyuanyin/Dome-Objaverse (CC-BY-4.0). v2 — 13,328 pairs: same source with character-prioritized selection (9,376), plus 1,986 rigged characters from Linzhan/Objaverse-XL-Rigged-Animated-Renders, license-filtered to CC-BY / CC-BY-SA / CC0 upstreams (3,952 pairs). Published: v1 pairs · v2 pairs.

  • Trainer: ai-toolkit, arch: qwen_image_2, control-image conditioning (block-causal reference attention). v1: rank 32, convrot8 int8 base, shift timesteps. v2: rank 64, bf16 full-precision base, weighted timesteps. Both: adamw8bit, lr 1e-4, caption dropout 0.05, 512², 2,500 steps.

  • Checkpoints: v1 in checkpoints/, v2 in checkpoints_v2/ — every 250 steps + samples, so the training curve is inspectable end to end.

Scope & limits

  • Trained on Objaverse objects; real photos generalize (see Real photos above) but identity fidelity on fur/fine detail is good-not-perfect. v2 targets this with character-prioritized data and full-precision training — OOD verification published in benchmarks/v2_curve/ (see v2 section for the verdict).

  • Trained at 512²; larger canvases inherit the style but were not trained on.

  • Angles interpolate reasonably between the trained 30° steps; exact azimuth/elevation numbers in prompts are not parsed — use the labels above.

  • The CLIP image–image benchmark above does not measure pose accuracy (see its section) — the demonstrated capability rests on the controlled visual A/B.

Credit to fal's Multiple-Angles LoRA for Edit-2511 for the caption grammar this format standardized on.

Version Detail

Qwen-Image-2.1
Prompt format <mva> {azimuth}, {elevation}[ close-up] Azimuth (12 steps of 30°) Label Azimuth front view 0° front-right view 30° front-right quarter view 60° right side view 90° back-right quarter view 120° back-right view 150° back view 180° back-left view 210° back-left quarter view 240° left side view 270° front-left quarter view 300° front-left view 330°

Project Permissions

Model reprinted from : https://huggingface.co/akhaliq/Qwen-Image-2.1-Multiple-Angles-LoRA

Reprinted models are for communication and learning purposes only, not for commercial use. Original authors can contact us to transfer the models through our Discord channel --- #claim-models.

Related Posts