← Back to Blog
Research RadarFace SwappingVisual EditingarXivSeptember 2026

Monthly arXiv Radar

September 2026 Face Swapping Papers: Lighting Harmony, Reference Attention, and Real-Time Dubbing

September offers one direct face-compositing method and two adjacent identity-editing systems. The shadow-harmonization pipeline deliberately limits what it can change so a bad estimate cannot recolor or brighten an approved face. RefGAP instruments reference attention across seven diffusion editors and corrects it at inference time. TBDub adapts a visual-dubbing model to production footage, then distills thirty denoising steps to two while measuring identity, lip sync, visual quality, and end-to-end throughput.

What This Month Signals

The geometry-driven harmonizer applies only a bounded, channel-uniform darkening field; on its analytic proxy, 56.5% of face pixels change, mean gain is 0.938, and no pixel is brightened, though the paper does not claim a real-image perceptual win. RefGAP calibrates two coefficients once, then improves identity similarity across seven image/video editors on 1,040 head-swap clips and 1,000 face-swap pairs; replacement success reaches at least 86% and 82%, respectively, with measurable pose and preservation trade-offs. TBDub's two-step student reaches 7.13 FPS at 512 by 512 on an H20, a 13.93-fold end-to-end speedup over its teacher, while human ratings retain the teacher's identity quality and slightly improve lip sync and visual quality. For buyers, the important distinction is risk: deterministic bounds, inference-time controls, and distilled generation have different failure modes and validation needs.

Paper 012026-09-15cs.CV

Geometry-Driven Shadow Harmonisation for Composited Faces: A Multiplicative, Albedo-Preserving Relighting Pipeline

Authors & Institutions

Vijesh KP

Independent researcher (no institutional affiliation stated in the paper)

What Problem It Solves

The paper targets the narrow case where donor color and identity are acceptable but form shadows are missing. It asks how to inject nose, socket, lip, and cavity shadows from approximate face geometry without brightening pixels, changing chromaticity, or synthesizing a new face.

Key Result

On the paper's analytic face heightfield with a known light, the default setting modifies 56.5% of face pixels, gives mean gain 0.938 overall and 0.890 on modified pixels, and sends 3.4% of pixels to the 0.82 floor. It guarantees zero brightened pixels and no mathematical hue-angle shift beyond floating-point noise. These figures characterize the bounded operator rather than perceptual quality: the paper provides no paired real-composite benchmark or baseline comparison. It cannot add highlights, correct an already dark donor, remove opposing source light, or avoid double-shadowing residual donor shading.

Abstract

Face swapping and face compositing pipelines routinely produce a face that is geometrically well aligned but photometrically implausible: the donor face carries flat, near-frontal studio illumination while the host body and background carry directional scene light. Most existing remedies re-synthesise the face through colour transfer, neural relighting, or inverse rendering, and therefore risk altering identity, skin tone, and texture. We present a conservative alternative: geometry-driven form-shadow injection. The pipeline never repaints the face. It estimates a per-pixel gain field $g\in[g_{\min},1]$ from a rasterised 3D face proxy and multiplies it channel-uniformly onto linear RGB, so the operator can only darken and cannot shift chromaticity. A dense landmark mesh is rasterised into a depth buffer, from which we derive surface normals, a cavity term, and screen-space cast shadows. Key-light direction is estimated from host-side cues (body, background, hair halo); on-face cues are downweighted because they recover the donor's lighting. Shadow magnitude is not matched to the host: it is set by a three-parameter transfer $(τ,σ,g_{\min})$. The shading field is divided by its 75th percentile over skin, then gated, scaled, clamped, smoothed, and re-clipped inside a feathered, skin-gated face mask. On an analytic face heightfield, the default $(τ,σ,g_{\min})=(0.90,0.45,0.82)$ modifies 56.5% of face pixels with mean gain 0.938 (0.890 on modified pixels) and drives 3.4% of pixels to the floor. Hue invariance is a corollary of the operator. We analyse the transfer in closed form, ablate its parameters, and discuss failure modes of a monotone, darkening-only formulation, including double-shadowing of non-flat donors.

Research Starting Point

A swapped face can be geometrically aligned and identity-correct yet still look pasted on because the donor carries flat studio light while the host scene has a directional key. Generative relighting can repair more effects, but it may also change skin tone, pores, texture, or identity after the face has already been approved. Production teams sometimes need a deliberately limited operator whose worst-case photometric change is known.

Method

A face detector and 468-point mesh produce a rasterized depth buffer, surface normals, cavity proxy, and screen-space cast shadows in a 512-pixel crop. Host-side body, background, neck, and hair-halo cues vote for key-light direction while donor-face cues are downweighted. The normalized shading map passes through a three-parameter threshold, strength, and minimum-gain transfer, then guided smoothing and a feathered skin mask. The same scalar gain is multiplied into all linear-RGB channels and clipped to a range from 0.82 to 1.0.

Paper Summary

This is a conservative post-compositing control with an understandable damage bound, not a complete relighting system. It is attractive when identity preservation outranks dramatic realism, provided real shots are reviewed for wrong light direction and double shadows.

Paper 022026-09-28cs.CV

Mind the RefGAP: Correcting Reference Attention in Diffusion-Based Visual Editing

Authors & Institutions

Yanan Wang

Mohamed bin Zayed University of Artificial Intelligence

Institute of Foundation Models

Shengcai Liao

United Arab Emirates University

Guangyi Liu

Mohamed bin Zayed University of Artificial Intelligence

Institute of Foundation Models

Xiaodan Liang

Mohamed bin Zayed University of Artificial Intelligence

What Problem It Solves

RefGAP aims to improve reference identity fidelity across different image and video editors without retraining and without sweeping a separate correction strength for every model. It explicitly balances stronger reference use inside the edit mask with suppression in regions that should be preserved.

Key Result

On 1,040 HeadSwapBench clips, identity similarity improves for all seven general-purpose editors and replacement success reaches at least 86% wherever defined; whole-head similarity also rises for all seven. On 1,000 FaceForensics++ face-swap pairs, identity similarity improves across every editor, with replacement and correct retrieval at least 82% and 75%. Applied to the specialized DirectSwap model, identity similarity rises from 0.7019 to 0.7450 while whole-frame LPIPS is nearly unchanged. Gains come with higher keypoint error for several systems and mixed non-edit preservation; background replacement does not transfer reliably. The unfused attention implementation can also add substantial memory or latency, so integration quality matters.

Abstract

Reference-guided diffusion editors struggle to faithfully reproduce user-provided references. We identify a potential bottleneck in diffusion editors: many methods provide limited reference-attention allocation. For example, in LoomVideo, edit-region queries assign less than 1% of their attention mass to the reference. We introduce RefGAP, a training-free correction that determines logit-offset magnitudes online at each layer from the reference-attention mass measured during the forward pass. Positive offsets to reference logits strengthen reference usage by edit-region queries, while negative offsets for keep-region queries limit reference-induced changes outside the edit. Two global coefficients control the correction; they are selected once on validation data from four development diffusion editors and held fixed. Across seven diffusion-based image/video editors, RefGAP improves identity fidelity in head swapping and face swapping. RefGAP achieves a fidelity-preservation trade-off comparable to separately tuned constant edit-side biases, without per-approach strength sweeps. Additional experiments on virtual try-on and background replacement evaluate transfer beyond identity editing.

Research Starting Point

Reference-guided diffusion editors may receive the correct source face yet devote almost no attention to it; LoomVideo edit-region queries allocate less than 1% of attention mass to the reference in the authors' measurement. Fixed attention boosts can strengthen identity but require per-model tuning and may leak reference appearance into hair, clothing, background, pose, or expression that should remain unchanged.

Method

During each attention layer, RefGAP measures the share of attention assigned to reference tokens separately for edit and keep queries. It derives an online logit offset from the observed mass: positive offsets strengthen reference keys in the edit region, while negative offsets constrain them in the keep region. Two global coefficients and a layer-selection rule are calibrated once on 40 head-swap clips using four development editors, then held fixed across seven editors, three held-out systems, face swapping, virtual try-on, and background replacement.

Paper Summary

Measuring actual reference attention is a useful model-agnostic control signal for identity edits, but stronger identity transfer can alter geometry and preserved regions. Teams need a task-specific quality gate, not identity similarity alone.

Paper 032026-09-05cs.CV

TBDub: Production-Oriented Visual Dubbing

Authors & Institutions

Bihan Li

TaoLive AIGC, Taobao & Tmall Group, Alibaba

Xinyang Li

TaoLive AIGC, Taobao & Tmall Group, Alibaba

Zeran Xu

TaoLive AIGC, Taobao & Tmall Group, Alibaba

Meiguang Jin

TaoLive AIGC, Taobao & Tmall Group, Alibaba

Junfeng Ma

TaoLive AIGC, Taobao & Tmall Group, Alibaba

What Problem It Solves

TBDub adapts an existing video diffusion dubbing stack to production footage and asks whether a two-step student can retain the teacher's perceived identity, synchronization, and visual quality. It evaluates reconstruction, audio response, human ratings, and the complete VAE-to-VAE generation path rather than quoting denoiser speed alone.

Key Result

On 38 TalkVid clips, the teacher improves all eight reported reconstruction, perceptual, identity, and synchronization metrics over X-Dub. In 114 ratings per method, X-Dub's lip-sync, identity, and visual-quality MOS values of 3.57, 2.83, and 2.88 rise to 3.71, 3.78, and 3.78 for the teacher. The student scores 3.85, 3.72, and 3.80, keeping identity close while leading lip sync and visual quality. At 512 by 512 on one H20, the student reaches 7.13 effective FPS versus 0.51 for the teacher, a 13.93-fold end-to-end speedup; the DiT stage alone is 42.49 times faster. The benchmark is small, excludes several preprocessing and encoding stages, and very low-resolution inputs remain a failure case.

Abstract

Visual dubbing must synchronize mouth motion with replacement speech while preserving identity, appearance, and temporal consistency. Although X-Dub provides a strong mask-free video-editing baseline, its application to livestream and generated-video content reveals limitations in production-domain robustness, temporal and motion stability, identity and oral-detail preservation, and inference efficiency. We present \textbf{TBDub}, a production-oriented extension of X-Dub that combines task-adaptive post-training with task-aware few-step distillation. Post-training adapts the video DiT using production-domain data, production-specific conditioning and filtering, and enhanced audio features to obtain a 30-step Teacher. Distillation adapts DMD/DMD2 to conditional video editing and compresses the Teacher into a two-step Student. On 38 TalkVid clips, the Teacher improves all eight reported reconstruction, perceptual, identity, and synchronization metrics over X-Dub. In the MOS evaluation, it improves lip-sync consistency, identity consistency, and visual quality over X-Dub by 0.14, 0.95, and 0.90 points, while the Student achieves the highest lip-sync and visual-quality scores and remains close to the Teacher in identity consistency. In paired end-to-end generation timing from the first VAE encode through the final VAE decode on a single NVIDIA H20 GPU at $512\times512$, the Student reaches 7.13 effective FPS and reduces total latency by $13.93\times$; the DiT stage alone is accelerated by $42.49\times$. The Student largely retains the Teacher's generation quality and audiovisual synchronization. The code is available on GitHub at \https://github.com/TaoLiveAIGC/TBDub, and the 30-step Teacher and two-step Student weights are available on Hugging Face at https://huggingface.co/TaoLiveAIGC/TBDub.

Research Starting Point

Visual dubbing must change mouth motion enough to express replacement speech while keeping identity, teeth, lip texture, pose, illumination, occlusion, and every non-speech region stable over time. X-Dub provides a capable mask-free video editor, but production livestream and generated-video inputs expose domain shift, flicker, oral-detail drift, and the latency of thirty denoising steps.

Method

Task-adaptive post-training mixes production-domain data with the inherited training distribution, strengthens audio features with HuBERT-large layers, degrades conditions asymmetrically, and uses motion dropout plus spatial and temporal weighted flow matching to create a 30-step teacher. A conditional DMD/DMD2 procedure distills that teacher into a differentiable two-step student, with dynamic fake-score training, staged region-selective supervision, causal first-latent handling, full-timestep sampling, FP32 guidance arithmetic, and FP32 EMA weights.

Paper Summary

Two-step visual dubbing can offer a much better quality-throughput point, but the reported 7.13 FPS is bounded to the generation core. Production sizing should include audio, face tracking, compositing, encoding, and long-video stability.