Deep Video Spoof Generation Based on Matting Technique
DOI:
https://doi.org/10.29304/jqcsm.2026.18.32718Keywords:
deep video spoofing, deepfake, video matting, image mattingAbstract
Deep video spoof generation has made significant progress by using well-designed face-swapping models with high-level identity transfer and visual realism. However, realistic deep video spoof generation from arbitrary videos remains difficult due to visible compositing artefacts both locally (e.g., hairline and semi-transparent regions) and temporally (flicker and drift). We consider deep video spoof generation, a focus on which is chiefly underpinned by matting, as an essential building block for boundary-accurate and temporally consistent insertion of manipulated foreground into real scenes with no requirement for the presence or recording of green screens. We highlight some key directions in image, background, and video matting: (1) background-referenced matting; (2) trimap-free and real-time portrait matting; and (3) temporally stable video matting methods that exploit spatio-temporal alignment as well as memory propagation. We also consider recent modelling trends, including transformer-based and diffusion-assisted matting, which provide better fine-detail reconstruction and interactive refinement. The analysed studies show that matting can alleviate key visible failure modes in spoof generation pipelines (which are improved alpha estimation quality, boundary sharpness, and temporal consistency) in all works we surveyed. Based on the observations, we finally summarise the open issues—robustness to occlusion and fast motion, generalisation from composited training to real-world scenes, computation overhead under high resolution, and unified evaluation protocols—and propose future directions for research towards more realistic and stable matting-based deep video spoof generation.
Downloads
References
R. Chen, X. Chen, B. Ni, and Y. Ge, “SimSwap: An efficient framework for high fidelity face swapping,” in Proc. ACM Int. Conf. Multimedia (ACM MM), 2020, pp. 2003-2011, doi: 10.1145/3394171.3413630.
X. Zhu et al., “One-shot face swapping on megapixels,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2021, doi: 10.1109/CVPR46437.2021.00480.
R. Dhanyalakshmi, G. Stoian, D. Danciulescu, and D. J. Hemanth, “A survey on face-swapping methods for identity manipulation in deepfake applications,” IET Image Processing, vol. 19, no. 1, art. no. e70132, 2025, doi: 10.1049/ipr2.70132.
S. A. R. Syed Abu Bakar, S. Waseem, Z. Omar, et al., “Exploring the advancements and challenges of deepfake face-swap: A survey,” Multimedia Tools and Applications, vol. 85, art. no. 14, 2026, doi: 10.1007/s11042-026-21285-8.
S. Sengupta, V. Jayaram, B. Curless, S. M. Seitz, and I. Kemelmacher-Shlizerman, “Background matting: The world is your green screen,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020, pp. 2291-2300, doi: 10.1109/CVPR42600.2020.00236.
S. Lin, L. Chen, Y. Wang, Z. Luo, and Y. Tai, “Real-time high-resolution background matting,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2021, pp. 8762-8771, doi: 10.1109/CVPR46437.2021.00866.
Z. Ke et al., “MODNet: Real-time trimap-free portrait matting via objective decomposition,” in Proc. AAAI Conf. Artificial Intelligence (AAAI), vol. 36, no. 1, 2022, pp. 1140-1147, doi: 10.1609/aaai.v36i1.19979.
J. Sun, J. Wang, Z. Tang, H. Guo, and C.-K. Tai, “Deep video matting via spatio-temporal alignment and aggregation,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2021, pp. 6975-6984, doi: 10.1109/CVPR46437.2021.00690.
P. Yang, S. Zhou, J. Zhao, Q. Tao, and C. C. Loy, “MatAnyone: Stable video matting with consistent memory propagation,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025, pp. 7299-7308.
J. Sun et al., “Ultrahigh resolution image/video matting with spatio-temporal sparsity,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023, pp. 14112-14121, doi: 10.1109/CVPR52729.2023.01356.
L. Huang et al., “SDMatte: Grafting diffusion models for interactive matting,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025, pp. 15229-15239, doi: 10.48550/arXiv.2508.00443.
G. Park et al., “MatteFormer: Transformer-based image matting via prior-tokens,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022, pp. 11696-11706, doi: 10.1109/CVPR52688.2022.01140.
J. Yao et al., “ViTMatte: Boosting image matting with pre-trained plain vision transformers,” Information Fusion, vol. 103, art. no. 102091, Mar. 2024, doi: 10.1016/j.inffus.2023.102091.
J. Li et al., “ProxyMatting: Transformer-based image matting via region proxy,” Knowledge-Based Systems, vol. 310, art. no. 112911, Feb. 2025, doi: 10.1016/j.knosys.2024.112911.
Y. Zhou et al., “Sampling propagation attention with trimap generation network for natural image matting,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 10, pp. 5828-5843, Oct. 2023, doi: 10.1109/TCSVT.2023.3260025.
X. Ye, Y. Liang, M. Tan, F. Feng, L. Wang, and H. Huang, “High-resolution natural image matting by refining low-resolution alpha mattes,” IEEE Transactions on Image Processing, vol. 34, pp. 3323-3335, 2025, doi: 10.1109/TIP.2025.3573620.
X. Fang, S.-H. Zhang, T. Chen, X. Wu, A. Shamir, and S.-M. Hu, “User-guided deep human image matting using arbitrary trimaps,” IEEE Transactions on Image Processing, vol. 31, pp. 2040-2052, 2022, doi: 10.1109/TIP.2022.3150295.
Y. Xu, B. Liu, Y. Quan, and H. Ji, “Unsupervised deep background matting using deep matte prior,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 7, pp. 4324-4337, Jul. 2022, doi: 10.1109/TCSVT.2021.3132461.
Y. Wang et al., “From composited to real-world: Transformer-based natural image matting,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 4, pp. 2097-2111, Apr. 2024, doi: 10.1109/TCSVT.2023.3300731.
B. Peng et al., “RGB-D human matting: A real-world benchmark dataset and a baseline method,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 8, pp. 4041-4053, Aug. 2023, doi: 10.1109/TCSVT.2023.3238580.
L. Dai et al., “Enabling trimap-free image matting with a frequency-guided saliency-aware network via joint learning,” IEEE Transactions on Multimedia, vol. 25, pp. 4868-4879, 2023, doi: 10.1109/TMM.2022.3183403.
G. Yao and A. Sun, “Multi-guided-based image matting via boundary detection,” Computer Vision and Image Understanding, vol. 243, art. no. 103998, Jun. 2024, doi: 10.1016/j.cviu.2024.103998.
A. Sun, J. Chang, and G. Yao, “DiffMatter: Different frequency fusion for trimap-free image matting via edge detection,” Computer Vision and Image Understanding, vol. 259, art. no. 104424, Sep. 2025, doi: 10.1016/j.cviu.2025.104424.
J. Yao, X. Wang, L. Ye, and W. Liu, “Matte anything: Interactive natural image matting with segment anything model,” Image and Vision Computing, vol. 147, art. no. 105067, Jul. 2024, doi: 10.1016/j.imavis.2024.105067.
D. C. Lepcha et al., “Image Matting: A comprehensive survey on techniques and applications,” International Journal of Image and Graphics, 2023, doi: 10.1142/S0219467823500110.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Nabaa Maad Abd AL-Kareem, Saud Jamila H

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.








