Low-Overhead Adaptive Multi-Scale Structure-Preserving Refinement for Pretrained Generative Photo Cartoonization

Main Article Content

Kavali Manisha
A. Soujanya

Abstract

Photo cartoonization transforms photographic content into a simplified artistic representation while attempting to
preserve scene geometry, salient contours, and recognizable object structure. Pretrained generative models provide an efficient
route to visually convincing stylization; however, local boundaries can be weakened during transformation, particularly in
texture-rich scenes and regions containing thin structural details. This work presents a low-overhead adaptive multi-scale
structure-preserving refinement framework that operates after pretrained cartoon generation without retraining the underlying
model. Structural evidence is extracted from the original image using Gaussian filtering at three spatial scales, adaptive Canny
edge estimation, confidence-map aggregation, and edge-density analysis. An adaptive fusion coefficient then regulates boundary
reinforcement according to scene complexity. Local luminance anchoring is subsequently used to control excessive darkening,
followed by luminance-channel contrast refinement in a perceptual color space. A reproducible pilot evaluation was conducted on
24 samples derived from 12 public images at a resolution of 512×512 pixels. The mean edge-preservation harmonic-mean score
increased from 0.6240 to 0.6429, corresponding to a 3.02 percent relative improvement, with improvements observed in 18 of 24
samples. The complete refinement introduced approximately 0.93 percent additional runtime, while the structural similarity index
measure changed marginally from 0.6230 to 0.6180. The findings indicate that targeted post-generation structural refinement can
strengthen important boundaries at modest computational cost while retaining the visual character of the pretrained cartoonization
backbone.

Article Details

How to Cite
[1]
Kavali Manisha and A. Soujanya, “Low-Overhead Adaptive Multi-Scale Structure-Preserving Refinement for Pretrained Generative Photo Cartoonization”, Int. J. Comput. Eng. Res. Trends, vol. 13, no. 3, pp. 1–12, Sep. 2026.
Section
Research Articles

References

X. Zhou, Y. Zheng, and J. Yang, “Bridging the metrics gap in image style transfer: A comprehensive survey of models and criteria,” Neurocomputing, vol. 624, p. 129430, 2025.

Y. Zhang, Y. Tian, and J. Hou, “CSAST: Content self-supervised and style contrastive learning for arbitrary style transfer,” Neural Networks, vol. 164, pp. 146–155, 2023.

J. Li, L. Wu, D. Xu, and S. Yao, “Soft multimodal style transfer via optimal transport,” Knowledge-Based Systems, vol. 271, p. 110542, 2023.

M. Liu, S. Lin, H. Zhang, Z. Zha, and B. Wen, “Intrinsic-style distribution matching for arbitrary style transfer,” Knowledge-Based Systems, vol. 296, p. 111898, 2024.

S. Kim, Y. Min, Y. Jung, and S. Kim, “Controllable style transfer via test-time training of implicit neural representation,” Pattern Recognition, vol. 146, p. 109988, 2024.

Y. Zheng, J. Jiao, F. Ye, Y. Zhou, and W. Li, “Fast style transfer for ethnic patterns innovation,” Expert Systems with Applications, vol. 249, p. 123627, 2024.

S. A. Khowaja, L. Nkenyereye, G. Mujtaba, I. H. Lee, G. Fortino, and K. Dev, “FISTNet: Fusion of style-path generative networks for facial style transfer,” Information Fusion, vol. 112, p. 102572, 2024.

Z. Zhang, Q. Zhang, J. Luan, M. Yang, Y. Wang, and L. Zhao, “SPAST: Arbitrary style transfer with style priors via pre-trained large-scale model,” Neural Networks, vol. 189, p. 107556, 2025.

Y. Dong, S. Liu, Y. Li, and L. Zheng, “Aesthetic-aware adversarial learning network for artistic style transfer,” Neurocomputing, vol. 646, p. 130431, 2025.

J. Li, Y. Xiang, H. Wu, S. Yao, and D. Xu, “Optimal transport-based patch matching for image style transfer,” IEEE Transactions on Multimedia, vol. 25, pp. 5927–5940, 2023.

M. Elad and P. Milanfar, “Style transfer via texture synthesis,” IEEE Transactions on Image Processing, vol. 26, no. 5, pp. 2338–2351, 2017.

M.-M. Cheng, X.-C. Liu, J. Wang, S.-P. Lu, Y.-K. Lai, and P. L. Rosin, “Structure-preserving neural style transfer,” IEEE Transactions on Image Processing, vol. 29, pp. 909–920, 2020.

Y. Jing, Y. Yang, Z. Feng, J. Ye, Y. Yu, and M. Song, “Neural style transfer: A review,” IEEE Transactions on Visualization and Computer Graphics, vol. 26, no. 11, pp. 3365–3385, 2020.

H. Ding, G. Fu, Q. Yan, C. Jiang, T. Cao, W. Li, S. Hu, and C. Xiao, “Deep attentive style transfer for images with wavelet decomposition,” Information Sciences, vol. 587, pp. 63–81, 2022.

S. Liu and T. Zhu, “Structure-guided arbitrary style transfer for artistic image and video,” IEEE Transactions on Multimedia, vol. 24, pp. 1299–1312, 2022.

Y. Shih, S. Paris, C. Barnes, W. T. Freeman, and F. Durand, “Style transfer for headshot portraits,” ACM Transactions on Graphics, vol. 33, no. 4, pp. 148:1–148:14, 2014.

Y. Men, Y. Yao, M. Cui, Z. Lian, and X. Xie, “DCT-Net: Domain-calibrated translation for portrait stylization,” ACM Transactions on Graphics, vol. 41, no. 4, pp. 140:1–140:9, 2022.

X. Gao and Y. Zhang, “SRAGAN: Saliency regularized and attended generative adversarial network for Chinese ink-wash painting style transfer,” Pattern Recognition, vol. 162, p. 111344, 2025.

L. M. Gladence, Y.-W. Lai, F.-T. Lee, M.-Y. Chen, and H.-T. Wu, “Generative adversarial network based on CNN classifier predicted scores for image style transfer,” Applied Soft Computing, vol. 178, p. 113303, 2025.

M. Ruder, A. Dosovitskiy, and T. Brox, “Artistic style transfer for videos and spherical images,” International Journal of Computer Vision, vol. 126, no. 11, pp. 1199–1219, 2018.

Z. Zhou, Y. Wu, and Y. Zhou, “Consistent arbitrary style transfer using consistency training and self-attention module,” IEEE Transactions on Neural Networks and Learning Systems, vol. 35, no. 11, pp. 16845–16856, 2024.

X. Kong, Y. Deng, F. Tang, W. Dong, C. Ma, Y. Chen, Z. He, and C. Xu, “Exploring the temporal consistency of arbitrary style transfer: A channelwise perspective,” IEEE Transactions on Neural Networks and Learning Systems, vol. 35, no. 6, pp. 8482–8496, 2024.

D. Chen, L. Yuan, J. Liao, N. Yu, and G. Hua, “Explicit filterbank learning for neural image style transfer and image processing,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 7, pp. 2373–2387, 2021.

J. Gao, Y. Sun, Y. Liu, Y. Tang, Y. Zeng, D. Qi, K. Chen, and C. Zhao, “StyleShot: A snapshot on any style,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 2, pp. 1215–1228, 2026.

K. Tang and C. Wang, “StyleRF-VolVis: Style transfer of neural radiance fields for expressive volume visualization,” IEEE Transactions on Visualization and Computer Graphics, vol. 31, no. 1, pp. 613–623, 2025.

Z. Hu, B. Ge, and C. Xia, “Multiangle feature fusion network for style transfer,” Image and Vision Computing, vol. 154, p. 105386, 2025.

Y. Yu, J. Wang, and N. Li, “Foreground and background separated image style transfer with a single text condition,” Image and Vision Computing, vol. 143, p. 104956, 2024.

Z. Zhou, F. Zhou, and G. Qiu, “Blind image quality assessment based on separate representations and adaptive interaction of content and distortion,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 4, pp. 2484–2497, 2024.

Y. Liu, Z. Ni, S. Wang, H. Wang, and S. Kwong, “High dynamic range image quality assessment based on frequency disparity,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 8, pp. 4435–4440, 2023.

I. Stępień and M. Oszust, “Three-branch neural network for no-reference quality assessment of pan-sharpened images,” Engineering Applications of Artificial Intelligence, vol. 139, p. 109594, 2025.

S. Pang, X. Chen, Y. Xie, H. Zhan, B. Yin, and Y. Lu, “Diff-TST: Diffusion model for one-shot text-image style transfer,” Expert Systems with Applications, vol. 263, p. 125747, 2025.

F. Zhang, B. Feng, Z. Xia, J. Weng, W. Lu, and B. Chen, “Conditional image hiding network based on style transfer,” Information Sciences, vol. 662, p. 120225, 2024.

H. Lee, J. Seol, S. Goo Lee, J. Park, and J. Shim, “Contrastive learning for unsupervised image-to-image translation,” Applied Soft Computing, vol. 151, p. 111170, 2024.

G. Iglesias, E. Talavera, and A. Díaz-Álvarez, “A survey on GANs for computer vision: Recent research, analysis and taxonomy,” Computer Science Review, vol. 48, p. 100553, 2023.

S. Huang, Q. Li, J. Liao, S. Wang, L. Liu, and L. Li, “Controllable image synthesis methods, applications and challenges: A comprehensive survey,” Artificial Intelligence Review, vol. 57, p. 336, 2024.