Articles in this Volume

Research Article Open Access
AI-Enhanced Noise Interference Mitigation and Signal Loss Recovery in Communication Transmission: A Case-Based Study
With the rapid development of information technology, communication systems are increasingly required to support stable, reliable, and real-time information transmission in complex environments. However, during practical communication transmission, signals are often affected by noise interference, channel fading, network jitter, multi-user interference, and packet loss. These problems may reduce speech intelligibility, increase bit error rates, cause video freezing, and weaken the continuity and reliability of communication services. Traditional enhancement techniques rely on stationary assumptions or fixed redundancy, and thus perform poorly under non-stationary noise or dynamically fluctuating network conditions. In recent years, deep learning-based methods have shown strong potential in learning nonlinear mappings from degraded to clean signals, offering better adaptability to diverse distortions. Nevertheless, the analysis shows that AI technology still faces practical deployment challenges, including high computational complexity, heavy dependence on large-scale labeled training data, and limited cross-scenario generalization capability, which require further optimization before such approaches can be reliably integrated into real-time audio/video communication systems.
Show more
Read Article PDF
Cite
Research Article Open Access
From Transformers to World Simulators: A Survey on Generative AI Video Production Technology
The rise of generative artificial intelligence has triggered a paradigm shift in the field of video production. This survey systematically examines the evolutionary progress of generative AI video production technologies, with a focus on analyzing the transition from CNN architectures to Transformer-dominated frameworks. Adopting a systematic literature review methodology, this study analyzes relevant academic papers and technical reports. Centered on three core research questions, this survey explores: (1) What architectural innovations have enabled the shift from CNN-based to Transformer-based video generation? (2) What are the current capabilities and limitations of cutting-edge models such as Sora? (3) What challenges and future directions lie ahead for this rapidly evolving field? Key findings demonstrate that the self-attention mechanism of Transformers fundamentally addresses the inherent long-range temporal dependency problem of CNNs, reducing temporal consistency error from approximately 42% for CNNs to roughly 15% for Transformers. Nevertheless, substantial challenges persist in multi-modal consistency, computational resource requirements (approximately 10¹⁸ FLOPs for generating a single 10-second video), and ethical concerns. This survey identifies world models and interactive generation as promising research frontiers for the domain.
Show more
Read Article PDF
Cite
Research Article Open Access
Adapting Pure-Text Large Language Models for Remote Sensing Classification via Lightweight Visual Adapter
Article thumbnail
Stable cross-domain feature alignment is indispensable for earth observation classification, which is fundamentally hampered by radiometric gaps between generic pre-training images and aerial remote sensing data. Vision-language pre-trained models exhibit strong zero-shot capability on ordinary photos yet suffer severe accuracy loss on satellite and aerial imagery. While LoRA tuning cuts partial training costs, backbone parameter fine-tuning still brings considerable GPU memory overhead in training. Relying on frozen DeepSeek V4 MoE text LLM and static SigLIP vision encoder, this study designs a slim cross-modal projection subnet to eliminate feature distribution gaps between modalities. Stacked residual MLPs constitute the sole learnable part, containing roughly 20M parameters for visual-text latent space matching. The model is evaluated collectively on EuroSAT, PatternNet and RSSCN7, covering nearly 60,000 aerial images with 55 separate scene classes. Recorded aggregate classification precision reached 99.80% across the unified multi-source testing pool. Compared with LoRA-dependent VL-ZSDA-RS benchmark schemes, the adjustable parameter scale shrinks by over half, alongside a 46% cut in peak GPU memory usage. Layer-wise ablation trials reflect unstable matching performance under shallow projection layouts; five stacked transformation layers deliver the most balanced tradeoff between computation overhead and inter-modal alignment quality. External trainable mapping subnetworks, as the test records suggest, unlock visual discrimination capacity for unmodified text-only large language models, supplying a low-hardware threshold tuning route for earth observation research groups constrained by computing resources.
Show more
Read Article PDF
Cite