Articles in this Volume

Research Article Open Access
Research on a Fine-Granularity Text-Image Emotional Recognition Algorithm Optimized Through Cross-Modality Feature Fusion
Article thumbnail
Emotion recognition is key tech for human-computer interaction and mental health check, it got big value in affective computing and smart healthcare. Old single-modal way only use one data source, so noise mess it up easy and it hard to catch full emotional meaning. But multimodal method mix visual, voice and text data together, use how they help each other to make result much better. Still, it have some trouble: feature talk bad between modes, meaning not match well, and too much noise stay around. The LRD method split weights and fused modal features side by side, which cut down model params a lot. Tests on CMU-MOSI and CMU-MOSEI datasets show the new model boost accuracy in multimodal emotion recognition and also keep model simpler.
Show more
Read Article PDF
Cite
Research Article Open Access
Neural Time Series Forecasting: A Problem-Driven Survey from Data Challenges to Model Choice
Surveys of time series forecasting usually proceed by model family, from recurrent networks to Transformers and foundation models. That chronology is useful, but it gives limited guidance when the practical question is why a forecast fails. This article instead organizes the literature around five recurring difficulties: nonlinear and nonstationary behavior, contamination and structural breaks, uncertainty, long contexts and cross-variable dependence, and limited target-domain data. Studies published between 2014–2025 are compared through the assumptions they make, the settings in which they work, and the failure modes they leave unresolved. The review covers recurrent and probabilistic models, decomposition methods, robust and uncertainty-aware forecasting, Transformer variants, theory-guided methods, and foundation models. Quantitative results are used only when the underlying protocol is sufficiently clear. Across these families, the evidence supports a restrained workflow: begin with a strong simple baseline, identify the dominant source of error, and add complexity only when it addresses that source.
Show more
Read Article PDF
Cite
Research Article Open Access
Review of Key Technologies and Application Scenarios of Neural Rendering
Article thumbnail
At prohibitive labor costs and low flexibility, physically based rendering is far from effective for the high-fidelity 3D content creation, thus in this paper, we attempt to give a survey of the area of neural rendering. We explain how deep neural networks learn implicit representation of scenes from visual data for efficient and photorealistic novel view synthesis. We review three common types of neural rendering methods: implicit representation (such as neural radiance fields (NeRF) and its subsequent architectures such as Instant-NGP), explicit representation (such as neural texture), and hybrids (such as NeuS and 3D Gaussian splatting). We explain their essence, linkages, advantages and disadvantages. For example, 3D Gaussian splatting achieves state-of-the-art results on the celebrated Blender dataset (33.55 dB PSNR, 0.967 SSIM) with real-time rendering 134 FPS, solving the efficiency limit of vanilla NeRF.Looking ahead, the combination of neural rendering with generative AI, large-scale 3D foundation models and lightweight architectures will lay a foundation for the emerging immersive internet, and digital twin landscape.
Show more
Read Article PDF
Cite
Research Article Open Access
AI for Preterm Brain Injury Imaging: From Automated Analysis to Clinical Translation
Preterm brain injury is a major cause of lifelong motor, cognitive, and behavioraldisability, yet its imaging phenotype changes rapidly with development and differs across cranialultrasound (cUS) and magnetic resonance imaging (MRI). Serial cUS is suited to bedsidesurveillance of hemorrhage and ventricular enlargement, whereas term-equivalent structural anddiffusion MRI provide more detailed assessment of tissue injury, maturation, and white matterorganization. Artificial intelligence (AI) can support image-quality assessment, lesion detection,segmentation, quantitative phenotyping, developmental-age estimation, and outcome prediction,but evidence for clinical translation remains uneven. This focused review maps imaging modalitiesand computational tasks to neonatal decisions and synthesizes studies published from 2005through 2025. Structural MRI segmentation has the most consistent technical evidence; cUSautomation, subtle-lesion detection, longitudinal multimodal prediction, and prospective clinicaluse remain less mature. Dataset scale alone is not sufficient because developmental stage,acquisition differences, nonrandom missing examinations, and delayed follow-up can influenceapparent performance. We therefore distinguish model accuracy from deployability and proposea minimum pipeline comprising input governance, infant-level external evaluation, calibrateduncertainty, safety-based abstention, human review, and post-deployment monitoring. Futureprogress depends on longitudinal multicenter cohorts, development-aware representation learning,and prospective studies that measure reporting efficiency, management changes, and safety. AIshould be judged by reproducible clinical benefit rather than internal benchmark accuracy.
Show more
Read Article PDF
Cite
Research Article Open Access
AI Billiard Coach: Real-Time Shot Assistance System
This study addresses the democratization of professional billiards training through the use of artificial intelligence and camera equipment. The study explores three core issues: (1) how to accurately detect the layout and specific positions of the balls and the table in varying lighting conditions from photos; (2) how to use generative AI to generate physics-based shooting suggestions; and (3) how to optimize the user interface for beginner accessibility. In terms of methodology, the system integrates computer vision algorithms based on OpenCV and constrained GPT-4 prompts, and combines the principles of classical collision mechanics. Experiments were conducted on 120 images submitted by users, and the accuracy of ball detection under medium lighting (300 lux) was 89.2%, while the consistency of the AI-generated suggestions with professional benchmarks reached 76.3%. The main limitations include angular errors caused by parallax (±4.7°) and color recognition failure of dark balls under warm lighting. The study concludes that AI training based on smartphones can effectively reduce the professional barriers of billiards, and future work will prioritize real-time video analysis and mobile deployment.
Show more
Read Article PDF
Cite
Research Article Open Access
Research and Analysis of Cross-Modal Attention Based on Pareto Optimal Fusion of Sketch Image Generation
Article thumbnail
Conditional Sketch Generation aims to convert abstract sketches into realistic images. This task is challenging due to the modality gap between sketches and photos, weak pixel-level correspondence, and the one-to-many mapping problem. This paper introduce two paradigm shifts within a single framework, replacing ControlNet's zero-convolution with deep cross-modal attention and noise prediction with L1 reconstruction loss. On this basis, a Dual-Path Non-Shared Encoder is presented. During the experiments on SketchCOCO, this paper uncovers three empirical phenomena with universal significance. Firstly, all evaluation metrics exhibit a synchronous improvement phase when w < 0.7 . Secondly, the guidance weight exhibits a stable Pareto interval, with model performance sustaining high stability for w ∈ [ 0.6,0.9 ] . Finally, through dense sampling with a step size of 0.1, w = 0.7 achieves the best trade-off across all metrics (FID = 177.90, Corr = 0.7160, SSIM = 0.6419). Meanwhile w = 0.9 yields the lowest LPIPS (0.1949) and w = 0.8 achieves the highest SSIM (0.6471). This work not only provides a lightweight yet powerful model for sketch generation, but also offers actionable insights for designing fusion mechanisms in cross-modal tasks.
Show more
Read Article PDF
Cite
Research Article Open Access
Semantic Response Mechanisms of L2 Writing Revision Behavior under Generative AI Feedback
Article thumbnail
Generative AI is changing the way L2 writing feedback is produced and how learners revise their texts. Writing feedback is gradually moving from one way correction to human AI semantic interaction. Based on this change, the study focuses on non English major undergraduates in Chinese universities who have passed CET-4. It builds an experimental process that combines initial drafts, ChatGPT feedback, revised drafts, revision logs, and teacher ratings. Inputlog or Google Docs is used to record the revision process. SBERT is used to calculate semantic distance between initial drafts and revised drafts. Coh Metrix indicators, feedback uptake annotation, and mixed effects models are then used to examine how different feedback types affect writing quality improvement. The results show that grammar and vocabulary feedback are more easily accepted, but their explanatory power for score improvement is limited. Cohesion and argument feedback have lower uptake rates, yet they contribute more clearly to task response, paragraph organization, and coherence. The relationship between semantic distance and writing improvement is nonlinear. Moderate semantic change is most likely to produce effective revision. The findings indicate that the value of generative AI feedback does not depend on revision quantity. It depends on whether learners can make high quality semantic responses while preserving their original writing intentions.
Show more
Read Article PDF
Cite
Research Article Open Access
Machine Learning Identification of Capital Structure Resilience in New Energy Firms under Energy Market Volatility Shocks
Article thumbnail
Global energy markets continue to fluctuate under geopolitical conflicts, supply demand mismatch, and the pressure of low carbon transition. These conditions place stronger external pressure on the financing stability and capital structure adjustment capacity of new energy firms. Photovoltaic and lithium battery firms usually have high R&D expenditure, long capacity expansion cycles, and strong demand for debt financing. High energy price volatility may affect their capital structure resilience through market expectations, financing costs, and cash flow pressure. Unbalanced panel data are constructed from A share listed photovoltaic and lithium battery firms from 2014 to 2024. High volatility energy price shocks are defined when the quarterly return volatility of Brent, WTI, or natural gas prices exceeds the 75th percentile of the sample period. Capital structure resilience labels are then constructed according to post shock leverage, interest coverage ratio, and cash to short term debt ratio. Logistic regression, random forest, XGBoost, and LightGBM are used for classification. SHAP is used to explain key feature contributions. The results show that XGBoost performs well in AUC, F1 score, and recall of low resilience firms. Short term debt ratio, cash to short term debt ratio, operating cash flow volatility, and Brent price volatility are important variables for distinguishing high resilience firms from low resilience firms. The analysis provides an explainable quantitative basis for financial risk identification, debt structure optimization, and investment screening under energy market shocks.
Show more
Read Article PDF
Cite
Research Article Open Access
Model Choice or Decision Threshold? A Decision-Aware Evaluation Protocol for Employee Attrition Prediction
Article thumbnail
Research on employee attrition prediction has grown quickly, yet most studies leave two questions open: whether the metric difference used to rank models is statistically reliable on the data concerned, and whether, once a business cost is attached to the retention decision, it is the choice of model or of decision threshold that determines what an organisation pays. This study proposes a four-step, decision-aware evaluation protocol addressing both and applies it to four datasets spanning 295 to 14,999 records, three public and the fourth from a real organisation. On the two smaller benchmarks the models cannot be separated, and over ten independent splits the leading model on the IBM set changes hands four times, so any single split's winner is an accident of partitioning. Where models are tied only the threshold reduces cost, and its optimum lies well below the conventional 0.5; where they are separated by a wide margin the model determines cost instead. The real dataset bounds that rule: there the models are separable, yet the cost parameters place the optimal policy so close to intervening on everyone that no model can improve on it. A sensitivity analysis over the cost parameters leaves these conclusions unchanged. The contribution is a reusable protocol rather than a new algorithm, with evidence that the dominant cost lever depends on the data and on whether the cost-optimal policy discriminates at all.
Show more
Read Article PDF
Cite
Research Article Open Access
Non-ideal Data Problems in Cooperative Spectrum Sensing
Article thumbnail
In recent years, as the development of wireless communication technologies and the increase in smart devices rapidly, the shortage of spectrum resources has become the important issue for future wireless communication systems. Cognitive Radio (CR) enabled spectrum utilization through Dynamic Spectrum Access (DSA) which allowing unlicensed users to access licensed spectrum without interference. Cooperative Spectrum Sensing (CSS) which is an important part of cognitive radio systems to improve the sensing reliability by allowing multiple nodes shared sensing results and made a joint decision. In practical scenarios, CSS systems always faced several non-ideal data issues which include noise uncertainty, data missing, transmission errors, hardware impairments, and synchronization problems. These problems will distort sensing statistics, to reduce the detection probability, increase false alarm probability, and reduce the benefit of gain from cooperative. Therefore, this paper investigates how non-ideal data affects CSS systems and reviews several representative methods, including adaptive threshold detection, robust detection methods, data recovery techniques, synchronization correction, and uncertainty-aware fusion. The comparison results indicate that robust and intelligent cooperative methods can provide better sensing performance than traditional energy detection method under practical non-ideal conditions, especially in terms of detection probability, false alarm control, and overall system robustness. In addition, this paper discusses potential future research directions for intelligent cooperative sensing and explores the emerging technology such as potential influence of 6G networks and IoT, which may influence the development of CSS system.
Show more
Read Article PDF
Cite