Current Issue : October-December Volume : 2026 Issue Number : 4 Articles : 5 Articles
Single image super-resolution (SR) is an important part of image processing, which aims to improve the spatial resolution of images. This is a typical ill-posed inverse problem. The main difficulty is that a low-resolution image block usually corresponds to multiple high-resolution image blocks. The existing methods cannot provide enough correlation to determine the unique high-resolution image block, which leads to artifacts and image distortion in the reconstructed image. To address this problem, a method (EHNet) is proposed to achieve super-resolution by using a Hybrid-Channel Fusion Block (HCFB) and an Enhanced Dual-Convolution Block (EDCB). The EDCB effectively enhances the network’s ability to capture image details and textures by combining local and global feature processing. The HCFB strengthens the information interaction between channels by combining channel segmentation with large-kernel convolution, fully explores feature dependencies, and thus optimizes the feature extraction effect. Experimental results show that the superresolution reconstructed image of EHNet achieves 32.59 dB PSNR and 0.9006 SSIM on the Set5 ×4 SR benchmark, outperforming several state-of-the-art SR methods. In addition, the model exhibits notable improvements in artifact suppression, and the reconstructed image’s subjective visual impact surpasses that of other current techniques....
Object detection (OD) technology, which identifies and classifies objects in images and videos, has been widely adopted across various fields. However, implementing OD faces challenges, including image preprocessing, labeling, model development, and deployment. To streamline these processes, we developed a Python-based software Ladder (Labeling and Detection Deployment for Entity Recognition). Ladder features a user-friendly graphic interface (GUI) that facilitates efficient labeling of training datasets, detection of new images, and model training. The software utilizes an interactive recurrent framework that begins with predictions from a pre-trained model for initial image labeling. Users can then add human labels, and these newly labeled images can be incorporated into the training data to retrain the model. In this study, we demonstrate an efficient development of a broken rice detection model using Ladder. The model employed a three-stage training process and demonstrated strong predictive performance (R2 = 0.99), with a mean absolute error (MAE) of 6.08 (95% CI: 5.18–6.97) and a root mean square error (RMSE) of 6.68 (95% CI: 5.93–7.46). Rice is one of the world’s most essential crops, with the rate of broken rice significantly affecting its price in the market and potential uses. This necessitates an efficient method for assessing the ratio of broken rice for breeding, production, and trading....
Multifunctional imaging that enables the extraction of multiple physical parameters from a single-shot image is highly desirable for compact and integrated optical systems. Here, we present a metalens-enabled double-helix point spread function (DH-PSF) for multifunctional imaging, capable of sensing both depth and wavelength information. By superposing Bessel beams with different transverse wave vectors, we design a DH-PSF metalens with two controllable degrees of freedom, whose rotation angle varies linearly with the axial position and the main lobe spacing is inversely proportional to the axial position, significantly enhancing the robustness of depth estimation. Furthermore, the dispersive property of the DH-PSF is deliberately exploited for wavelength sensing, as the lobe spacing exhibits a linear dependence on the incident wavelength at a fixed depth. We experimentally demonstrate single-shot three-dimensional (3D) imaging and wavelength discrimination, with estimation errors maintained below 5%. This work provides new insights into the construction and engineering of DH-PSFs and establishes a compact and versatile framework for 3D imaging and spectral sensing, offering new opportunities for multifunctional imaging devices in machine vision and microscopy....
The processing of two-dimensional (2D) spectral images constitutes a critical and multifaceted discipline in contemporary astronomical data analysis. As spectroscopic instruments evolve towards higher multiplexing, resolution, and sensitivity, the raw 2D data captured by detectors present increasingly complex challenges that transcend simple onedimensional extraction. This review provides a systematic and comprehensive examination of the methodological evolution in this field over the past two decades. It gathered relevant studies by searching mainstream academic repositories and general search engines with the core keyword ‘2D Spectral Image’, and selected qualified references according to accessibility and research relevance. We categorize the landscape into three major paradigms: (1) physics-based modeling and algorithmic correction techniques for geometric distortion, scattered light, and sky background; (2) data-driven machine learning and deep learning approaches for image correction, spectral classification, and faint signal detection; and (3) the development of open-source software pipelines that democratize advanced processing. A central contribution of this review is a detailed comparative analysis of the performance metrics, underlying assumptions, and practical limitations of prominent algorithms. We highlight the transformative impact of convolutional neural networks (CNNs) and vision transformers (ViTs) on tasks such as celestial object classification and exoplanet detection, while also acknowledging the enduring importance of robust physical models for calibration and uncertainty quantification. The discussion culminates in an assessment of persistent challenges—including computational scalability, model generalizability, and interpretability—and outlines promising future directions at the intersection of AI, statistical inference, and large-scale survey science....
Extracting spatio‐temporal cues from neighbouring frames is challenging in video super‐resolution (VSR). Although deformable alignment‐based VSR methods have shown promise in aligning neighbouring frames with the reference frame, most existing methods rely on one or a few traditional convolutions to estimate motion offsets for spatio‐temporal alignment, restricting receptive field size and alignment accuracy. To address these limitations, we propose an effective spatio‐temporal alignment network (ESTA‐Net) for VSR. The core component of our method is the group convolution‐based alignment module (GCBAM), which utilises cascaded group convolutions to learn offsets across both the original and downsampled resolutions. By employing group convolutions rather than traditional convolutions, GCBAM enables the deformable alignment to achieve a wider receptive field with lower computational cost, thereby improving the accuracy of offset estimation. Additionally, the bi‐scale alignment strategy within GCBAM enhances robustness to complex and large‐scale motions. Furthermore, we introduce an attentionbased feature enhancement module (AFEM) to refine the aligned features, focusing on critical details to improve reconstruction quality. Extensive experiments on standard benchmarks show that our ESTA‐Net achieves superior VSR performance against other advanced methods, while maintaining a good equilibrium between model size and performance....
Loading....