Infrared small target detection is a critical capability for remote sensing, fire prevention, and surveillance, but existing methods struggle with tiny, low-contrast targets that lack distinct shape or texture. A new two-stage deep learning network developed by researchers at the Harbin Institute of Technology's Research Center for Space Optical Engineering promises to change that, achieving state-of-the-art detection rates across three public datasets while significantly reducing both missed detections and false alarms in complex imaging environments.
The challenge is formidable. Infrared targets often occupy fewer than 81 pixels—typically under 9×9—and exhibit extremely low energy with signal-to-noise ratios around 3. They lack prominent shape or texture information, causing them to be easily submerged in background clutter. While deep learning methods have improved performance, most focus exclusively on target features while neglecting background information, the vast majority of the image, leading to severe class imbalance between positive and negative samples.
On June 30, 2026, the team published (DOI: 10.34133/remotesensing.1046) their findings in the Journal of Remote Sensing. Their proposed Diffusion-Enhanced Dense Mamba Network (DEDM-Net) addresses the critical challenge of detecting infrared small targets that are easily overwhelmed by background clutter. This technology directly impacts forest fire prevention, surveillance early warning systems, and military threat assessment—applications where missed detections or false alarms can have severe consequences.
The two-stage network achieves a synergistic effect greater than the sum of its parts. The first stage employs a dual-path diffusion model with a novel blind processing module that predicts each pixel using only surrounding information—never the pixel itself—preventing extremely small targets from being misclassified as background. The second stage introduces a dense nested Mamba architecture based on the state space model, which captures long-range correlations across global and local features with linear computational complexity—a significant advantage over conventional Transformers. A cross-stage prediction fusion module further integrates features from both stages, improving contour segmentation accuracy.
Evaluated on three public datasets—NUAA-SIRST (427 images), NUDT-SIRST (1,327 images at 256×256), and IRSTD-1k (1,000 images at 512×512)—DEDM-Net outperformed 11 state-of-the-art methods. On NUDT-SIRST, it achieved 93.40% IoU, 93.28% nIoU, 98.37% detection probability, and a remarkably low false-alarm rate of just 3.75×10⁻⁶. On IRSTD-1k, it reached 73.71% IoU and 93.89% detection probability with only 11.10×10⁻⁶ false alarms.
“Infrared small targets are extremely challenging because they lack shape and texture—they're essentially just a few bright pixels in a sea of background,” said corresponding author Dr. Shikai Jiang. “By modeling both the target-free background and potential target regions simultaneously, our diffusion-enhanced approach effectively amplifies what matters while suppressing what doesn't. The Mamba architecture then provides the global context needed to distinguish true targets from bright clutter.”
While DEDM-Net achieves superior accuracy, the diffusion-based two-stage design increases inference time compared to single-stage networks. Future work will focus on model distillation, mixed-precision inference, and faster samplers to reduce the required diffusion steps. The approach holds promise for real-time surveillance systems, autonomous drone navigation in low-visibility conditions, and early wildfire detection networks.
The framework could also inspire new thinking about how generative models and state-space architectures can be combined for other challenging computer vision tasks where target-background separation is critical. For industries relying on remote sensing and surveillance, this research represents a significant step toward more reliable, automated detection systems that can operate effectively in the most demanding environments.


