# LongBench v2 / 66ee3c6d821e116aacb20f6f

task_id: 063060d1-c3b0-592c-a263-5380965c5233
task_key: train--66ee3c6d821e116aacb20f6f
task_revision_id: 1

{"choice_A":"SRDTrans focus on handling both spatial redundancy, while the self-inspired model focuses exclusively on time-lapse temporal data.","choice_B":"The SRDTrans model lacks dynamic training data, limiting its performance in low frame rate situations.","choice_C":"SRDTrans utilizes a temporal sampling technique that is more computationally efficient.","choice_D":"The self-inspired learning model requires fewer network parameters and can therefore generalize better than SRDTrans.","context":"Nature Methods\nnature methods\nhttps://doi.org/10.1038/s41592-024-02400-9\nArticle\nSelf-inspired learning for denoising live-cell \nsuper-resolution microscopy\nLiying Qu1,14, Shiqun Zhao2,14, Yuanyuan Huang1,14, Xianxin Ye2, Kunhao Wang2, \nYuzhen Liu1, Xianming Liu \n \n 3, Heng Mao4, Guangwei Hu \n \n 5, Wei Chen \n \n 6, \nChangliang Guo \n \n 2, Jiaye He \n \n 7,8, Jiubin Tan9, Haoyu Li \n \n 1,9,10,11, \nLiangyi Chen \n \n 2,12,13 & Weisong Zhao \n \n 1,9,10,11 \nEvery collected photon is precious in live-cell super-resolution (SR) \nmicroscopy. Here, we describe a data-efficient, deep learning-based \ndenoising solution to improve diverse SR imaging modalities. The method, \nSN2N, is a Self-inspired Noise2Noise module with self-supervised data \ngeneration and self-constrained learning process. SN2N is fully competitive \nwith supervised learning methods and circumvents the need for large \ntraining set and clean ground truth, requiring only a single noisy frame for \ntraining. We show that SN2N improves photon efficiency by one-to-two \norders of magnitude and is compatible with multiple imaging modalities for \nvolumetric, multicolor, time-lapse SR microscopy. We further integrated \nSN2N into different SR reconstruction algorithms to effectively mitigate \nimage artifacts. We anticipate SN2N will enable improved live-SR imaging \nand inspire further advances.\nFluorescence microscopy of live cells requires gentle imaging condi-\ntions and adequate spatiotemporal resolution to record authentic \nbiological information, thus the photon budget is usually limited. By \nencoding SR information via specific optics and fluorescent on–off \nindicators, the SR techniques have further strengthened the spatial \nresolution1 and enabled previously unappreciated, intricate structures \nto be observed2–5. However, for a finite number of fluorophores within \nthe cell volume, the increase in spatial resolution leads to the rise in \nillumination intensity or exposure time by orders of magnitude to \nmaintain the signal-to-noise ratio (SNR)6. Furthermore, any increase \nin spatial resolution must be matched with an increase in temporal \nresolution to prevent motion artifacts7,8. This is particularly challenging \nfor live-cell SR imaging to accumulate sufficient photons.\nBecause of the hardware-limited photon budget, computa-\ntionally boosting the SNR is essential to maximize the utilization of \nsensor-collected photons. By modeling the image formation process, \nclassical denoising algorithms based on numerical filtering and math-\nematical optimization can remove the noise from fluorescence images \nto a certain extent9–12. However, the corresponding carried assumptions \nare not dependent on the specific content of the imaging data and \ncannot fulfill the optimal performance. Therefore, to capture the full \nstatistical complexity of data, the field has witnessed a sudden surge \nReceived: 22 January 2024\nAccepted: 31 July 2024\nPublished online: xx xx xxxx\n Check for updates\n1Innovation Photonics and Imaging Center, School of Instrumentation Science and Engineering, Harbin Institute of Technology, Harbin, China. 2State Key \nLaboratory of Membrane Biology, Beijing Key Laboratory of Cardiometabolic Molecular Medicine, Institute of Molecular Medicine, National Biomedical \nImaging Center, School of Future Technology, Peking University, Beijing, China. 3School of Computer Science and Technology, Harbin Institute of \nTechnology, Harbin, China. 4School of Mathematical Sciences, Peking University, Beijing, China. 5School of Electrical and Electronic Engineering, \nNanyang Technological University, Singapore, Singapore. 6School of Mechanical Science and Engineering, Advanced Biomedical Imaging Facility, \nHuazhong University of Science and Technology, Wuhan, China. 7National Innovation Center for Advanced Medical Devices, Shenzhen, China. 8Shenzhen \nInstitute of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China. 9Key Laboratory of Ultra-precision Intelligent Instrumentation of \nMinistry of Industry and Information Technology, Harbin Institute of Technology, Harbin, China. 10Frontiers Science Center for Matter Behave in Space \nEnvironment, Harbin Institute of Technology, Harbin, China. 11Key Laboratory of Micro-Systems and Micro-Structures Manufacturing of Ministry of \nEducation, Harbin Institute of Technology, Harbin, China. 12PKU-IDG/McGovern Institute for Brain Research, Beijing, China. 13Beijing Academy of Artificial \nIntelligence, Beijing, China. 14These authors contributed equally: Liying Qu, Shiqun Zhao, Yuanyuan Huang. \n e-mail: weisongzhao@hit.edu.cn\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\neach pixel of the camera is sampled independently; (p4) The imaging \nsystem can be regarded as an equivalent low-pass filter. In this case, \neach SR image is composed of many PSFs multiplied with the corre-\nsponding fluorophore brightness at different spatial positions, and \neach PSF occupies at least larger than an area of 2 × 2 pixels12. Consider-\ning this spatial redundancy (p1 and p2), we can directly resample one \nSR image as two subimages by an unusual form of binning we regard as \n‘diagonal resampling’ (Fig. 1a). More specifically, we consider every four \nadjacent pixels (2 × 2) as one unit, and every two diagonal pixels (top \nleft and bottom right, magenta color in Fig. 1a; top right and bottom \nleft, green color in Fig. 1a and Extended Data Fig. 1a) are averaged as one \nnew pixel. Because of the independence of each pixel and the spatial \nsymmetry on the two diagonal directions (p2 and p3), we can gener-\nate statistically independent image pairs that share identical details \nbut different noise realizations for the N2N configuration (Methods). \nFundamentally, assuming identical contents inside one smallest unit, \nour method can be seen as a relaxation form of the temporal resam-\npling approach22–24, which requires the entire contents of two frames \nbeing identical. Compared to other spatial generation methods33–35, \nthis resampling in diagonal axes creates the most content-similar \nimage pair (Supplementary Fig. 2a), providing a more stable process \nwithout the need for additional registration or calibration to create \nperfectly matched image content. Notably, although originated from \nthe spatial redundancy of SR images, our approach is not sensitive to \nchanges in pixel size and can produce stable denoising results even for \nundersampled images (until 260 nm pixel size with 150 nm resolution; \nExtended Data Fig. 2a,b).\nBecause the produced two subimages are two times smaller, we \nfurther adapt an interpolation method to rescale them to the original \nstructural scale (Methods). Different from the conventional spatial \nmethods, we apply a Fourier interpolation to the resulting two sub-\nimages, which is based on the fact that the optical transfer function \n(OTF) of an SR microscope has only finite support (p4). Padding the \nFourier-transformed image out of its OTF support with zeros does not \nalter its information content36, and after back-transforming, this pad-\nded image will have doubled pixel numbers in both x and y axes, identi-\ncal to the original SR image (Fig. 1a and Extended Data Fig. 1a). Without \nthis operation, the network may produce structural artifacts for the \nscale difference between the training and test datasets (Extended Data \nFig. 2e). On the other hand, the spatial methods (such as bilinear inter-\npolation) create more pixels according to the noise-corrupted pixels, \nand it will be problematic when meeting the background areas without \nfluorescence signal (Supplementary Fig. 2b), because these smoothly \ncreated pixels do not conform to the randomness of noise, which may \ninfluence the learning process, potentially producing background \nartifacts (Extended Data Fig. 2e).\nSelf-constrained learning process. Beyond the full form of N2N, we \nfurther design a self-constrained learning process to constrain the \ntraining process and denoising variance (Fig. 1a). This is based on a \nsimple intuition that the denoising predictions of the same underly-\ning signal should be identical. In our case, the generated data pair has \nmatched content but different noise components, thus ideally, the \ncorresponding two denoised results should have no differences. Spe-\ncifically, the two generated images are successively and individually fed \ninto the network, and the two resulting predictions can be calculated \nby the loss function shown in equation (1) to execute the training stage.\n1 = ‖ ̃\nx1 −x2‖1 + ‖ ̃\nx2 −x1‖1 + λ‖ ̃\nx1 −\ñ\nx2‖1,\n(1)\nwhere ̃\nx represents the network outcome of input x, and || ||1 refers to \nthe l1 norm. We used the l1 norm as the distance calculation (Methods). \nThe first two terms are the conventional N2N losses, and the last term \nis our self-constrained loss with a constraint weight λ. By enforcing this \nconstraint, we found that the low-frequency components manifest a \nof data-driven methods, producing the unprecedented restoration \nresults13. After iterative training on a dataset with ground truth (GT) \nlabels, deep neural networks (DNNs) can learn the mapping between \nnoisy images and their clean counterparts14–16. Intuitively, for super-\nvised learning, the collection of abundant content-matched clean \nimages is crucial, which is of great challenge to live-cell applications, \nespecially under the SR scale.\nTo denoise images without clean ones, Noise2Noise (N2N)17 learns \na mapping between pairs of independently degraded versions of the \nsame image, and its performance can approach that of supervised \nlearning methods. Nevertheless, the need for the twin noisy pairs is \nstill against the live-cell SR applications. By leveraging the pixel-wise \nindependence of noise, several approximated forms of N2N have \nbeen developed to denoising without paired data18–21. However, these \napproximations may lead to downgraded performances. To realize the \nfull form of N2N configuration, the DeepCAD22 used a temporally inter-\nleaved self-supervised data generation process to create the required \nnoisy data pairs, by assuming that the two adjacent frames in a video of \ncontinuous imaging can be considered with the same underlying con-\ntent. Unfortunately, although this assumption is routinely satisfied in \ncalcium imaging22,23 or other fast-imaging applications24, it is still hard \nto accomplish in live-cell SR techniques considering the commonly \ncompromised temporal resolution and increased spatial resolution.\nOn the other hand, SR imaging usually has a more-than-sufficient \nsampling rate, at least over the Nyquist sampling theorem to protect \nthe enriched spatial information. Thus, we create a self-supervised data \ngeneration strategy based on this spatial redundancy, using a diagonal \nresampling step followed by a Fourier interpolation. Compared to \nthe previous resampling methods22–24, this spatially interleaved data \ngeneration is more universal and stable, and it produces almost no bias \nin our tests. Beyond this spatially self-supervised N2N realization, we \nalso develop a self-constrained learning process to further enhance \nthe denoising performance and data efficiency. Conceptually, the two \nnetwork predictions of the generated noisy pair will be constrained \nto one identical expectation, and this process shrinks the learning \nand predictive uncertainty, inherently increasing the efficiency and \neffectiveness. Together, our self-inspired N2N (SN2N) reaches or even \noutperforms supervised learning-based denoising methods, especially \nwhen only a single frame is used for training.\nTo showcase the broad applicability, we apply SN2N directly to \ndata obtained on two commercial spinning-disk confocal-based struc-\ntured illumination microscopes (SD-SIM)25,26, two commercial stimu-\nlated emission depletion (STED)27,28 microscopes and a high-resolution \nconfocal microscope. Here, we show the method enables high-quality, \nlong-term, multicolor three-dimensional (3D) live-cell SR imaging \n(five-dimensional (5D) in xyz color time). We also show benefits for \nexpansion microscopy (ExM)29. Beyond that, we also integrate our \nSN2N framework into the existing SR reconstruction procedures, \nincluding iterative deconvolution on SD-SIM and STED, SR optical \nfluctuation imaging (SOFI) reconstruction30,31 and SIM10, effectively \nmitigating the artifacts embedded in the SR images and enhancing the \nphoton efficiencies. These organic integrations deliver the fact that \nour SN2N can serve as a practical framework for further strengthening \nthe performance of live-cell SR microscopy. With fully open-sourced \ncode, we also expect our SN2N will be applicable to other fields for \nrandom noise removal.\nResults\nPrinciple of SN2N\nSelf-supervised data generation. Before designing the data gen-\neration strategy, we intend to list several physical properties (p) of SR \nmicroscopy32: (p1) To fulfill the increased spatial resolution, the sample \nrate is routinely finer than the Nyquist theorem. (p2) The point spread \nfunction (PSF) represents the highest frequency of the system and \n \nis usually with spatial symmetry inside 2 × 2 pixels; (p3) Commonly, \n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\ngreater propensity of being stable than the high-frequency ones \n(Extended Data Fig. 2g,h). Thus, to avoid over-smoothed results, we \nroutinely set the constraint weight as 1. Notably, our SN2N is model \nindependent, thus we used the commonly used U-Net37 to highlight its \nability (Fig. 1a, Extended Data Fig. 1a and Methods).\nData augmentation for low-level tasks. Although there are plenty of \nstrategies to increase the training data amount, they are designed for \nhigh-level classification tasks and not suitable for low-level denoising \napplications38. Accordingly, we develop random patch transformations \nin multiple dimensions (Patch2Patch) to further improve the data \nefficiency (Extended Data Fig. 1b). Specifically, the randomly selected \nregions of interest on each image will be rotated or flipped, or will \nremain still and subsequently interchanged with other regions of inter-\nest from the same frame or a different frame from different times or \nexperiments (Methods and Supplementary Video 1). This Patch2Patch \ncan create more imaging results without changing the inherent noise \nproperties and hence it can effectively reduce the required data bulk.\nBenchmarking with known structures\nSimulation validations. To quantitatively test SN2N’s performance, we \nfirst validated it on synthetic microtubule imaging data with 150 nm \nresolution and three different Poisson and Gaussian noise levels \n(Methods and Supplementary Fig. 1). Our SN2N solution is superior \n1/5 data\nFull data\n1/50 data\nxy2\nStep 3. Self-constrained learning\nSelf-inspired process\na\nd\nb\nRaw\nGT\nPURE\nACsN\nN2V\nSupervised\nSN2N (w/o c)\nSN2N\nc\nSTD: 48.32\nSTD: 16.41\nSTD: 12.40\nSTD: 15.97\nSTD: 5.12\nSTD: 11.87\nSTD: 5.42\ne\nf\nSteps 1–2. Self-supervised data generation\nSelf-constraining\nSelf-supervising\nSelf-supervising\nStep 1. Diagonal resampling\nStep 2. Fourier rescaling\n=\n=\n1\n2\n3\n4\n+\n2 + 3\nNet (xy1)\nNet (xy2)\niFFT\nxy1\nxy2\nFFT\nPadding\n1 + 4\nConvolution\nSkip connection\nUpsampling\nDownsampling\nxy1\nCorrelation\n0.2\n0.4\n0.6\n0.8\n1.0\n0\nSpatial frequency  (µm–1)\n0\n2.5\n5.0\n7.5\n10.0\n10\n15\n20\n25\nPSNR\nSN2N\nk = 1.2\nSN2N\n(w/o c)\nk = 2.9\nk = 0\nRaw\nN2V\nk = 2.7\nSupervised\nk = 4.3\nSN2N\n(w a)\nk = 0.9\nRaw\nPURE\nACsN\nN2V\nSN2N\nSN2N (w/o c)\nSupervised\n1.0\n0.8\n0.6\n0.4\n0.2\n0\nSSIM\nLevel 1\nLevel 2\nLevel 3\n1/7\nFig. 1 | Workflow and simulation validation of SN2N. a, Overview of SN2N. \nSteps 1–2, self-supervised data generation. The single-frame input image of \nH × W pixels is diagonally resampled to two images of H/2 × W/2 pixels. Then, \nthe two resulting images are rescaled back to two images of H × W pixels with a \nFourier interpolation. Step 3, self-constrained learning process, that is, the two \nrescaled images serve as both inputs and labels, and the corresponding two \npredicted images from a classical U-Net architecture are constrained to minimize \nthe difference between them (Methods). b, Validation of SN2N using synthetic \nmicrotubule structures (Methods). The synthetic structures were convoluted \nwith a 150 nm PSF and downsampled two times (pixel size, 32.5 nm) as GT. The \nnoisy images were created by further injection of Poisson noise and Gaussian \nnoise. From left to right, noisy (top) and GT images (bottom), PURE, ACsN, N2V, \nsupervised, SN2N without self-constrained loss (SN2N w/o c) and SN2N denoising \nresults. c, Data uncertainty results of b. Ten independent frames of identical \ncontents under the same imaging conditions were fed to the trained SN2N \nnetwork, and the STD of the ten resulting predictions was calculated as the data \nuncertainty. Marked numbers are the average values of STD maps. d, FRC analysis \nof the denoising images. e, SSIM values of various denoising methods under \ndifferent noise levels (n = 10). f, PSNR values of networks trained by different \namounts of data under the level-1 condition (n = 10). Full data dimensions are \n2,048 × 2,048 × 50. k denotes the slope (red lines) of PSNR values along the data \nincrement. Error bars indicate the s.e.m. Experiments were repeated ten times \nindependently with similar results. Scale bar, 1 µm (b).\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nto existing methods in two aspects. First, SN2N is denoising effectively, \nespecially for ultralow-SNR conditions. Under the low-SNR condition \n(level 1), classical denoising methods, for example, Poisson unbiased \nrisk estimate (PURE)9 and automatic correction of sCMOS-related \nnoise (ACsN)11, failed to eliminate the noise and left overly blurred \nunderlying structures (Fig. 1b). The approximated form of N2N, \n \nNoise2Void (N2V)19, can reduce the noise, but the predicted images \nfrom ten repetitive acquisitions are still with large standard deviation \n(STD) values (for example, 15.97 in Fig. 1b,c), reflecting the limited \nperformances. Removing the self-constrained learning process (SN2N \nw/o c), the improvement of our method against N2V is relatively limited \nand cannot compete with the supervised learning method using suf-\nficient training data. The full SN2N truly approaches the performance \nof supervised learning with similar STD and even better Fourier ring \ne\nSTD\n20\n15\n10\n5\n0\nSN2N\nSN2N\n(w/o c)\nN2V\nPURE\nSD-SIM\nSD-SIM\nSN2N-1f (full aug)\nf\ng\nl\nSN2N-1f \n(full aug)\nSupervised-1f\n(w/o aug)\nGT\nSN2N-1f \n(full aug)\nSN2N-1f\n(w/o aug)\nSupervised-1f\n(w/o aug)\nSN2N-500f\nSupervised-\n500f\nGT\nSD-SIM\nSD-SIM\nSN2N-1f \n(full aug)\nGT\nh\ni\nj\nk\nn\no\nSupervised-1f\n(w/o aug)\nSD-SIM\nm\na\nSD-SIM\nd\nSpatial frequency (µm–1)\nCorrelation\n0.2\n0.4\n0.6\n0.8\n1.0\n0\n25\n20\n15\n10\n5\n0\nb\n240\n210\n180\n150\n120\nDistance (nm)\nLRQ\n0.20\n0.25\n0.30\n0.35\n0.40\n0.10\n0\nc\nSSIM\n0.2\n0.4\n0.6\n0.8\n1.0\n0\nSN2N\nSN2N\n(w/o c)\nN2V\nPURE\nSD-SIM\nRaw\nSupervised\n-1f \nSN2N-1f\n0\n0.2\n0.4\n0.6\n0.8\n1.0\nSSIM\nRaw\nSupervised\n-1f \nSN2N-1f \n0\n0.2\n0.4\n0.6\n0.8\n1.0\nSSIM\nData\nuncertainty\nModel\nuncertainty\nSTD\n4\n6\n8\n10\n12\n14\n16\nSupervised (w/o aug)\nSN2N (w/o aug)\nSupervised-1f (w/o aug)\nSN2N-1f (w/o aug)\nSN2N-1f\nSN2N-500f\nSupervised-1f (w/o aug)\nSN2N-1f (full aug)\nSupervised-1f (w/o aug)\nSN2N-1f (full aug)\nSupervised-1f (w/o aug)\nSN2N-1f (full aug)\nTraining data amount (frame)\n1\n500\n100\n10\nSSIM\n0.5\n0.6\n0.7\n0.8\n0.9\n1.0\nExposure time\n1×\n10×\n5×\n2×\n0.5\n0.6\n0.7\n0.8\n0.9\n1.0\nSSIM\nW/o\naug\nFull \naug\nP2P\naug\nBasic\naug\nSSIM\n0.5\n0.6\n0.7\n0.8\n0.9\n1.0\nSD-SIM\n90\n120\n150\n180\n210 240 (nm)\nPURE\nN2V\n100×\nexposure\nSN2N\nSN2N (w/o c)\n1/7\n100× exposure\nPURE\nN2V\nSN2N (w/o c)\nSN2N\nFig. 2 | Systematical evaluations in SD-SIM experiments using known \nstructures. a, Benchmarking on the Argo-SIM slide under the SpinSR10 SD-SIM \nsystem. Representative images are presented below the corresponding intensity \nprofiles indicated by the white line. Top row, low-SNR SD-SIM images (left), and \nPURE (middle) and N2V (right) denoising results; bottom row, our SN2N without \nself-constrained loss (w/o c), full SN2N denoising results and high-SNR (with 100× \nexposure) SD-SIM images. Intensity profiles of the double-line pair are displayed \nas insets. b, The LRQ values of images in a. c, Average SSIM values (n = 10). d, FRC \nanalysis of images in a. e, The STD of the denoising images predicted from ten \nrepetitively collected noisy images. f, Qdot 525 (QD525)-labeled microtubules in \nfixed COS-7 cells imaged by SD-SIM (left) and SN2N (right) trained with one frame \nand full augmentation (‘SN2N-1f (full aug)’). g, Enlarged region enclosed by the \nwhite boxes in f. From left to right: Raw image under SD-SIM, GT image, supervised \nlearning method (‘Supervised-500f’), SN2N results trained with 500 frames \n(‘SN2N-500f’), supervised learning (‘Supervised-1f (w/o aug)’), SN2N (‘SN2N-1f \n(w/o aug)’) trained with one frame and no augmentation, and SN2N trained \nwith one frame and full augmentation (‘SN2N-1f (full aug)’). h–j, Average SSIM \nvalues of different training data amounts (n = 10) (h), different models trained by \nimages under different exposure time (n = 10) (i) and different data augmentation \nstrategies (n = 10) (j). k, Data and model uncertainties quantified by the STD of \nten independent frames and ten independently trained models, respectively. \nl,m, Representative denoising data in fixed COS-7 cells of lysosomes labeled with \nLAMP1-EGFP (l) and mitochondria labeled with Tom20–mGold1 (m). n,o, SSIM \nvalues of imaging results in l (n) and m (o) (n = 10). In the box blots, the center line \nindicates the median, box limits indicate the 25th and 75th percentiles, and the \nwhiskers represent the maximum and minimum values; error bars indicate the \ns.e.m. Experiments were repeated ten times independently with similar results; \nscale bars, 1 µm (a and f), 500 nm (g) and 2 µm (l and m).\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\ncorrelation (FRC)39, structural similarity (SSIM)40, peak signal-to-noise \nratio (PSNR) and root-mean-square error (RMSE) metrics (Fig. 1b–e \nand Extended Data Fig. 3b). On the other hand, when examining the \nhigher SNR conditions (level 2 and level 3), all methods can achieve \nacceptable denoising results (Fig. 1e and Extended Data Fig. 3a) and \nthe improvements are less impressive, in which the results of N2V and \nSN2N without constraint (SN2N w/o c) reach closer to the full SN2N \n(Supplementary Table 1).\nSecond, SN2N is data efficient and can be trained even on one \nsingle image. The requirement of large datasets for DNNs to capture \nthe accurate data distribution of high noise level is moderated by our \nself-constrained learning process. To test this, we measured the slopes \nof SSIM, PSNR and RMSE metrics of different learning-based methods \nby decreasing the training data size by a factor of 5 and 50 (one frame \nonly; Fig. 1f, Extended Data Fig. 3c,d and Supplementary Table 2). As \nthe data pool shrunk, the denoising quality of the supervised learning \nmethod dropped quickly, for example, k = 4.3 for PSNR (Fig. 1f), and this \nvariation also dramatically effected the performances of N2V (k = 2.7) \nand SN2N w/o c (k = 2.9). In contrast, it is clearly observed that the full \nSN2N moderated the influence (k = 1.2). Our self-constrained loss helps \nthe network to learn the denoising process more efficiently, which is \nfurther strengthened by the developed Patch2Patch (k = 0.9, SN2N \nwith augmentation, SN2N w a). Fundamentally, these two aspects are \nhighly correlated with the data and model uncertainties of the DNNs41, \nin which the denoising effectiveness reflects the data uncertainty, and \ndata efficiency represents the model uncertainty. Using ten acquisi-\ntions’ predictions and ten repetitively trained models’ predictions, we \nmeasured the data and model uncertainties, respectively. Consistently, \nit can be seen that our SN2N outperforms other methods (Extended \nData Fig. 3e and Supplementary Figs. 3 and 4). Finally, we applied \nN2V19, parametric probabilistic Noise2Void (PPN2V)42, Self2Self (S2S)43, \nRecorrupted2Recorrupted (R2R)44, Noise2Fast (N2F)33 and our SN2N \non the single-frame simulation dataset (Supplementary Fig. 5a), and we \nfound that SN2N reached the closest to the GT, reflecting the advanced \ndenoising ability using a limited data amount.\nExperimental evaluation using standard sample under SD-SIM. SR \nconfocal microscopy can double the spatial resolution by narrowing \nits pinhole size, but correspondingly the number of photons reaching \nthe detector is severely restricted25. Although the photon reassignment \nconcept has mitigated this inherent low-SNR condition25, the photon \nefficiency of SR confocal microscopy still needs to be improved for \ncapturing fast long-term suborganelle dynamics in multiple dimen-\nsions (5D in xyz-color-time), especially for its paralleled version, that \nis, SD-SIM26. Next, we experimentally evaluated our SN2N with GT \nunder a commercial SD-SIM system (Olympus SpinSR10 with an sCMOS \ncamera; Methods). Compared to the diffraction-limited confocal mode \n(Extended Data Fig. 4d), the integrated optical reassignment module \nand ×3.2 second-stage magnification system of SpinSR10 bring heavily \nreduced SNR conditions. Using a commercial Argo-SIM slide (Meth-\nods), we measured the variance across the vertical straight line, and \nonly SN2N could draw the expected brightness profile with minimum \nfluctuations, even smaller than that of image under 100× exposure \n(Fig. 2a). Furthermore, the double-line structures are highlighted by \nSN2N from noise with the highest contrast (Fig. 2a and Supplementary \nVideo 2). Based on several metrics including line restoration quality \n(LRQ; Methods)12 according to the known line structures, SSIM against \nthe 100× exposure generated GT, FRC and STD values from repetitive \nacquisitions, as well as visual inspection, we found SN2N successfully \nrestored real-world collected images with superior stability and quality \n(Fig. 2b–e, Extended Data Fig. 4a and Supplementary Table 3).\nExperimental evaluation by biological samples under SD-SIM. To \nexamine the broad applicability of SN2N, we applied it on microtubules \nand outer membranes of lysosomes and mitochondria in fixed COS-7 \ncells under another commercial SD-SIM system (GATACA Systems, \nLive-SR with an sCMOS camera; Methods) (Fig. 2f,l,m). Generally, the \nultimate goal of unsupervised learning-to-denoise is to approach the \nsimilar performance of supervised learning without requiring the \ndata pairs and large training set. In this case, we focused on compar-\ning our SN2N against the supervised learning method in the denois-\ning performance and data efficiency (Fig. 2f,l,m), in which the 500× \nexposure result is considered as the GT. When only using one frame, the \nsupervised learning method produces obscure structures (Fig. 2g,l,m), \nindicating an underfitting configuration, and the performance grows \ndramatically when increasing the data amount from 1 to 500 frames \n(Fig. 2g,h). In contrast, we found the expansion of the data pool has \nlittle effect on the denoising performance of SN2N (Fig. 2g,h), and \nthe integration of our Patch2Patch data augmentation (Methods) \nhelps the one-frame-trained model to approach the performance of \nthe 500-frame-trained model (Fig. 2j and Supplementary Table 5). \nFurthermore, by taking advantage of its intrinsic data efficiency and \nfurther reasonable data augmentation, SN2N exploits the full potential \nof available data, producing superior model and data uncertainties \n(Fig. 2k). Interestingly, we found that the one-frame-trained SN2N \nis less affected by the SNR degradations (Fig. 2i and Supplementary \nTable 6), prompting us to explore the trickier denoising task of multi-\ncolor live-cell SD-SIM data.\nRL-SN2N on SD-SIM unlocks fast long-term imaging across 5D\nMulticolor live-cell SR imaging. The Richardson–Lucy (RL) deconvo-\nlution45,46 has been routinely applied on SD-SIM to enhance the contrast \nand resolution and is especially useful in scenarios containing strong \nout-of-focus signals or requiring precise segmentation. However, \nit is prone to artifacts under live-cell imaging conditions (Fig. 3b). \nThus, beyond denoising, we integrate the RL deconvolution with our \nSN2N (RL-SN2N; Methods) for simultaneously achieving the artifacts \nremoval and resolution enhancement (Fig. 3a,b). To avoid breaking the \nFig. 3 | Multicolor live-cell SR imaging enabled by RL-SN2N on SD-SIM.  \na, Workflow of RL-SN2N (Methods). In the training stage, RL deconvolution was \napplied after self-supervised data generation to create the RL-SN2N training \nset. In the inference stage, the denoised data are used for segmentation and \ndownstream analysis. b, A representative example for dual-color SR imaging \nof mitochondria (green) and ER (magenta) labeled with Tom20–mCherry \nand Sec61β-EGFP in live COS-7 cells under SD-SIM (top left), SD-SIM after RL \ndeconvolution (top right, RL SD-SIM), SN2N result (bottom left) and RL-SN2N \nresult (bottom right), alongside the enlarged region of the yellow dashed box.  \nc, A representative example for four-color imaging of the mitochondria (green), \nER (gray), lysosomes (red) and GA (blue) labeled with MitoTracker Deep Red \nFM, Sec61β-EGFP, Lamp1–mCherry and Golgi-BFP in live COS-7 cells under raw \nSD-SIM (right) and RL-SN2N (left). d, The yellow dashed box in c is enlarged and \nshown at seven time points under different configurations. From top to bottom: \nRaw SD-SIM, RL-SN2N, RL SD-SIM segmentation (by Otsu hard threshold) and \nRL-SN2N segmentation (by Otsu hard threshold) results. The lines in different colors \nindicate different interaction events (E1–E4). e, Segmentation results of the \nmitochondria (orange, by Mitonet) and ER (by ERnet) under SD-SIM (left), and \nRL-SN2N (right). The ERnet segmentations contain tubules (cyan), sheets (yellow) \nand sheet-based tubules (magenta, SBTs). f, Trajectories of lysosomes exhibiting \ndirected motion (green), free diffusion motion (blue) and confined motion (red), \nand ER tubules (black). g, Distribution of the estimated α values of lysosomes \nversus their temporal average distances to ER tubules. h, Average mitochondrial \ndiameters (n = 100; Methods). i, Average numbers of events surpassing the MOC \nthreshold (>0.26) of Lys–Mito (n = 46). DM (green), directed motion; FDM (blue), \nfree diffusion motion; CM (red), confined motion. In the box blots, the center line \nindicates the median, box limits indicate the 25th and 75th percentiles, and the \nwhiskers represent the maximum and minimum values; error bars indicate the \ns.e.m. Experiments were repeated ten times independently with similar results; \nscale bars, 2 µm (c–e), and 5 µm (b).\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\na\nc\nd\nSegmentation\nE1: Lys-GA with Mito; E2/E3/E4: Lys-GA with ER\n13.6 s\n19.5 s\n32.3 s\n49.3 s\n60.3 s\n63.7 s\n83.3 s\nRL-SN2N\nSD-SIM\nRL-SN2N 30.0 s\nSD-SIM 30.0 s\nRL-SN2N\nRL SD-SIM\nb\ne\nf\ng\nh\ni\nStep 3. Self-constrained\nlearning\nRL-SN2N\nSteps 1–2. Self-supervised data generation\nStep 1. RL\ndeconvolution\nNoisy data\nData pair\nRL data pair\nStep 1. Diagonal resampling\n& Fourier ×2\nStep 2. RL deconvolution\nxy1\nxy2\nxy1\nNet (xy1)\nNet (xy2)\nxy2\nStep 3. Organelle\nsegmentation\nStep 2. SN2N\ninference\nTrained model\nTraining stage\nStep 4. Downstream\nanalysis\nDenoised data\nStatistics\nSmart algorithm\nSegmentation\nInference stage\nSD-SIM\nRL-SN2N\nNumber of events\n50\n60\n70\nDM FDM CM\nDirected motion \nConfined motion \nERnet tubules\nFree diffusion\nMito diameter (µm)\nLys-Mito\nER-Mito\n0\n0.5\n1.5\n1.0\nContacted\nNot contacted\nMitonet\nERnet SBTs\nERnet sheets\nERnet tubules\n0\n0.4\n0.8\n1.2\n1.6\nMean distances (µm)\nMSD coefficient (α)\n0\n0.1\n0.2\n0.3\n0.4\nConfined motion \nFree diffusion \nDirected motion \nER-Lys\nLys-Mito\nRL-SN2N\nSN2N\nRL SD-SIM\nSD-SIM\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nRL-SN2N\nSD-SIM\nSD-SIM\nNucleus\n3D rendering\nx\nyz\n~46 µm\n~46 µm\n~6 µm\n0 h 10 min\n0 h 30 min\n0 h 35 min\n0 h 40 min\n0 h 50 min\n0 h 55 min\n1 h 5 min\n1 h 0 min\n1 h 10 min\n0 h 0 min\nSD-SIM\nRL-SN2N\na\n2D-UNet\n3D-UNet\nb\nc SD-SIM\nSegmetation\ne\nf\ng\nyz\nSD-SIM\nRL-SN2N\nSD-SIM\nRL-SN2N\nSD-SIM\nRL-SN2N\n5 µm\n0 µm\n(z)\nxz\nSegmentation\nRL-SN2N\n5 µm\n0 µm\n(z)\nRL-SN2N\n1 h 0 min\n1 h 10 min\n1 h 0 min\n1 h 10 min\nSD-SIM\n500 nm\ny\nz x\n7.6 µm\n7.6 µm\nConvolution\nDownsampling\nUpsampling\nSkip connection\nh\n1 h 50 min\n2 h 55 min\ni\n1 h 10 min\nER\nMito\nd\n39 min\n33 min\n3 min\n0 min\n2.8 µm\n0.4 µm\n(z)\nyz\nxz\nxz\nyz\nyz\nxz\nxz\nyz\n0 min\nRL-SN2N\nxz\nyz\n12 min\nxz\nyz\n27 min\nSD-SIM\nxz\nyz\nxz\n57 min\nyz\nxz\n42 min\nyz\nxz\n27 min\nyz\n5 µm\n0 µm\n(z)\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\npixel-wise noise independence, the deconvolution is executed after \nthe self-supervised data generation step during the training stage, \nand at the inference phase, the trained SN2N model is fed with the \ndirectly deconvolved images. We first validated the performance of \nRL-SN2N using the Argo-SIM slide, and its LRQ contrast substantially \noutperformed SN2N (Extended Data Fig. 4b,c).\nAlso, we expect the resulting high-quality SR images can \nempower precise segmentation, facilitating automated analysis at \nsuborganelle-level precision from massive data (see the prototype \nin Fig. 3a and Extended Data Fig. 5a). The first representative case is \ndual-color imaging of endoplasmic reticulum (ER) and outer mito-\nchondrial membrane (OMM) labeled live COS-7 cells (Supplementary \nVideo 3). We can observe that our RL-SN2N effectively extracts the \nfluorescence signal from the noisy background with further enhanced \ncontrast (Extended Data Fig. 5b,c), hence many relative movements \nbetween OMM and ER can be dissected. Predictably, the direct hard \nthreshold segmentation (Methods) on the low-SNR data produced \nhighly broken structures and strong false positives. Our RL-SN2N mini-\nmized such noise-induced segmentation artifacts, allowing inspection \nof fast mitochondrial fission events near the ER–mitochondria contact \nsites (Extended Data Fig. 5c).\nNext, we extended the application to four-color live-cell imaging \nof lysosomes (Lys), Golgi apparatus (GA), mitochondrial matrix protein \n(Mito) and ER (Fig. 3c and Supplementary Video 4). Consistently, our \nRL-SN2N provided plausible reconstructions for monitoring multiple \norganelle interactions from obscure captures. The increases in SNR \nand contrast were stable during time-lapse SR imaging (Fig. 3d), which \nled to distinct observations of Lys–GA (GA being surrounded by Lys) \nwith Mito (event 1, E1) and Lys–GA with ER (E2–E4) interactions. In \nthese events, we found that Mito or ER settled in a circle could anchor \nonto a moving Lys–GA and be pulled out for spatial movements as \nit followed the trajectory of Lys–GA (segmentations in Fig. 3d)47,48. \nFurthermore, we applied two recently developed learning-based seg-\nmentation methods, that is, the ERnet49 and Mitonet50, on our ER and \nMito data (Methods). Although better than the hard threshold segmen-\ntations, we found these well-trained networks were still influenced by \nthe noise conditions (Fig. 3e and Extended Data Fig. 5d–g). On the other \nhand, empowered by our RL-SN2N, ERnet could precisely draw the \nER topology of tubules, sheets and sheet-based tubules (Fig. 3e). The \ntrajectories of Lys and spatial masks of ER tubules (Methods and Fig. 3f) \nenabled us to examine the correlation between motions of lysosomes \nand the ER network. By calculating the mean square displacement \n(MSD) of Lys and their distances to ER, the Lys with confined motion \nbehaviors were mostly located adjacent to ER48 (Fig. 3g and Extended \nData Fig. 5h,i,k). Mitonet-generated masks allowed us to calculate the \nMander’s overlap coefficient (MOC) of Lys with the nearest Mito for \nidentifying potential functional sites of Lys–Mito contact sites24, that \nis, MOC > 0.26 as a potential event (Methods and Extended Data Fig. 5j), \nin which the Lys with directed motion behaviors exhibited a relatively \nlarger number of events compared to the free diffusion and confined \nmotions (Fig. 3i). We also quantified the mitochondrial diameters at \nthe ER–Mito or Lys–Mito contact sites (Methods and Fig. 3h), reflecting \nthe fact that the constricted loci in Mito preferentially associated with \nthe ER–Mito contact48. Together, in the absence of tedious manual pro-\ncesses, our RL-SN2N on SD-SIM system simplified the dissection of the \nsynergy of different organelles at a suborganelle scale (Extended Data \nFig. 5a). Comparatively, N2V19, PPN2V42, S2S43, R2R44, N2F33, DeepCAD22, \nDeepSeMi21 and SRDTrans34 cannot provide valid denoising results from \ndata under such ultralow-SNR conditions (Supplementary Fig. 6a).\n3D extension for volumetric SR imaging across 5D. With roots in the \nconfocal microscopy, SD-SIM can perform 3D SR imaging with further \nenhanced axial contrast by the 3D deconvolution. To fully utilize the \naxial information, we extend the RL deconvolution and U-Net from \nits 2D37 version to its 3D mode (Fig. 4a and Extended Data Fig. 4f,g). \nBy our RL-SN2N, 3D mitochondrial networks were visualized, and, as \nexpected, it revealed tori cross-sections of OMM genuinely, which \nwere ambiguous in the raw SD-SIM results (Fig. 4b and Supplementary \nVideo 5). Various types of OMM structures, that is, from tubular to a \nseries of different structures including fragments, small vesicles and \nspheroids, can be resolved by our RL-SN2N (Fig. 4d). The segmentation \nalso helps to highlight the hollow structures of OMM networks, which \nare hardly distinguishable in the raw images (Fig. 4c). Furthermore, \nbenefiting from SNR and contrast reinforcements, we can capture the \nFig. 4 | 3D RL-SN2N on SD-SIM unlocks fast long-term imaging across 5D.  \na, The extension of 2D U-Net (top) to its 3D mode (bottom). b–d, 3D OMM network \nimaging of Tom20–mCherry-labeled live COS-7 cells. b, Color-coded volumes of \nraw SD-SIM (top) and RL-SN2N (bottom). The color-coded axial views (yz and xz \nplanes) indicated by the yellow dashed lines are provided alongside. Magnified  \nview of yellow boxed regions on xz section is shown at the bottom left of the  \nimages. c, 3D rendering views of the white boxed region in b under raw SD-SIM \n(1st column), segmentation of SD-SIM (2nd column), SN2N (3rd column) and \nsegmentation of SN2N (4th column). d, 2D slices under raw SD-SIM (top) and  \nRL-SN2N (bottom). e,f, 4D imaging of OMM network (Tom20–mCherry) in live COS-\n7 cells. e, Representative color-coded volumes and their xz and yz cross-sections \nat six time points. f, Magnified views and their xz and yz cross-sections of the white \nboxed region in e under RL-SN2N at four time points. The red and white arrows \nindicate the mitochondrial fission and before fission, respectively. g, 5D imaging \nof mitochondria (green, mGold-Mito-N-7), ER (magenta, DsRed-ER) and nucleus \n(blue, SPY650-DNA) in live COS-7 cells. Representative 3D rendering views of the \ncell mitosis process at ten time points under RL-SN2N, except for the first (left part) \nand the last views being under raw SD-SIM. h, Magnified views from yellow boxes in \ng under raw SD-SIM (left) and RL-SN2N (right). i, Two representative time points of \nmitochondria (green) and ER (magenta) after mitosis. Experiments were repeated \nfive times independently with similar results; scale bar, 5 µm (b and i), 1 µm (c, d and \nh), 10 µm (e) and 2 µm (f).\nFig. 5 | SN2N and RL-SN2N permit long-term live-cell STED imaging. a, Results \nof live HeLa cells labeled with SiR-tubulin under STED (center pixel, STED \ndepletion laser power as 50%, left) and its SN2N result (STED-SN2N, right). Data \nfrom ref. 52 (Methods). b, Magnified views of the white boxed region in a under \ndifferent imaging configurations. FRC-measured resolution values are labeled. \nc, A sketch of different pixel assembly strategies: center pixel (‘Cent. pix. ‘), 5 × 5 \nsum and APR. d, Fluorescence intensity profiles along the white arrow in b of \nconfocal/STED (center pixel, left) and their SN2N results (right). e, FRC analysis \nof the images in b. f, Average FWHM values (n = 6, positions). g, STED snapshots \nof live COS-7 cells labeled with SiR-Tubulin (top), LifeAct-EGFP (middle) and \nSec61β–EGFP (bottom) under a commercial STED microscope (Leica; Methods). \nh, Three representative time points of STED (left) and SN2N (right) counterparts \nof magnified views of the white-boxed regions in g. i,j,m, A representative \nexample of PKMO-labeled live COS-7 cells imaged under different conditions \nusing another commercial STED microscopy (Abberior). i, Three STED frames \nunder high depletion power (86%) and long duration time (100 μs per pixel). \n j, Representative frames of STED (left) and SN2N results (right) at high depletion \npower (86%) and short duration time (10 μs per pixel). k, Photobleaching analysis \nof STED images used in i, j and m. l, Fluorescence profiles along the white arrow \nin j. m, Long-term imaging of STED (left) and RL-SN2N (right) at low depletion \npower (41%) and short duration time (10 μs per pixel). n, Magnified views of the \nwhite boxed regions in m. o,p, Representative montages of the mitochondrial \nfusion (o) and fission (p) events. The yellow and blue arrows highlight the regions \nof mitochondrial fusion and fission, respectively, while the white arrows indicate \nthe moments before events. In the box blots, the center line indicates the median, \nbox limits indicate the 25th and 75th percentiles, and the whiskers represent the \nmaximum and minimum values; error bars indicate the s.e.m. Experiments were \nrepeated five times independently with similar results; scale bar, 2 µm (a), 1 µm \n(b, g and h) and 500 nm (i and m–o). a.u., arbitrary units.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nSN2N\n243 nm\nConfocal\n267 nm\nCenter pixel\nSTED-SN2N\nCenter pixel\nSTED\nSTED 50%\nAPR\nCent. pix.\n5 × 5 sum\na\nb\nc\nSN2N\nd\ne\nf\n0\n1.0\nIntensity (a.u.)\nDistance (nm)\nCenter pixel confocal/STED\n0\n1.0\nDistance (nm)\nIntensity (a.u.)\nSTED power (%)\nFWHM (nm)\nSTED (%)\nSpatial frequency (µm–1)\n0\n2.5\n12.5\n5.0\n7.5\n10.0\nCorrelation\ng\ni\nj\nk\nl\nm\nn\nRL-SN2N\no\n0 min 0 s\n1 min 40 s\nSTED\nSN2N\nRL-SN2N\nSTED 0%\nSTED 8%\nSTED 16%\nSTED 30%\nSTED 50%\nSTED\nCenter pixel\n148 nm\nSN2N\nCenter pixel\n91 nm\nSTED -ISM\n166 nm\nSTED-ISM\n173 nm\nSN2N\nCenter pixel\n101 nm\nSTED\nCenter pixel\n155 nm\nSN2N\nCenter pixel\n104 nm\nSTED-ISM\n203 nm\nSTED\nCenter pixel\n193 nm\nSTED -ISM\n206 nm\nSN2N\nCenter pixel\n109 nm\nSTED\nCenter pixel\n210 nm\nConfocal\nCenter pixel\n216 nm\nSN2N\nCenter pixel\n118 nm\nISM\n5 × 5 APR\n213 nm\nSTED 0%\nh\nCenter pixel confocal/STED-SN2N\n0\n1\nHigh power\nlong duration\nHigh power\nshort duration\nLow power\nshort duration\n7 min 55 s\n3 min 45 s\n5 min 50 s\n0 min 0 s\n0 min\nSTED\nSN2N\n3 min\n6 min\n3 min\nSN2N\nSTED\n0 min\n6 min\nSTED\nSN2N\n6 min\n3 min\nSN2N\nSTED\n0 min\n1 min 40 s\n3 min 45 s\n7 min 55 s\n0 min 0 s\n0 min 0 s\n2 min 0 s\n2 min 30 s\n3 min 45 s\n4 min 10 s\n18 min 0 s\n16 min 15 s\n16 min 40 s\n24 min 35 s\nIntensity (a.u.)\nDistance (nm)\nA\n0\n0.2\n0.4\n0.6\n0.8\n1.0\n300\n0\n250\n200\n150\n100\n50\n76 nm\nIntensity (a.u.)\nTime (min)\n0\n0.2\n0.4\n0.6\n0.8\n1.0\n0\n10\n20\n30\nCenter pixel \nconfocal/STED\nCenter pixel \nSN2N\nSN2N\nSTED\n5 × 5 APR\n5 × 5 APR\n5 × 5 APR\n5 × 5 APR\n5 × 5 sum\n5 x 5 sum\n300\n0\n250\n200\n150\n100\n50\n300\n0\n250\n200\n150\n100\n50\n0\n20\n40\n60\n80\n50\n100\n150\n200\n250\n300\n0\nISM/STED-ISM\nCenter pixel confocal/STED\nCenter pixel SN2N\nConfocal/STED\nLow power short duration\nHigh power short duration\nHigh power long duration\n3 min 45 s\n3 min 45 s\n1 min 40 s\n1 min 40 s\n7 min 55 s\n0 min 0 s\n0 min 0 s\n0 min 0 s\n0 min 0 s\n1 min 40 s\n3 min 45 s\n5 min 50 s\n7 min 55 s\nSTED\nSTED\nSN2N\nSN2N\n80\n0\n5\n8\n12\n16\n19\n30\n50\np\n1/7\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nfour-dimensional (4D) OMM network dynamics across hours (Fig. 4e, \nSupplementary Fig. 10 and Supplementary Video 5). In our RL-SN2N \nresults, the kiss-and-run events happening in 3D are correctly identi-\nfied (Fig. 4f), which might be misinterpreted as the fissions in 2D slices \nunder low-SNR and contrast conditions. Finally, we extended our test \nto the challenging 5D SR imaging, revealing the dynamics of ER, Mito \nand nuclei during the entire cell mitosis process in all three dimensions \nover a long duration of 3 h (Fig. 4g, Supplementary Fig. 11 and Supple-\nmentary Video 6). Empowered by RL-SN2N, we clearly monitored the \nsegregation process (Fig. 4h) and the interactions (Fig. 4i) of dense \nER–Mito networks.\nAdditionally, according to the simulations, we found our \napproach is not sensitive to pixel size (Extended Data Fig. 2a). Thus, \nour success in SD-SIM equipped with an sCMOS camera (~38 nm pixel) \nmade us wonder whether our SN2N and RL-SN2N can be applied to \ndiffraction-limited SD-confocal microscopy (~65 nm pixel) or SD-SIM \nwith an EMCCD camera (~94 nm pixel) (Methods). On the Argo-SIM \nslice, our SN2N removed the readout noise of SD-confocal images \n(Extended Data Fig. 4d,e). Likewise, by applying RL-SN2N (with further \n2× upsampling) on two-color whole-cell volumes from the EMCCD \nSD-SIM, we successfully recorded two intermediate phases of cell mito-\nsis, in which the 3D mitochondrion and nuclear protein distributions \ncame into presence from background noise (Extended Data Fig. 6).\nEmpowering long-term live-cell STED imaging\nSimilarly to confocal SR microscopy, the increase in spatial resolution of \nSTED microscopy from depletion brings dramatically decreased SNR51. \nBecause the depletion is driven by de-excitation through stimulated \nemission, enlarging the depletion laser power will inherently decrease \nthe emission brightness of fluorescence labels and cause adverse effects \nsuch as photobleaching and phototoxicity, preventing long-term moni-\ntoring of samples51,52. Thus, the SN2N extraction of structures from few \nphotons gives the possibility to tackle the contradiction for imaging live \ncells with both high spatial resolution and SNR. First, we systematically \nevaluated SN2N’s denoising performances under different depletion \npowers (Fig. 5a,b and Supplementary Fig. 12), in which the data52 were \ncollected from a custom STED setup incorporating a SPAD array detec-\ntor with 5 × 5 elements for adaptive pixel-reassignment (APR; Methods). \nBy assembly of SPAD signals collected from the center pixel, the spatial \nresolution of STED is further improved but most photons would be \ndropped. From ref. 52, the 5 × 5 APR STED results, or equivalently the \nSTED image scanning microscopy (STED-ISM), can fully exploit the \ncollectible photons to preserve SNR (Fig. 5c). Differently, without these \ncollected photons from the circumjacent 24 pixels, our SN2N effectively \nrecovered the live microtubule structures from only photons collected \nfrom the center pixel, especially under a high depletion power (Fig. 5b). \nThe two-peak draws of microtubule intersections, FRC curves and \nfull width at half maximum (FWHM) measurements also highlight the \nincrement process of spatial resolution without sacrificing SNR after \nSN2N denoising (Fig. 5d–f).\nIt is straightforward to apply our SN2N on a commercial STED \nsystem (Leica, TCS SP8 STED 3X; Methods) for time-lapse imaging of \nmicrotubule-, actin- and ER-labeled COS-7 cells (Fig. 5g,h and Supple-\nmentary Video 7), routinely enabling high-quality live-cell SR imaging. \nBeyond that, we also recorded the long-term dynamics of mitochon-\ndrial cristae (PK Mito Orange, PKMO53 labeled) using another com-\nmercial STED system (Abberior Instruments, STEDYCON; Methods) \nunder various imaging conditions (Supplementary Video 7). To acquire \nhigh-quality STED images in situ, the high depletion power and long \npixel duration time were applied and hereby instantly extinguished \nthe fluorescence signal before five frames of recording (Fig. 5i). The \nhigh depletion power with short pixel duration time could delay the \nbleaching effects without loss of spatial resolution but create noisy \ncaptures, in which our SN2N effectively restored this SNR degrada-\ntion (Fig. 5j,k). The measured cristae-to-cristae distance reflects the \nachievable resolution of ~76 nm (Fig. 5l). Further turning down the \ndepletion laser power will offer significantly less susceptibility to pho-\ntobleaching but short of spatial resolution maximization (Fig. 5k,m). \nFortunately, our RL-SN2N can lift the dropped resolution in the absence \nof amplified photobleaching (Fig. 5m,n). Short exposure and low illu-\nmination energy facilitated us to record the perplexing locomotion of \ncristae during mitochondrial fusion (Fig. 5o) and fission (Fig. 5p) over \nhalf an hour. In contrast, conventional STED produced noisy SR images \nunder the same conditions (Fig. 5m).\nImproving reconstruction efficiency of SOFI\nSOFI30 can routinely break the diffraction limit by exploiting the natural \ntemporal fluctuations of fluorescence emissions under optical systems \nin their native states. However, the statistical uncertainty of reconstruc-\ntions from short sequences may dramatically affect image continuity \nand homogeneity, which leads it to generally requiring hundreds of raw \nimages (~1,000 frames for 2nd order) to preserve structural integrity31. \nTo increase its reconstruction efficiency, we integrated our SN2N solu-\ntion into the SOFI reconstruction pipeline (Fig. 6a). Specifically, the raw \nimage sequence is calculated by the nth-order cumulant (core SOFI), and \nthe resulting image is followed by the RL-SN2N procedure (Methods). \nUsing SIM as a reference, the 2nd-order SOFI-SN2N is first validated on \na wide-field microscope (Methods). We found that the strong snowflake \nartifacts in conventional 20-frame SOFI were effectively eliminated, and \nthe original microtubule structures were highlighted with high axial \ncontrast (Fig. 6b, Extended Data Fig. 7a and Supplementary Video 8). \n \nThe two-peak analyses, FRC metrics and FWHM measurements all \ndemonstrate the massively improved temporal resolvability and effec-\ntively increased spatial resolution of SOFI-SN2N (Fig. 6b–d). Next, we \nextended our SOFI-SN2N to different orders of cumulants executed \non a commercial SD-confocal microscope (Fig. 6e–h and Methods). \nConsistently, both visual examination and FRC analysis exhibited that \nthe 2nd-order SOFI-SN2N can produce artifact-free results from only \n20 frames (Fig. 6e–g). On the other hand, the increase of resolution \nfrom higher-order cumulants brings a cost of more frames needed, and \nour SOFI-SN2N enables efficient 3rd-order and 4th-order SOFI recon-\nstructions from 50 frames and 100 frames, respectively (Fig. 6f and \nExtended Data Fig. 7b–e). The FWHM- and FRC-measured resolution \nvalues (Fig. 6h) also evince the spatial resolution enhancement. Finally, \nthe recording of OMM dynamics during mitochondrial fusion during \n10 min gives us a glance at live-cell SN2N-SOFI SR imaging (Fig. 6i,j and \nSupplementary Video 8).\nIn fixed-cell experiments, the temporal sampling22,24 (first and \nsecond 20 frames for SOFI reconstruction) can be applied directly \n(Extended Data Fig. 7f). Interestingly, due to the inconsonant temporal \nfluctuation behaviors, the temporal sampling leads to statistical dif-\nferences between the adjacent frames and produces strong artifacts. \nIn contrast, our spatial sampling is impressed with sturdily extracting \nthe microtubule structures for requiring no temporal consistency \n(Extended Data Fig. 7f).\nSN2N on expanded samples\nIn addition to the use of fluorescence fluctuations, the ExM29 is another \nsystem-agnostic SR modality by artificially enlarging the size of samples \nto break the diffraction limit. However, considering the determined \nnumber of fluorophores, this space extension of samples will result \nin the geometrically decreased SNR according to the expansion times \n(Extended Data Fig. 7g–k). Under a wide-field microscope (Methods), we \noffered the ExM-SN2N enabling high-quality SR imaging of ~110 nm and \n~67 nm resolutions for 2× and 4× expansions (Extended Data Fig. 7g–i), \nrespectively. Under different noise levels, the ExM results after SN2N \ndenoising exhibited significantly improved signal-to-background \nratios. Similarly, the SN2N outcomes of the cells expanded by 4.5 times \nyielded (Extended Data Fig. 7j,k) noise-eliminated results, revealing \nthe complex ER tubule structures.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nFRC (nm)\nSN2N\nSOFI\nConf.\n100\n150\n200\n250\n300\n350\n100\n150\n200\n250\n300\n350\nFWHM (nm)\nConf.\nSN2N\nSOFI\na\nb\nc\nd\ni\nj\ne\nf\n15\n3\n6\n9\n12\n18\n0\n0\n0.2\n0.4\n0.6\n0.8\n1.0\nCorrelation\nSpatial frequency (µm–1)\nWide-field\nSOFI-20f\nSN2N-20f\n2D-SIM\nSOFI-SN2N\ng\nConfocal\n4th SN2N 2,000f\n2nd SOFI 20f\n2nd SN2N 20f\n3rd SOFI 50f\n3rd SN2N 50f\n4th SOFI 100f\n4th SN2N 100f\n2nd SOFI 20f\n2nd SOFI-SN2N 20f\n2nd SOFI-SN2N 20f\nSteps 1-3. Self-supervised data generation\nTraining stage\nInference stage\nStep 4. Self-constrained learning\nStep 1. Nth order\ncumulantion (SOFI)\nStep 2. Diagonal resampling\n& Fourier ×4\nStep 3. RL\ndeconvolution\nSOFI\nxy1\nxy2\nSOFI data pair\nNet (xy1)\nNet (xy2)\nStep 1. Direct SOFI\nreconstruction\nStep 2. SN2N inference\nTrained model\nDistance (nm)\nIntensity (a.u.)\n130 nm\n65 130 195 260 325\n0\n390\n0\n1.0\n65 130 195 260 325\n0\n390\nDistance (nm)\nIntensity (a.u.)\n0\n1.0\n65 130 195 260 325\n0\n390\nDistance (nm)\n0\n1.0\nIntensity (a.u.)\n2D-SIM\n2nd SOFI-SN2N 20f\nWide-field\n2nd SOFI 20f\nh\n120\n140\n160\n180\n200\n220\n240\n260\n280\nFRC (nm)\nWF\nSOFI\nSN2N\nSIM\nFWHM (nm)\nWF\nSOFI\nSN2N\nSIM\n120\n140\n160\n180\n200\n220\n240\n260\n280\n0\n1.0\nIntensity (a.u.)\n132 nm\nDistance (nm)\n65 130 195 260 325\n0\n390\n2nd SN2N 20f\n2nd SOFI 20f\nConfocal\nSpatial frequency (µm–1)\n15\n0\n0.2\n0.4\n0.6\n0.8\n1.0\n3\n6\n9\n12\n18\n0\nCorrelation\nConfocal\n2nd SOFI 20f\n2nd SN2N 20f\n10 min 40 s\n11 min 20 s\n12 min 40 s\nxy1\nxy2\n12 min 0 s\n12 min 0 s\n12 min 0 s\n12 min 0 s\n7 min 20 s\n8 min 0 s\n2nd SOFI 20f \n12 min 0 s\nConfocal\n2nd SOFI\n3rd SOFI\n4th SOFI\n1/7\n1/7\nFig. 6 | Integration of SN2N and SOFI massively improves the SR \nreconstruction efficiency. a, Workflow of SOFI-SN2N (Methods). In the training \nstage, self-supervised data generation was applied after nth-order SOFI and \nbefore Fourier upsampling. b, Cross-validation of SOFI-SN2N. Snapshots of \nmicrotubules in a COS-7 cell labeled with QD525 under wide-field microscopy \n(top left), 2D-SIM (bottom left), 2nd-order SOFI using 20 frames (2nd SOFI 20 f, \ntop right) and its SN2N result (2nd SOFI-SN2N 20 f, bottom right). The intensity \nprofiles and multiple Gaussian fitting indicated by the white arrows are provided. \nc, FRC analysis of the images in b. d, Average FWHM (top) and FRC (bottom) \nvalues (n = 5, measurements). WF, wide-field. e, Results of microtubules in a COS-\n7 cell labeled with QD525 under a SD-confocal microscopy reconstructed by 2nd-\norder SOFI using 20 frames. f, Zoomed views from the white box in e. First row: \nconfocal image (left) and 4th-order SOFI using 2,000 frames denoised by SN2N \n(4th SOFI-SN2N 2,000 f, right); the other rows, from top to bottom: 2nd-, 3rd- and \n4th-order SOFI reconstructions (left) using 20, 50 and 100 frames, respectively, \nand their SN2N results (right). g, FRC analysis of raw SD-confocal image, 2nd \nSOFI 20 f reconstruction, and its SN2N result. h, Average FWHM (left) and FRC \n(right) values of the images in f (n = 5, measurements). i, A representative live \nCOS-7 cell labeled with Skylan-S-TOM20 imaged by 2nd SOFI-20f of SD-confocal \nmicroscopy and its SN2N result. j, Zoomed views of OMM structures. Top three \nrows, from top to bottom: magnified views of the white boxed region in i under \nSD-confocal microscopy, 2nd SOFI 20 f reconstruction, and 2nd SOFI-SN2N 20 f \nresult; bottom three rows: montages of a representative mitochondrial fission \nevent. The yellow and white arrows highlight the mitochondrial fission and \nbefore fission, respectively. In the box blots, the center line indicates the median, \nbox limits indicate the 25th and 75th percentiles, and the whiskers represent the \nmaximum and minimum values; error bars indicate the s.e.m. Experiments were \nrepeated three times independently with similar results; scale bars, 2 µm (b, i and \nj), 5 µm (e) and 1 µm (f).\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nArtifact removal for live-cell SIM\nAlthough SIM is recognized to have a higher photon efficiency than \nother SR modalities, it still requires an adequate SNR for each raw \nimage to prevent random reconstruction artifacts10,12,24. Therefore, \nto minimize the artifacts, we included the SN2N module into SIM \nreconstruction (‘raw resampling’, Extended Data Fig. 8a). The gen-\nerated twin nine raw images were individually reconstructed by the \nHiFi-SIM54 procedure (high-fidelity SIM; Methods) to create the train-\ning sets, and at inference stage, the SIM images are taken as input by \nusing the original nine raw images. First, we systematically evaluated \nthe performance of our SIM-SN2N using the BioSR open-sourced \ndataset55 under high (H), medium (M) and low (L) SNR levels. Com-\nparing to the GT ultrahigh-SNR reconstructions, our SIM-SN2N effec-\ntively disentangled the real features from artifacts and produced \nstable SSIM values for various conditions and samples (Extended \nData Fig. 8b–g and Supplementary Video 9). Then, the 2D-SIM imag-\ning of mitochondrial cristae exhibited strong background artifacts, \nwhich was further amplified as the emission fluorescence progres-\nsively decreased (Extended Data Fig. 8h,i). The suppression of arti-\nfacts in SIM-SN2N supports its superior performance in obtaining \nhigh-fidelity SR images from raw images of low SNR (Extended Data \nFig. 8j and Supplementary Video 9).\nIn particular, when meeting ultrafast SIM imaging under \nultralow-SNR conditions, the resampled dataset might encounter \nreconstruction failures, in which the parameter estimation is highly \nunstable. On the other hand, the frequency-component reassembly \nstep in SIM reconstruction is usually accompanied by an artificial 2× \nupsampling, which breaks the pixel independency, and this prevents \nthe direct application of SN2N on SIM images. To meet this challenge, \nwe conducted a new strategy of spatial sampling by defining 3 × 3 pixels \nas one unit and a new average calculation of four directions followed \nby the Fourier 3× operation (‘SIM resampling’, Extended Data Fig. 8k). \nThis one-pixel interval enabled the direct artifact removal on the SIM \nreconstructed images only with a slight drop of SSIM metrics (Extended \nData Fig. 8l–q). Tested on the ultrafast (188 Hz, ~0.6 ms exposure per \nraw frame) TIRF-SIM experiments, the ‘raw sampling’ failed to offer \neligible SIM reconstructions, and in contrast, the ‘SIM resampling’ \nprovided high-quality images along 6,800 consecutive SR frames \n(Extended Data Fig. 8r–t).\nDiscussion\nIn the N2N ecosystem, theoretically, infinite data pairs are essential to \napproach the supervised learning methods’ performance, because of \nthe need for averaging training sets to remove the zero-mean noise. \nUsing twin images from our self-supervised data generation, the devel-\noped self-constrained learning process further generalizes this N2N \nconcept to remove noise with randomness, and also relaxes the need \nfor an infinite data amount17. As a result, our SN2N is competitive with \nsupervised learning methods but overcomes the need for a large train-\ning dataset and clean GT. Indeed, we showed a single noisy frame is \nfeasible for training. We applied SN2N to the photoelectric detector \ndirectly captured data, including two commercial SD-SIM systems \nwith two different types of cameras, one custom-built and two com-\nmercial STED microscopes, one ExM under a wide-field microscope \nand one commercial SD-confocal microscope, demonstrating its \nextensive application value and superior performance. We further \nintegrated SN2N into the prevailing SR reconstructions, including \nRL deconvolution, SOFI and SIM for artifact removal, enabling effi-\ncient reconstructions from limited photons by one-to-two orders \nof magnitude.\nA concern of learning-based recovery is that the spatially \ndenoised features might distort the temporal signals in a nonlinear \nmanner. Recording cells labeled with cytosolic Ca2+ indicators56, we \nfound SN2N acted without nonlinearly perturbing the amplitudes \nof different Ca2+ transients, indicating that the SN2N denoising is \nquantitatively accurate (Extended Data Fig. 9a–c). Furthermore, for \nfast imaging applications (Extended Data Fig. 9d,e), although SN2N \nperformed superiorly in spatial denoising, DeepCAD’s smaller tem-\nporal signal fluctuations led us to integrate its temporal resampling \nmethod into SN2N (SN2N (temporal)) for further strengthening the \ntemporal stability.\nAlthough SN2N was shown to work well overall, it would be appro-\npriate to discuss limitations observed in the current version. Here, we \nprovided three representative failure cases. The first example is for the \nultralow-SNR experimental data with strong baseline signal. Without \npercentile normalization (Methods), the resulting predictions will \nexhibit small background fluctuations (Supplementary Fig. 9c). The \nsecond failure case is the unlimited increasing of noise level and, pre-\ndictably, using a fixed amount of data, the SN2N progressively exhib-\nited noisy output (Extended Data Fig. 2i). The last failure case is the \ngeneralization errors inherited from the mechanism of unsupervised \nlearning-to-denoise. The retraining or the transfer learning for the new \ndataset is usually needed. Specifically, the extraction of structures from \nmodels trained by higher SNR data would exhibit strong background \nartifacts caused by the misleading information learned between noise \nand structures (Supplementary Fig. 9a,b). The outcomes from models \ntrained by lower SNR data are free from this failure (Supplementary \nFig. 9a,b). Furthermore, we directly fed the four-color live-cell imaging \ndata (Fig. 3c) to the simulation dataset trained model (Supplementary \nFig. 9d), and this rough test using data with different features and \nnoise levels resulted in stronger hallucinations. Although SN2N can be \nexecuted without clean data, the internal mechanism of learning-based \ndenoising methods is substantially different from the classical ones. \nIn an abstract sense, SN2N is learning to extract structures from noisy \ninput according to the training set, while the numerical algorithms \nusually intend to remove the noise. We could moderate this issue by \nthe iterative execution of the SN2N model (SN2N2) to shrink the gap \nbetween training and test sets (Extended Data Fig. 10a–f). In addition, \nthis network extraction is also influenced by the structural scale of \ninput images, and the upsampling/downsampling operation should be \napplied to match the pixel size before SN2N inference (Extended Data \nFig. 10g–k), otherwise the network would produce erroneous results \n(Extended Data Fig. 10h,i).\nRandom noise is an unavoidable obstacle in fluorescence \nmicroscopy, especially for the live-cell SR recording. Our SN2N and \nits extensions provide powerful solutions for routine 2D–5D imaging \nof suborganelle dynamics at ultrahigh spatiotemporal resolution and \nhigh fidelity for long durations. We anticipate that the elimination \nof noise could benefit precise structure segmentation and facilitate \nautomatic multi-parameter analysis, establishing a panoramic view \nof the organelle interaction systems49,50. SN2N is a model-agnostic \nsolution. The 2D/3D U-Nets used in this work are simple end-to-end \nbaselines, and the extensions to other advanced networks, such as the \nroutinely used residual/dense blocks57, adversarial training strategy58 \nand transformer-based solutions59, are straightforward. Furthermore, \nthe incorporations of SN2N with network-based deconvolution and \nSIM are expected to enable stable and efficient SR reconstructions for \nthe real-time potential. Finally, the loss realizations of our SN2N have \nmany variants, and the distance calculation by SSIM or in the Fourier \ndomain could lead to additional improvements.\nNotably, our self-constrained learning process with self-supervised \ndata generation have the potential to reduce the data uncertainty of \nlearning-based microscopy in general. The predictive fluctuations \ninduced by the noise in the input data might be effectively shrunk by \nthe integration of our self-constrained learning process. Beyond the \nhigh-resolution fluorescence imaging showcased in this work, we also \nexpect our SN2N can be applied to other sensitive modalities, includ-\ning stimulated Raman scattering microscopy, multiplexed ion beam \nimaging and cryogenic electron microscopy, to further increase the \nimaging throughput and quality in general.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nOnline content\nAny methods, additional references, Nature Portfolio reporting sum-\nmaries, source data, extended data, supplementary information, \nacknowledgements, peer review information; details of author contri-\nbutions and competing interests; and statements of data and code avail-\nability are available at https://doi.org/10.1038/s41592-024-02400-9.\nReferences\n1.\t\nSchermelleh, L. et al. Super-resolution microscopy demystified. \nNat. Cell Biol. 21, 72–84 (2019).\n2.\t\nSchermelleh, L. et al. Subdiffraction multicolor imaging of the \nnuclear periphery with 3D structured illumination microscopy. \nScience 320, 1332–1336 (2008).\n3.\t\nLawo, S., Hasegan, M., Gupta, G. D. & Pelletier, L. Subdiffraction \nimaging of centrosomes reveals higher-order organizational \nfeatures of pericentriolar material. Nat. Cell Biol. 14, 1148–1158 \n(2012).\n4.\t\nSzymborska, A. et al. Nuclear pore scaffold structure analyzed by \nsuper-resolution microscopy and particle averaging. Science 341, \n655–658 (2013).\n5.\t\nXu, K., Zhong, G. & Zhuang, X. Actin, spectrin, and associated \nproteins form a periodic cytoskeletal structure in axons. Science \n339, 452–456 (2013).\n6.\t\nMishin, A. & Lukyanov, K. Live-cell super-resolution fluorescence \nmicroscopy. Biochemistry 84, 19–31 (2019).\n7.\t\nGodin, A. G., Lounis, B. & Cognet, L. Super-resolution microscopy \napproaches for live cell imaging. Biophys. J. 107, 1777–1784 (2014).\n8.\t\nValli, J. et al. Seeing beyond the limit: a guide to choosing the \nright super-resolution microscopy technique. J. Biol. Chem. 297, \n100791 (2021).\n9.\t\nLuisier, F., Vonesch, C., Blu, T. & Unser, M. Fast interscale wavelet \ndenoising of Poisson-corrupted images. Signal Process. 90, \n415–427 (2010).\n10.\t Huang, X. et al. Fast, long-term, super-resolution imaging with \nHessian structured illumination microscopy. Nat. Biotechnol. 36, \n451–459 (2018).\n11.\t\nMandracchia, B. et al. Fast and accurate sCMOS noise  \ncorrection for fluorescence microscopy. Nat. Commun. 11,  \n94 (2020).\n12.\t Weisong, Z. et al. Sparse deconvolution improves the resolution \nof live-cell super-resolution fluorescence microscopy. Nat. \nBiotechnol. 40, 606–617 (2022).\n13.\t Belthangady, C. & Royer, L. A. Applications, promises, and pitfalls \nof deep learning for fluorescence image reconstruction. Nat. \nMethods 16, 1215–1225 (2019).\n14.\t Weigert, M. et al. Content-aware image restoration: pushing the \nlimits of fluorescence microscopy. Nat. Methods 15, 1090–1097 \n(2018).\n15.\t Chen, J. et al. Three-dimensional residual channel attention \nnetworks denoise and sharpen fluorescence microscopy image \nvolumes. Nat. Methods 18, 678–687 (2021).\n16.\t Fang, L. et al. Deep learning-based point-scanning super- \nresolution imaging. Nat. Methods 18, 406–416 (2021).\n17.\t Lehtinen, J. et al. Noise2Noise: learning image restoration without \nclean data. In Proceedings of the 35th International Conference \non Machine Learning (eds Dy, J. & Krause, A.) 2965–2974 (PMLR, \n2018).\n18.\t Batson, J. & Royer, L. Noise2self: blind denoising by self-supervision. \nIn Proceedings of the 36th International Conference on Machine \nLearning (eds Chaudhuri, K. & Salakhutdinov, R.) 524–533 (PMLR, \n2019).\n19.\t Krull, A., Buchholz, T. -O. & Jug, F. Noise2void—learning \ndenoising from single noisy images. In Proceedings of the IEEE/\nCVF Conference on Computer Vision and Pattern Recognition, \n2129–2137 (2019).\n20.\t Eom, M. et al. Statistically unbiased prediction enables accurate \ndenoising of voltage imaging data. Nat. Methods 20, 1581–1592 \n(2023).\n21.\t Zhang, G. et al. Bio-friendly long-term subcellular dynamic \nrecording by self-supervised image enhancement microscopy. \nNat. Methods 20, 1957–1970 (2023).\n22.\t Li, X. et al. Reinforcing neuron extraction and spike inference \nin calcium imaging using deep self-supervised denoising. Nat. \nMethods 18, 1395–1400 (2021).\n23.\t Lecoq, J. et al. Removing independent noise in systems \nneuroscience data using DeepInterpolation. Nat. Methods 18, \n1401–1408 (2021).\n24.\t Qiao, C. et al. Rationalized deep learning super-resolution \nmicroscopy for sustained live imaging of rapid subcellular \nprocesses. Nat. Biotechnol. 41, 367–377 (2023).\n25.\t Muller, C. B. & Enderlein, J. Image scanning microscopy. Phys. Rev. \nLett. 104, 198101 (2010).\n26.\t Hayashi, S. & Okada, Y. Ultrafast superresolution fluorescence \nimaging with spinning disk confocal microscope optics. Mol. Biol. \nCell 26, 1743–1751 (2015).\n27.\t Hell, S. W. & Wichmann, J. Breaking the diffraction resolution \nlimit by stimulated emission: stimulated-emission-depletion \nfluorescence microscopy. Opt. Lett. 19, 780–782 (1994).\n28.\t Vicidomini, G. et al. Sharper low-power STED nanoscopy by time \ngating. Nat. Methods 8, 571–573 (2011).\n29.\t Sun, D.-E. et al. Click-ExM enables expansion microscopy for all \nbiomolecules. Nat. Methods 18, 107–113 (2021).\n30.\t Dertinger, T., Colyer, R., Iyer, G., Weiss, S. & Enderlein, J. Fast, \nbackground-free, 3D super-resolution optical fluctuation imaging \n(SOFI). Proc. Natl Acad. Sci. USA 106, 22287–22292 (2009).\n31.\t Zhao, W. et al. Enhanced detection of fluorescence fluctuations \nfor high-throughput super-resolution imaging. Nat. Photonics 17, \n806–813 (2023).\n32.\t Born, M. & Wolf, E. Principles of Optics, 7th Edn (Cambridge \nUniversity Press, 1999).\n33.\t Lequyer, J., Philip, R., Sharma, A., Hsu, W. -H. & Pelletier, L. A fast \nblind zero-shot denoiser. Nat. Mach. Intell. 4, 953–963 (2022).\n34.\t Li, X. et al. Spatial redundancy transformer for self-supervised \nfluorescence image denoising. Nat. Comput. Sci. 3, 1067–1080 \n(2023).\n35.\t Chen, X. et al. Self-supervised denoising for multimodal \nstructured illumination microscopy enables long-term \nsuper-resolution live-cell imaging. PhotoniX 5, 1–22 (2024).\n36.\t Stein, S. C., Huss, A., Hähnel, D., Gregor, I. & Enderlein, J. Fourier \ninterpolation stochastic optical fluctuation imaging. Opt. Express \n23, 16154–16163 (2015).\n37.\t Ronneberger, O., Fischer, P. & Brox, T. U-net: convolutional \nnetworks for biomedical image segmentation. In International \nConference on Medical Image Computing and Computer-Assisted \nIntervention, 234–241 (2015).\n38.\t Yun, S. et al. Cutmix: regularization strategy to train strong \nclassifiers with localizable features. In Proceedings of the IEEE/\nCVF Conference on ICCV, 6023–6032 (2019).\n39.\t Nieuwenhuizen, R. P. et al. Measuring image resolution in optical \nnanoscopy. Nat. Methods 10, 557–562 (2013).\n40.\t Wang, Z., Bovik, A. C., Sheikh, H. R. & Simoncelli, E. P. Image \nquality assessment: from error visibility to structural similarity. \nIEEE Trans. Image Process. 13, 600–612 (2004).\n41.\t Lakshminarayanan, B., Pritzel, A. & Blundell, C. Simple and \nscalable predictive uncertainty estimation using deep ensembles. \nAdv. Neural Inf. Process. Syst. 30, 6402–6413 (2017).\n42.\t Prakash, M., Lalit, M., Tomancak, P., Krul, A. & Jug, F. Fully \nunsupervised probabilistic Noise2Void. In 2020 IEEE 17th \nInternational Symposium on Biomedical Imaging (ISBI), 154–158 \n(2020).\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\n43.\t Quan, Y., Chen, M., Pang, T. & Ji, H. Self2self with dropout: \nlearning self-supervised denoising from single image. In \nProceedings of the IEEE/CVF Conference on CVPR, 1890–1898 \n(2020).\n44.\t Pang, T., Zheng, H., Quan, Y. & Ji, H. Recorrupted-to-recorrupted: \nunsupervised deep learning for image denoising. In Proceedings \nof the IEEE/CVF Conference on CVPR, 2043–2052 (2021).\n45.\t Richardson, W. H. Bayesian-based iterative method of image \nrestoration. J. Opt. Soc. Am. 62, 55–59 (1972).\n46.\t Lucy, L. B. An iterative technique for the rectification of observed \ndistributions. Astron. J. 79, 745–754 (1974).\n47.\t Guo, Y. et al. Visualizing intracellular organelle and cytoskeletal \ninteractions at nanoscale resolution on millisecond timescales. \nCell 175, 1430–1442 (2018).\n48.\t Lu, M. et al. The structure and global distribution of the \nendoplasmic reticulum network are actively regulated by \nlysosomes. Sci. Adv. 6, eabc7209 (2020).\n49.\t Lu, M. et al. ERnet: a tool for the semantic segmentation and \nquantitative analysis of endoplasmic reticulum topology. Nat. \nMethods 20, 569–579 (2023).\n50.\t Sekh, A. A. et al. Physics-based machine learning for subcellular \nsegmentation in living cells. Nat. Mach. Intell. 3, 1071–1080 (2021).\n51.\t Harke, B. et al. Resolution scaling in STED microscopy. Opt. Express \n16, 4154–4162 (2008).\n52.\t Tortarolo, G. et al. Focus image scanning microscopy for sharp \nand gentle super-resolved microscopy. Nat. Commun. 13, 7723 \n(2022).\n53.\t Liu, T. et al. Multi-color live-cell STED nanoscopy of mitochondria \nwith a gentle inner membrane stain. Proc. Natl Acad. Sci. USA 119, \ne2215799119 (2022).\n54.\t Wen, G. et al. High-fidelity structured illumination microscopy  \nby point-spread-function engineering. Light Sci. Appl. 10,  \n70 (2021).\n55.\t Qiao, C. et al. Evaluation and development of deep neural \nnetworks for image super-resolution in optical microscopy.  \nNat. Methods 18, 194–202 (2021).\n56.\t Zhang, Y. et al. Mitochondria determine the sequential \npropagation of the calcium macrodomains revealed by the \nsuper-resolution calcium lantern imaging. Sci. China Life Sci. 63, \n1543–1551 (2020).\n57.\t Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J. & Maier-Hein, K. H.  \nnnU-Net: a self-configuring method for deep learning-based \nbiomedical image segmentation. Nat. Methods 18, 203–211 (2021).\n58.\t Mirza, M. & Osindero, S. Conditional generative adversarial nets. \nPreprint at https://arxiv.org/abs/1411.1784 (2014).\n59.\t Cao, H. et al. Swin-Unet: Unet-like pure transformer for medical \nimage segmentation. In European Conference on Computer \nVision, 205–218 (2022).\nPublisher’s note Springer Nature remains neutral with regard to \njurisdictional claims in published maps and institutional affiliations.\nSpringer Nature or its licensor (e.g. a society or other partner) holds \nexclusive rights to this article under a publishing agreement with \nthe author(s) or other rightsholder(s); author self-archiving of the \naccepted manuscript version of this article is solely governed by the \nterms of such publishing agreement and applicable law.\n© The Author(s), under exclusive licence to Springer Nature America, \nInc. 2024\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nMethods\nSN2N framework\nSN2N core. In principle, the N2N does not require a clean target to \ntrain a denoising model approaching supervised learning performance. \nOriginally, N2N relies on the assumption that the noise is zero-mean \nand the training set contains infinite noisy data pairs sharing identical \ncontents with independent noise components. Our SN2N upgrades \nthis to remove random noise with increased data efficiency. First, to \nremove the need for data pairs, we designed a diagonal resampling \nstrategy to generate such twin images from one single frame using \n \nthe inherent spatial redundancy of SR images. Specifically, we assem-\nbled every four pixels (as follows) in an SR image as one unit: \n \nand then we applied a special diagonal binning operation on these four \npixels to create two new pixels for the N2N data pair. We averaged the \ntwo diagonal pixels, \n and \n, as one new pixel; and the other two \ndiagonal pixels, \n and \n, as the other new pixel. Throughout the \ninput image, two 2× subsampled images were created. Second, to \nrescale back to their original image size, we adapted a Fourier interpola-\ntion method to upsample the resulting image pair by two times. Thus, \nwe transformed the two subsampled images into the Fourier domain \nand applied zero padding to extend the borders in the Fourier space \nby half of the original sizes in each direction. Then, the inverse Fourier- \ntransformed image was twice the size of the input image, identical to \nthe original SR image. Additionally, to prevent boundary artifacts, we \nextended the image (in spatial domain) with mirror-symmetric \nhalf-copies before the Fourier padding, and the resulting image was \nFourier upsampled and then cropped back to the original size.\nThird, originally in N2N, the infinite training set can execute the \naverage of noise to its true zero-mean result. However, in the case of \nlimited data for training, there would be inevitable differences between \nthese means across the different realizations of noise. Thus, the insuf-\nficient training set will introduce large predictive uncertainty, inher-\nently resulting in limited denoising performance. To minimize this \nuncertainty, we designed a self-constrained learning process while \nconsidering the identical underlying content from the noisy data pair. \nSpecifically, the noisy data pair were successively and individually fed \ninto the network, and the two resulting predictions were calculated by \nthe following loss function to execute the training stage.\n1 =\n1\n2 + λ ‖ ̃\nx1 −x2‖1 +\n1\n2 + λ ‖ ̃\nx2 −x1‖1 +\nλ\n2 + λ ‖ ̃\nx1 −\ñ\nx2‖1,\nwhere ̃\nx represents the network outcome of the input x, and ‖‖1 refers \nto the l1 norm. The first two terms are the conventional N2N loss. \nBecause the corresponding two denoised results should have no dif-\nferences, we included the last term as our self-constrained loss with \nconstraint weight λ to enforce the consistency of the predictions. 2 + λ \nis a normalization factor that ensures the gradient magnitude remains \nwithin a stable range when adjusting the weight of λ.\nPatch2Patch data augmentation. We developed a data augmentation \npipeline, namely Patch2Patch, using multidimensional random patch \ntransformations to increase the data amount, which is adaptable to \nboth 2D (xy) and 3D (xyz) datasets (Extended Data Fig. 1b). Within the \nPatch2Patch framework, there are three modes for augmentation along \nthe temporal axis, in a single image and between different experiments. \nThese designs will effectively increase variations of noise realizations \nand structural features without altering the intrinsic characteristics \nof data. Mode 1: For image/volume dataset with additional temporal \ndimension (xy-t/xyz-t), Patch2Patch defaults to create augmentation \nalong the temporal dimension. The randomly chosen patches at time \npoint A (stack A, image A) are exchanged with the patches from ran-\ndomly selected time point B (stack A, image B) at the same locations. \nMode 2: For the augmentation of a single image/volume, two distinct \npatches at different positions are randomly swapped. Mode 3: For \naugmentation between different experiments, the randomly chosen \npatches at experiment A (image A) are exchanged with the patches \nfrom randomly selected experiment B (image B) at randomly picked \nlocations. To further increase the data amounts, we also randomly \napplied one of six geometric transformations for each pair of patches: \nno transformation, horizontal flip, vertical flip, 90° rotation to the left, \n90° rotation to the right and 180° rotation.\nNetwork architecture. Because our SN2N can be applied on any \nDNNs, we simply adopted the widely recognized U-Net37 architecture \nto showcase the strengths of SN2N (Extended Data Fig. 1a). The 2D \nU-Net architecture consists of a 2D encoder module (contracting path), \na 2D decoder module (expanding path) and four skip connections \nbridging the encoder and decoder. Both the 2D encoder and decoder \nmodules are structured into four distinct blocks. Each encoder block \nis equipped with two 3 × 3 convolutional layers, followed by a leaky \nrectified linear unit and a 2 × 2 max pooling operation with a stride of \ntwo in both dimensions. Conversely, each decoder block contains two \n3 × 3 convolutional layers, followed by a leaky rectified linear unit and \na 2D bilinear interpolation. Batch normalization60 is integrated after \neach convolutional layer. The skip connections serve to concatenate \nlow-level and high-level feature maps, enhancing the preservation of \nspatial information.\nFor denoising tasks involving 3D datasets (xyz), we directly shifted \nthe U-Net from the 2D version to its 3D extension to better leverage the \naxial spatial information. With all its internal operations tailored to a \n3D framework, the 2D U-Net was transformed into the 3D U-Net. In this \nwork, we changed the convolution operations from 3 × 3 to 3 × 3 × 3, the \nmaximum pooling from 2 × 2 to 2 × 2 × 2, and the interpolation opera-\ntions from 2D bilinear to 3D trilinear interpolation.\nLearning and inference processes. Regarding the training data gen-\neration, the percentile image normalization is first applied before \ntraining to remove the baseline background and moderate the large \nintensity gap between the bright and dim fluorescence signals, which \nis defined for an image x as:\nN(x; Ilow, Ihigh) =\nx −perc(x, Ilow)\nperc(x, Ihigh) −perc(x, Ilow) ,\nwhere perc(x, I) is the I-th percentile of all pixel values of x, and Ilow and \nIhigh represent the lowest and highest values, respectively. Notably, for \nsome data under ultralow-SNR conditions, a wavelet-based background \nsubtraction12 was executed before this percentile normalization. In this \nwork, the Ilow and Ihigh were assigned as 0% and 99.999%, respectively, in \nmost applications. Specially, for the ultralow-SNR data in Fig. 3c with \nultrahigh baseline signal and a number of hot pixels, we set Ilow and Ihigh \nas 20% and 99.9%, respectively.\nFor the small dataset, we first executed the Patch2Patch pre- \naugmentation pipeline on the raw dataset to enlarge the training set \n(step. 1; Extended Data Fig. 1). Then, a sliding window approach was \nused to generate small patches (128 × 128 or 128 × 128 × 16 tiles by \ndefault) suitable as network input for training (step. 1; Extended Data \nFig. 1). The interval of the sliding window was customizable to adjust \ndifferent image processing requirements (64 pixels by default). Dur-\ning this step, a background patch rejection was performed on the fly, \nin which the patches with averaged intensity twice lower than that of \nthe entire image/volume were filtered. For these small patches, the \nspatial diagonal resampling strategy was applied to produce pairs of \ntwice smaller subimages, each sharing identical content but different \n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nnoise realizations (step. 2; Extended Data Fig. 1). After that, the Fourier \nupsampling was used to create paired SN2N data to match the original \nstructural scale. Additionally, the conventional data augmentation \nstrategies such as rotation and flipping could be included to further \nincrease the network generalization (step. 3; Extended Data Fig. 1). \nFinally, the self-constrained learning process utilized the basic U-Net \nnetwork and selected either the 2D U-Net or the 3D U-Net based on the \ninput data dimensions (step. 4; Extended Data Fig. 1).\nWe set the constraint weight λ as 1 routinely in most of our experi-\nments. The models were optimized utilizing the Adam optimizer61 with \na learning rate of 2 × 10−4, and the first-moment and second-moment \nestimates were regulated with exponential decay rates of 0.5 and 0.999, \nrespectively. For every training iteration, the batch size is set to the \nmaximum number reaching the graphics processing unit memory \nlimit, in which the training stage was conducted using NVIDIA GeForce \nRTX 4090 graphics processing units. Finally, in the inference process, \nthe raw data after the percentile normalization are directly fed into \nthe trained SN2N network for predictive analytics. For input data sizes \nthat exceed the memory limit (most of the volumetric data), the inputs \nare automatically cropped into dozens of subvolumes and fed into \nthe SN2N network, and the predictions are stitched back together to \ngenerate the final denoised results.\nIntegrations of the SN2N framework with SR reconstructions\nSN2N-enhanced RL deconvolution. We have integrated RL decon-\nvolution45,46 with our SN2N framework, namely RL-SN2N (Fig. 3a). This \napproach leverages the contrast and resolution enhancement capabili-\nties of RL deconvolution while mitigating the artifacts that arise from \ndeconvolution under low-SNR conditions. To preserve the stochastic \nindependence of noise at the pixel level, we perform RL deconvolution \nafter the initial self-supervised data generation. We used the acceler-\nated RL deconvolution:\nd j+1 = d j ⋅(hT ⋅\ng\nh⋅d j )\nv j = d j+1 −y j\nα j+1 =\n∑v j⋅v j−1\n∑v j−1⋅v j−1\ny j+1 = d j+1 + α j+1 ⋅(d j+1 −d j)\n,\nwhere y n + 1 is the image after n + 1 iterations; g is the input image; and \nh is the PSF. The d and v are the intermediate variables to help the \nacceleration process. The adaptive acceleration factor α was intro-\nduced by Biggs & Andrews62, representing the length of an iteration \nstep, which can be estimated directly from experimental results. The \nRL deconvolution was executed by the corresponding theoretically \ncalculated 2D/3D PSFs by the Gaussian kernel approximation, and the \niterations were usually selected by ~10–15 times. During the inference \nphase, the trained RL-SN2N model is fed with the directly deconvolved \nimage/volume.\nSN2N-enhanced SOFI. We integrated our SN2N solution into the SOFI \nreconstruction pipeline (SOFI-SN2N; Fig. 6a) to increase its imaging \nefficiency30. After the acquisition of the image sequence, we calculated \nthe nth order of auto-correlation cumulants, that is, the 2nd-, 3rd- and \n4th-order SOFI, on each pixel along time. Then, we performed the \nself-supervised data generation and RL deconvolution successively. \nTo match the improved spatial resolution, in SOFI-SN2N, we applied 4× \nFourier upsampling on the diagonal resampled data pair rather than \nthe routinely used 2× upsampling. Finally, after RL deconvolution, we \nused an intensity linearization step by taking the n-th root directly \nto minimize the nonlinear effects of SOFI. In the inference stage, the \ntrained SOFI-SN2N model is fed with the SR image reconstructed by the \nnth order SOFI followed by 2× Fourier interpolation before RL decon-\nvolution and intensity linearization.\nSN2N-enhanced SIM. We designed two strategies for artifact removal \nof SR-SIM images. Mostly, we first performed the self-supervised data \ngeneration to every modulated raw frame (nine frames) for creating \nthe twin SIM sequences. After that, we individually reconstructed \nthe paired SR-SIM images with the HiFi-SIM54 pipeline. Finally, these \ntwo resulting SIM images were considered as the training set of our \nSIM-SN2N. In the test stage, the ordinary SIM reconstruction was \ndirectly inputted into the SIM-SN2N network.\nBeyond that, the self-supervised data generation might influence \nthe parameter estimation, especially under ultralow-SNR conditions \nfor ultrafast imaging, leading to failed reconstruction. On the other \nhand, the SIM reconstruction brings the inherent artificial upsampling \n(twofold), and hence the adjacent pixels are highly correlated. To create \nindependent data pairs from SIM reconstruction, we performed the \nresampling step within 3 × 3 pixels (Extended Data Fig. 8). Specifically, \nwe assembled every nine pixels as one unit: \nand then we applied a special binning operation on these nine pixels \nto create two new pixels for the N2N data pair. We averaged the four \ncorner pixels, \n, \n, \n and \n, as one new pixel; and the other four \nintermediate pixels, \n, \n, \n and \n, as the other new pixel. Through-\nout the input image, two 3× subsampled images were created. Followed \nby a 3× Fourier upsampling, we finalized the required data pair directly \nfrom an SR-SIM image.\nSN2N-empowered automated subcellular segmentation  \nand tracking\nSubcellular segmentation. We used the Otsu63 method to auto-\nmatically determine the hard thresholds for identifying the corre-\nsponding organelle features in images. Additionally, to achieve more \nprecise segmentation, we also utilized pretrained models from several \nlearning-based approaches, for example, ERnet49 for ER structures \nand Mitonet50 for Mito shapes. Following the segmentation of ER, we \neliminated isolated pixels in the binarized masks and then extracted the \nskeleton structures for the network topology construction, in which \nthe nodes represent intersections in the skeleton graph, and edges \nrepresent connections between these nodes. After Mito segmenta-\ntion, we computed the connected domains within the binary masks \nand identified the skeletons and key points. These key points were \nsubsequently categorized into junctions or end points based on their \nrespective topological positions (Extended Data Fig. 5d).\nTracking of Lys. Tracking of Lys was performed using TrackMate \n(7.11.1)64. To characterize their dynamic behaviors, we computed the \nMSD for the trajectories across all time points65. The calculated MSD \ncurves are approximated with a power-law function65:\nMSD(t) = Γ × t α,\nwhere t represents the time interval, Γ is the proportionality factor that \nrelates to both particle motion dynamics and the physical properties of \nthe system, and α characterizes the different modes of particle move-\nments66. Then, a logarithmic transformation is applied to the MSD \nformula, followed by a linear regression to estimate α:\nlog(MSD) = α × log(t) + log(Γ ).\nIn our analysis, the Lys movements were classified into confined \nmotion (α < 0.85), free diffusion (0.85 ≤ α ≤ 1.2) and directed move-\nment (α > 1.2)66.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nMito diameter. To estimate the Mito diameter at a selected point, we \ndrew the tangent line of the nearest Mito contour. Using the perpen-\ndicular line of this tangent line, we obtained the other intersected point \nto the opposite Mito contour. Finally, we approximated the diameter \nby measuring the distance between the initially selected point and this \nintersection point67.\nER–Lys distance. We identified the edges of ER tubules from the \nERnet-generated masks, and the distances were calculated based on \nthe centroid coordinates of Lys and the most adjacent ER edges. Then, \nthe final ER–Lys distances were quantified by averaging the distances \nacross the entire trajectory.\nER–Mito distance. The ER–Mito contact level was quantified by the \nminimum distance from each Mito to the edges of ER tubules, in which \nthe distances of 0 were distributed as ‘contacted’ and larger than 0 as \n‘not contacted’. The Mito diameter was estimated at the point with the \nclosest ER–Mito distance.\nLys–Mito interactions. We calculated the MOC to examine the Lys–\nMito interactions. The masks of Mito and Lys were convolved with the \nPSF of the SD-SIM system, and the MOC values were calculated from \nthe resulting masks at each time point. Subsequently, we calculated \nthe ratio of the mean value and STD of the MOC curve and obtained \nan empirical threshold of 0.26, in which the associated MOC values \nlarger than this threshold were identified as functional sites in spe-\ncific instances of Lys–Mito contact24 (Extended Data Fig. 5j). With \nMOC values below the threshold, intersection points on the Mito edge \nwere identified from the connecting line of Lys and Mito centroids. \nWe then used it as the selected point for the Mito diameter estima-\ntion. For events of MOC values higher than the threshold, we drew the \nperpendicular line of the connecting line from the two intersection \npoints between the edges of Lys and Mito, and the intersection of this \nperpendicular line and the Mito edge was picked as the selected point \nfor the Mito diameter estimation.\nPerformance metrics\nSimulations of microtubule filaments. To perform benchmarks \nwith simulated ground truth, we created the microtubule-like struc-\ntures. Specifically, we utilized the ‘insertShape’ (MATLAB function) \nto sketch multiple lines with random orientations on a blank canvas \nof 4,096 × 4,096 pixels (Supplementary Fig. 1). To simulate the effect \nof incomplete labeling observed in practical experiments, we applied \na small Gaussian mask (σ as 2 pixels), introducing random notches \nalong the lines (indicated by red circles in Supplementary Fig. 1). Then, \nto mimic the curviness of cytoskeleton, an elastic deformation was \napplied to bend the straight lines to curved lines in both x and y dimen-\nsions, in which the resulting image has a pixel size of 16.25 nm. The syn-\nthetic structures were convolved with a 150-nm PSF and downsampled \nby twofold, for 2,048 × 2,048 pixels with a 32.5-nm pixel size (simulated \nblurred ground truth). Finally, we assigned a number of photons to the \nsynthetic structures with intensity emissions of 100 a.u. (level 3), 50 a.u. \n(level 2), 25 a.u. (level 1), 19 a.u., 13 a.u., 7 a.u. and 1 a.u. Then, the Poisson \nnoise injection is followed by an addition of the Gaussian readout noise:\ny = Poisson(photon × x) + n,\nwhere y is the final noisy image, and x is the downsampled image. ‘Pois-\nson’ denotes the Poisson noise injection; ‘photon’ is the photon num-\nber; and n represents the Gaussian noise with a fixed variance value.\nPixel-wise metrics. In this work, the SSIM40, PSNR and RMSE were \nused as metrics to evaluate the pixel-level consistency between recon-\nstructed images and GT images. For the fixed samples, we directly \nenlarged the exposure time to acquire the high-SNR data. To remove \npotential small baseline background and noise, the GT images were \ncreated from these high-SNR images by subtracting a constant back-\nground value and subsequently filtering a small Gaussian kernel. To \nqualify data and model uncertainties, we adopted the STD using the ten \npredictions from ten repetitively collected inputs or ten repetitively \ntrained models:\nSTD = ∑\nx,y\n√\n√\n√\n∑\nn\ni=1 (Ii −⟨I ⟩1∼n)\n2\nn −1\n,\nwhere n is the sequence length (default as 10); Ii(x, y) represents the \nintensity of the ith image in the sequence, and ⟨I(x, y)⟩1∼n is the averaged \nintensity of the sequence.\nFRC resolution. The calculation of FRC resolution requires two inde-\npendent frames of identical contents under the same imaging condi-\ntions39. In case of confocal, SD-SIM and STED imaging, we repetitively \nacquired the same content twice. For SOFI, these two frames were \ngenerated by splitting the raw image sequence into two image subsets, \nfor example, the first 20 frames and the last 20 frames, and reconstruct-\ning them independently.\nLRQ. To quantitatively evaluate the reconstruction quality of parallel \nlines, we used the LRQ12 metric:\nLRQ =\n2 × avg(region0)\navg(region1) + avg(region2) .\nThe ‘region1’, ‘region2’ and ‘region0’ represent the parallel lines and \nthe region in between, and ‘avg’ indicates the mean intensity of pixels \nwithin the corresponding area. To avoid overconfident determination, \nwe used a threshold of 0.2 to ascertain the successful separation of \nparallel lines.\nPrediction uncertainty estimation. An optimized denoising algorithm \ncan be characterized by minimal data and model uncertainties. Data \nuncertainty indicates the algorithm’s ability to efficiently reduce noise \n(minimize errors), whereas model uncertainty reflects the model’s \nefficiency in utilizing data (sustain performance under limited or \nvarying data conditions)68. To estimate data uncertainty, we collected \nten independent frames of identical contents under the same imag-\ning conditions and fed them to the trained SN2N network. The STD of \nthe ten resulting predictions was calculated as the data uncertainty. \nRegarding the model uncertainty, we repetitively trained the DNNs \nten times and inputted the same data into these ten models. The STD \nof the resulting predictions served as a measure of model uncertainty.\nFWHM measurements. The FWHM values were estimated from the \nGaussian fittings of the manually picked intensity profiles. Particularly, \nthe profiles and values plotted in Extended Data Fig. 4a were auto-\nmatically created by using the LuckyProfiler ImageJ plugin69, which \ncan autonomously identify and quantify the optimal FWHM locations \nwithin images. It enabled us to select the necessary regions for FWHM \ncalculations and apply Gaussian fitting algorithms.\nCompared methods\nSupervised learning. For the fixed-sample experiments, we man-\nually created the GT by enlarging the exposure time to acquire the \nhigh-SNR data, which served as labels. To remove potential low baseline \nbackground and noise, the GT images were then created from these \nhigh-SNR images by subtracting a constant background value and \nsubsequently filtering a small Gaussian kernel. The network configura-\ntions and training hyperparameters of the ‘supervised learning’ used \nin this work were identical to those of SN2N.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nOther self-supervised methods. We evaluated the denoising perfor-\nmance of the simulation dataset (Fig. 1b and Supplementary Fig. 5a) and \nmulticolor SD-SIM live-cell imaging dataset (Supplementary Fig. 6a) \nagainst five and eight self-supervised denoisers, respectively. The \ntested denoisers include N2V19, PPN2V42, S2S43, R2R44, N2F33, DeepCAD22, \nDeepSeMi21 and SRDTrans34. In particular, for the ultralow-SNR SD-SIM \nlive-cell data, we applied percentile normalization before the data gen-\neration step to remove the strong baseline background, with the nor-\nmalization bounds set to a minimum of 20% and a maximum of 99.9%.\nN2V. To control variables and enhance the N2V denoising effect, we \nutilized the data generation code from N2V with a neighborhood radius \nof 5, combined with the SN2N network architecture. We kept other train-\ning settings consistent with SN2N, featuring a learning rate of 2 × 10–4, \n50 epochs, a batch size of 32 and a patch size of 128.\nPPN2V. For PPN2V, we used the default settings as published: the noise \nmodel set to the Gaussian mixture model, a learning rate of 1 × 10–5, 200 \nepochs, 50 iterations per epoch, a batch size of 4 and a patch size of 100.\nS2S. For S2S, we used the default published settings of 150,000 itera-\ntions and a learning rate of 1 × 10–4. According to the final predicted \nresults, the optimal iteration numbers for the simulation dataset and \nSD-SIM live-cell dataset were 200 and 2,000, respectively.\nR2R. For R2R, we used the published settings of a 1 × 10–3 learning rate \nand 50 epochs. We optimized the denoising performance by reducing \nthe batch size from the default 128 to 32. For the simulated microtubule \ndataset, the noise level (noise STD) is set as 25, and for SD-SIM live-cell \ndata experiments, it is set as 50.\nDeepCAD. For DeepCAD, the only modification was to increase the \nbatch size from 1 to 32. The learning rate was set as 1 × 10–3, and the \nnumber of iterations was set at 50 by default.\nN2F. For N2F, the learning rate is set as 1 × 10–3. The network training \nuses a custom loop that includes an early stopping mechanism, which \nhalts training when there is no further improvement in PSNR over a \ncertain period.\nDeepSemi. For DeepSemi, we use the default published settings of a \n1 × 10−4 learning rate, 100 epochs and a batch size of 2.\nSRDTrans. For SRDTrans, we adhered to the published settings of a \n1 × 10−4 learning rate, 30 epochs and a batch size of 2.\nSD-SIM setup\nWe used two commercial SD-SIM systems to capture SR confocal images. \nWe validated the denoising performance using a commercial fluorescent \nsample (the Argo-SIM slide, Argolight) with GT patterns consisting of \nfluorescing double-line pairs (spacing from 0 nm to 390 nm, λex = 300–\n550 nm; http://argolight.com/products/argo-sim/) under the SpinSR10 \nsystem and conducted live-cell imaging tests using the Live-SR system.\nSpinSR10 system. The SpinSR10 system is a commercial SD-SIM \nsystem (SpinSR10, Olympus) equipped with a wide-field objective \n(×100/1.49 oil, APON, Olympus) and an sCMOS camera (ORCA Fusion, \nHamamatsu). Four laser beams of 405 nm, 488 nm, 561 nm and 640 nm \nwere combined with the SD-SIM. The detection optical path adopted a \nfurther ×3.2 magnification, and the total magnification was ×320. We \ncaptured SD-SIM images in its SoRa (SR) mode and collected confocal \nimages by switching to its conventional SD-confocal mode.\nLive-SR system. The Live-SR SD-SIM system is based on an inverted \nfluorescence microscope (IX81, Olympus) equipped with a wide-field \nobjective (×100/1.3 oil, Olympus), a scanning confocal system (CSU-X1, \nYokogawa) and a Live-SR module (GATACA Systems). Four laser beams \nof 405 nm, 488 nm, 561 nm and 647 nm were combined with the SD-SIM. \nThe images were captured by either an sCMOS camera (C14440-20UP, \nHamamatsu) or an EMCCD camera (iXon3 897, Andor).\nSTED setup\nWe acquired STED images from two commercial STED systems.\nAbberior STED. Long-term activities of mitochondrial cristae (PKMO53 \nlabeled) were recorded by a commercial STED microscope (STEDYCON, \nAbberior Instruments) equipped with a wide-field objective (×100/1.45, \nCFI Plan Apochromat Lambda D, Nikon). PKMO was excited at a wave-\nlength of 561 nm, and STED was performed using a pulsed depletion \nlaser at a wavelength of 775 nm with gating of 1 ns to 7 ns and dwell times \nof 10 μs. A pixel size of 25 nm was used for STED recording and each \nline was scanned one or ten times (line accumulations). The pinhole \nwas set to 0.7 to 1.0 Airy units.\nLeica STED. Other live-cell STED images were obtained using a gated \nSTED microscope (TCS SP8 STED 3X, Leica) equipped with a wide-field \nobjective (×100/1.40 oil, HCX PL APO, Leica). The excitation and deple-\ntion wavelengths were 488 nm and 592 nm for the Sec61β-GFP and \nLifeAct-GFP, 594 nm and 775 nm for the Alexa Fluor 594, 635 nm and \n775 nm for the Alexa Fluor 647, and 651 nm and 775 nm for SiR-tubulin. \nThe detection wavelength range was set to 495–571 nm for GFP, 605–\n660 nm for Alexa Fluor 594, 657–750 nm for SiR and 649–701 nm for \nAlexa Fluor 647. For comparison, confocal images were acquired in the \nsame field before the STED imaging. All images were obtained using \nthe LAS AF software (Leica).\nSOFI setup\nWide-field microscopy. The three phases of structured illumination \nunder the same orientation can be averaged to a uniform wide-field \nillumination. Taking advantage of that, we use the SIM setup described \nabove to generate the wide-field images by integrating three frames \n(corresponding to three phases of structured illumination) on the \ncamera plane, which enables more flexible cross-validation of SIM and \nSOFI-SN2N results. In other words, we use the identical commercial \ninverted fluorescence microscope (IX83, Olympus) equipped with \nan objective (×100/1.7 HI oil, APON, Olympus) and an sCMOS (Flash \n4.0 V3, Hamamatsu) camera to capture the wide-field images for our \nSOFI-SN2N reconstruction.\nSD-confocal microscopy. A commercial SD-confocal microscope \nsystem (Dragonfly SD system, Andor) based on an inverted fluorescence \nmicroscope (DMi8, Leica) with a wide-field objective (×100/1.3 oil, Plan \nApo, Leica) is used in this work. Four laser beams of 405 nm, 488 nm, \n561 nm and 647 nm were combined with the SD-confocal microscope. \nThe images were captured by an sCMOS camera (Zyla 4.2 Plus, Andor).\nSIM setup\nThe SIM system is based on a commercial inverted fluores-\ncence microscope (IX83, Olympus) equipped with an objective \n(×100/1.49 oil, UAPON, Olympus, for 2D-SIM; ×100/1.7 HI oil, APON, \nOlympus, for TIRF-SIM) and a multiband dichroic mirror (DM, \nZT405/488/561/640-phase R; Chroma) as described previously34. In \nshort, laser light with wavelengths of 488 nm (Sapphire 488LP-200) \nand 561 nm (Sapphire 561LP-200, Coherent) and acoustic optical tun-\nable filters (AOTF, AA Opto-Electronic, France) were used to combine, \nswitch and adjust the illumination power of the lasers. A collimating \nlens (focal length of 10 mm, Lightpath) was used to couple the lasers \nto a polarization-maintaining, single-mode fiber (QPMJ-3AF3S, Oz \nOptics). The output lasers were then collimated by an objective lens \n(CFI Plan Apochromat Lambda ×2 NA 0.10, Nikon) and diffracted by \n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nthe pure phase grating that consisted of a polarizing beam splitter, \na half-wave plate, and the SLM (3DM-SXGA, ForthDD). The diffrac-\ntion beams were then focused by another achromatic lens (AC508-\n250, Thorlabs) onto the intermediate pupil plane, where a carefully \ndesigned stop mask was placed to block the zero-order beam and other \nstray light and to permit passage of ±1 ordered beam pairs only. To max-\nimally modulate the illumination pattern while eliminating the switch-\ning time between different excitation polarizations, a home-made \npolarization rotator was placed after the stop mask34. Next, the light \npassed through another lens (AC254-125, Thorlabs) and a tube lens \n(ITL200, Thorlabs) to focus on the back focal plane of the objective \nlens, which interfered with the image plane after passing through the \nobjective lens. Emitted fluorescence collected by the same objective \npassed through a dichroic mirror, an emission filter and another tube \nlens. Finally, the emitted fluorescence was split by an image splitter \n(W-VIEW GEMINI, Hamamatsu) before being captured by an sCMOS \n(Flash 4.0 V3, Hamamatsu) camera.\nExM setup\nWe used the commercial inverted fluorescence microscope (IX83, \nOlympus) equipped with a wide-field objective (×100/1.49 oil, UAPON, \nOlympus) and an sCMOS (Flash 4.0 V3, Hamamatsu) camera to capture \nthe wide-field images of the expanded samples.\nImaging sample preparation\nCell maintenance and preparation. COS-7 cells (American Type \nCulture Collection, CRL-1651) and HeLa cells (American Type Cul-\nture Collection, CCL-2) were cultured in high-glucose DMEM (Gibco, \n21063029) supplemented with 10% FBS (Gibco) and 1% 100 mM sodium \npyruvate solution (Sigma-Aldrich, S8636) in an incubator at 37 °C \nwith 5% CO2 until ~75% confluency was reached. MCF7 cells were cul-\ntured in MEM (Thermo Fisher, 11095072) supplemented with 10% FBS, \n0.01 mg ml−1 human recombinant insulin (Sigma, I9278), and 1% 100 mM \nsodium pyruvate solution. For the SD-SIM/SD-confocal/STED imag-\ning experiments, 35-mm glass-bottomed dishes (Cellvis, D35-14-1-N) \nwere used. For the wide-field and 2D-SIM imaging experiments, cells \nwere seeded onto coverslips (H-LAF 10 L glass; reflection index, 1.788; \ndiameter, 26 mm; thickness, 0.15 mm; customized) coated with 0.01% \npoly-l-lysine solution (Sigma, P4707) for 10 min and washed twice with \nsterile water before seeding transfected cells.\nLive-cell samples for SD-SIM and SIM. To label late endosomes or \nlysosomes, we incubated COS-7 cells in 50 nM LysoTracker Deep Red \n(Thermo Fisher Scientific, L12492) for 45 min and washed them three \ntimes in PBS before imaging. To label mitochondria, COS-7 cells were \nincubated with 250 nM MitoTracker Green FM (Thermo Fisher Scien-\ntific, M7514) and 250 nM MitoTracker Deep Red FM (Thermo Fisher \nScientific, M22426) in HBSS containing Ca2+ and Mg2+ or no phenol red \nmedium (Thermo Fisher Scientific, 14025076) at 37 °C for 15 min before \nbeing washed three times before imaging. To perform nuclear staining \non COS-7 cells, SPY650-DNA (Cytoskeleton, CY-SC501) was diluted to \n1:1,000 in PBS for ~1 h and washed three times in PBS.\nTo label cells with genetic indicators, COS-7 cells were transfected \nwith LifeAct-EGFP/LAMP1-EGFP/LAMP1–mCherry/Tom20–mCherry/\nSec61β-EGFP/Golgi-BFP/mGold-Mito-N-7/DsRed-ER70. The transfec-\ntions were executed using Lipofectamine 2000 (Thermo Fisher Sci-\nentific, 11668019) according to the manufacturer’s instructions. After \ntransfection, cells were plated in precoated coverslips. Live cells were \nimaged in a complete cell culture medium containing no phenol red \nin a 37 °C live-cell imaging system. For the calcium lantern imaging in \nSD-SIM, the calcium signal was stimulated with a micropipette contain-\ning 10 μmol l−1 5′-ATP-Na2 solutions (Sigma-Aldrich, A1852)56.\nSamples for STED imaging. To label the ER-tubule/actin/microtubule \nin live cells, COS-7 cells were either transfected with Sec61β-EGFP/\nLifeAct-EGFP, or incubated with SiR-Tubulin (Cytoskeleton, CY-SC002) \nor PKMO53 for ~20 min without washing before imaging.\nImmunofluorescence for SOFI. The COS-7 cells were grown in 35-mm \nglass-bottomed dishes overnight and rinsed with PBS, then immedi-\nately fixed with prewarmed 4% paraformaldehyde (Santa Cruz Biotech-\nnology, sc-281692) for 10 min. After three washes with PBS, cells were \npermeabilized with 0.1% Triton X-100 (Sigma-Aldrich, X-100) in PBS for \n10 min. Cells were blocked in 5% BSA/PBS for 1 h at room temperature \n(RT). Mouse anti-Tubulin DM1a (Sigma, T6199) was diluted to 1:100 and \nstained cells in 2.5% BSA/PBS blocking solution for 2 h at RT. The cells \nwere then washed with PBS five times for 10 min per wash and stained \nwith biotin-XX goat anti-mouse IgG antibody (Invitrogen, B2763). The \ncells were then washed with PBS five times for 10 min per wash and \nstained with QD525 streptavidin conjugate (Invitrogen, Q10143MP) \nfor 60 min. Finally, cells were washed five times with PBS and imaged.\nLive-cell samples for SOFI. To label the mitochondria in live cells, \nCOS-7 cells were transfected with Skylan-S-TOM20 (ref. 71). The trans-\nfections were executed using Lipofectamine 2000 (Thermo Fisher \nScientific, 11668019) according to the manufacturer’s instructions. \nAfter transfection, cells were plated in glass-bottomed dishes. The \nSkylan-S was under sequential illumination with a 405-nm laser (low \npower) when imaging. In addition, live cells were imaged in complete \ncell culture medium containing no phenol red in a 37 °C live-cell \nimaging system.\nSample preparation for ExM\nSample expansion. The sample expansion was performed as previ-\nously described29,72. The labeled cells were incubated with 0.1 mg ml−1 \nof Acryloyl-X (AcX, Thermo, A20770) diluted in PBS overnight at RT \nand washed three times with PBS. To prepare the gelation solution, \nfreshly prepared 10% (wt/wt) N,N,N′,N′-tetramethylethylenediamine \n(Sigma, T7024) and 10% (wt/wt) ammonium persulfate (Sigma, \nA3678) were added to the monomer solution (1× PBS, 2 M sodium \nchloride, 2.5% (wt/vol) acrylamide (Sigma, A9099), 0.15% (wt/vol) \nN,N′-methylenebisacrylamide (Sigma, M7279) and 8.625% (wt/vol) \nsodium acrylate (Sigma, 408220)) to a final concentration of 0.2% \n(wt/wt) each. Next, the cells were embedded with the gelation solu-\ntion first for 5 min at 4 °C, and then for 1 h at 37 °C in a humidified \nincubator. The gels were immersed into the digestion buffer (50 mM \nTris, 1 mM EDTA, 0.1% (vol/vol) Triton X-100, and 0.8 M guanidine HCl, \npH 8.0) containing 8 units per ml proteinase K (NEB, P8107S) at 37 °C \nfor 4 h, and then placed into double-distilled water to expand. Water \nwas changed 4–5 times until the expansion process reached a plateau. \n \nBy determining the gel sizes of before and after the expansion, we \nquantified the expansion factor to be 4.5 times. The gels were immo-\nbilized on poly-d-lysine-coated cover glass with a thickness of no. 1.5 \nfor further imaging.\nα-tubulin immunostaining. COS-7 cells were seeded in a Lab-Tek \nII chamber slide (Nunc, 154534). Cells were firstly extracted in the \ncytoskeleton extraction buffer (0.2% (vol/vol) Triton X-100, 0.1 M PIPES, \n1 mM EGTA, and 1 mM magnesium chloride, pH 7.0) for 1 min at RT. Next, \nthe extracted cells were fixed with 3% (w/vol) formaldehyde and 0.1% \n(vol/vol) glutaraldehyde for 15 min, reduced with 0.1% (wt/vol) sodium \nborohydride in PBS for 7 min, and washed three times with 100 mM \nglycine. Then, the cells were permeabilized with 0.1% (vol/vol) Triton \nX-100 for 15 min, and blocked with 5% (wt/vol) BSA in 0.1% (vol/vol) \nTween 20 for 30 min. For antibody staining, the cells were incubated \nwith monoclonal rabbit anti-α-tubulin antibody (EP1332Y, 1:250 dilu-\ntion, Abcam, ab52866) in antibody dilution buffer (2.5% (wt/vol) BSA \nin 0.1% (vol/vol) Tween 20) overnight at 4 °C, washed three times with \n0.1% (vol/vol) Tween 20, incubated with Alexa Fluor 488-conjugated \nF(ab')2-goat anti-rabbit secondary antibody (1:1,000 dilution, Thermo, \n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nA11070) in antibody dilution buffer for 2 h at RT and washed three times \nwith 0.1% (vol/vol) Tween 20.\nSec61β-GFP transfection. COS-7 cells were seeded in a Lab-Tek II cham-\nber slide (Nunc, 154534) and cultured to reach around 50% confluence. \nFor transient transfection of Sec61β-GFP in a single well, 500 ng plasmid \nand 1 μl of X-tremeGENE HP (Roche) were diluted in 20 μl Opti-MEM \nsequentially. The mixture was vortexed, incubated for 15 min at RT and \napplied to cells. Twenty-four hours after transfection, the cells were \nwashed three times with PBS and fixed as described in the α-tubulin \nimmunostaining experiment.\nOpen-sourced datasets\nSTED images with different excitation/depletion laser powers. We \nused the published dataset from ref. 52 for testing our SN2N’s denoising \nperformance on STED images. A custom STED system incorporating a \nSPAD array detector with 5 × 5 elements was used for data collection. A \nseries of images (Fig. 5a) were collected from living HeLa cells labeled \nwith SIR-tubulin under gradually increased STED power (0%, 5%, 8%, 12%, \n16%, 19%, 30%, 50% and 80% depletion laser powers). Another series of \nimages (Supplementary Fig. 7) were acquired from α-tubulin-labeled \nfixed HeLa cells with gradually increased STED power (0%, 10%, 20%, \n25%, 30%, 40% and 90% depletion laser powers). With 0% depletion laser \npower, the system was switched to a conventional point-scanning confo-\ncal mode. We used the direct sum and the reassigned sum of signals from \n25 elements as confocal/STED and ISM/STED-ISM images, respectively.\nBioSR SIM dataset. We used the open-sourced dataset, the BioSR \ndataset from ref. 55, for evaluating the denoising performance on SIM \nimages. The clathrin-coated pits, microtubules and ER data under \ndifferent noise levels were used as SIM images of high-, medium- and \nlow-SNR conditions in this work.\nImage rendering and processing\nWe used the ‘biop-12colors’ color map to color code the 3D volumes in \nFig. 4b,e,f, Extended Data Figs. 4g and 10a–e and Supplementary Figs. 10 \nand 11c. The 3D volumes in Fig. 4g–i and Extended Data Fig. 6a,c,e were \nrendered using the Microscape software (https://www.microscape.\nxyz/). All data processing was achieved using Python scripts, MATLAB \nand Fiji/ImageJ. All figures were prepared with MATLAB, Fiji/ImageJ, \nMicrosoft Visio and OriginPro.\nReporting summary\nFurther information on research design is available in the Nature \nPortfolio Reporting Summary linked to this article.\nData availability\nWe provided two representative datasets from Figs. 1 and 3b available \nat https://github.com/WeisongZhao/SN2N/. All other data that sup-\nport the findings of this study are available from the corresponding \nauthor upon request.\nCode availability\nVideos were produced with Microsoft PowerPoint and our lightweight \nMATLAB framework, which is available at https://github.com/Wei-\nsongZhao/img2vid/. The percentile normalization method has been \nwritten as a Fiji/ImageJ plugin and can be found at https://github.com/\nWeisongZhao/percentile_normalization.imagej/. The tutorials and \nthe updated version of our SN2N can be found at https://github.com/\nWeisongZhao/SN2N/.\nReferences\n60.\t Ioffe, S. & Szegedy, C. Batch normalization: accelerating \ndeep network training by reducing internal covariate shift. In \nInternational Conference on Machine Learning, 448–456 (2015).\n61.\t Kingma, D. P. & Ba, J. Adam: A method for stochastic optimization. \nPreprint at https://arxiv.org/abs/1412.6980 (2014).\n62.\t Biggs, D. S. & Andrews, M. Acceleration of iterative image \nrestoration algorithms. Appl. Opt. 36, 1766–1775 (1997).\n63.\t Otsu, N. A threshold selection method from gray-level \nhistograms. IEEE Trans. Syst. 9, 62–66 (1979).\n64.\t Ershov, D. et al. TrackMate 7: integrating state-of-the-art \nsegmentation algorithms into tracking pipelines. Nat. Methods 19, \n829–832 (2022).\n65.\t Qian, H., Sheetz, M. P. & Elson, E. L. Single particle tracking. \nAnalysis of diffusion and flow in two-dimensional systems. \nBiophys. J. 60, 910–921 (1991).\n66.\t Ba, Q., Raghavan, G., Kiselyov, K. & Yang, G. Whole-cell scale \ndynamic organization of lysosomes revealed by spatial statistical \nanalysis. Cell Rep. 23, 3591–3606 (2018).\n67.\t Damenti, M., Coceano, G., Pennacchietti, F., Boden, A. & Testa, I.  \nSTED and parallelized RESOLFT optical nanoscopy of the tubular  \nendoplasmic reticulum and its mitochondrial contacts in neuronal \ncells. Neurobiol. Dis. 155, 105361 (2021).\n68.\t Zhao, W. et al. Quantitatively mapping local quality of \nsuper-resolution microscopy by rolling Fourier ring correlation. \nLight Sci. Appl. 12, 298 (2023).\n69.\t Li, M. et al. LuckyProfiler: an ImageJ plug-in capable of \nquantifying FWHM resolution easily and effectively for \nsuper-resolution images. Biomed. Opt. Express 13, 4310–4325 \n(2022).\n70.\t Lee, J. et al. Versatile phenotype-activated cell sorting. Sci. Adv. 6, \neabb7438 (2020).\n71.\t Zhang, X. et al. Development of a reversibly switchable \nfluorescent protein for super-resolution optical fluctuation \nimaging (SOFI). ACS Nano 9, 2659–2667 (2015).\n72.\t Tillberg, P. et al. Protein-retention expansion microscopy of cells \nand tissues labeled using standard fluorescent proteins and \nantibodies. Nat. Biotechnol. 34, 987–992 (2016).\nAcknowledgements\nWe thank the assistance of T. Liu from Z. Chen’s laboratory at the \nPeking university for STED imaging of PKMO-labeled mitochondrial \ncristae. This work was supported by the National Key Research and \nDevelopment Program of China (grant no. 2022YFC3400600 to \nL.C.), the National Natural Science Foundation of China (grant nos. \n32422052 to W.Z., 62305083 to W.Z., T2222009 to H.L., 32227802 to \nL.C., 21927813 to L.C., 81925022 to L.C., 92054301 to L.C., 32301257 \nto S.Z., 32071458 to H. M.), the Young Elite Scientists Sponsorship \nProgram by China Association for Science and Technology (grant \nno. 2023QNRC001 to W.Z.), and the Heilongjiang Provincial \nPostdoctoral Science Foundation (grant no. LBH-Z22027 to W.Z.), \nthe Natural Science Foundation of Heilongjiang Province (grant no. \nYQ2021F013 to H.L.), the Beijing Natural Science Foundation (grant \nno. Z20J00059 to L.C.), the Nanyang Assistant Professorship Start-up \nGrant, and National Research Foundation of Singapore (grant no. \nNRF-CRP29-2022-0003 to G.H.) and the Guangdong Basic and \nApplied Basic Research Foundation (grant no. 2022A1515011683  \nto J.H.). L.C. acknowledges support by the High-performance \nComputing Platform of Peking University.\nAuthor contributions\nW.Z. conceived the research; L.Q. implemented the corresponding \nsoftware; S.Z., X.Y. and K.W. performed the experiments and collected \nthe data; Q.L. analyzed the data and prepared the figures; L.Q. and \nY.H. prepared the videos; X.L., H.M., G.H., W.C., C.G., J.H., J.T., H.L. \nand L.C. participated in discussions during the development of the \nmanuscript; W.Z. and L.Q. wrote the manuscript with input from \nall authors; W.Z., H.L. and L.C. supervised the project. All authors \nparticipated in the discussions and data interpretation.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nCompeting interests\nL.C., H.L., W.Z. and L.Q. have a pending patent application on the presented \nframework. The remaining authors declare no competing interests.\nAdditional information\nExtended data is available for this paper at  \nhttps://doi.org/10.1038/s41592-024-02400-9.\nSupplementary information The online version contains supplementary \nmaterial available at https://doi.org/10.1038/s41592-024-02400-9.\nCorrespondence and requests for materials should be addressed to \nWeisong Zhao.\nPeer review information Nature Methods thanks  \nLaurence Pelletier, Yide Zhang, Jiji Chen and the other,  \nanonymous, reviewers for their contribution to the peer  \nreview of this work.\nReprints and permissions information is available at  \nwww.nature.com/reprints.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 1 | See next page for caption.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 1 | Workflow of SN2N and network architectures.  \na, Detailed flow diagram of SN2N (Methods). Steps 1-3, self-supervised data \ngeneration. First, the data pre-augmentation (optional) is performed using  \nthe Patch2Patch (random patch transformations in multiple dimensions) \nstrategy. After that, a sliding window approach is employed to generate small \npatches suitable for input into the network for training. Subsequently, the  \nspatial diagonal resampling strategy followed by Fourier upsampling is used  \nto create paired SN2N data. Additionally, basic augmentations such as rotation \nand flipping (optional) are applied to the generated data pairs. Step 4:  \nself-constrained learning process. SN2N utilizes the classical U-Net network \nand selects either the 2D U-Net or 3D U-Net based on the input data dimensions. \nThe generated paired images are considered as one training example, and the \nresulting two predictions are used to calculate the loss for back propagation.  \nb, Patch2Patch (P2P) pipeline (Methods). It includes three available modes  \nfor augmentation along the temporal axis, in a single image, and between \ndifferent experiments.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 2 | See next page for caption.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 2 | Systematical tests of different components of SN2N. \na, SN2N denoising results under different pixel sizes with same resolution. From \ntop to bottom: Raw images, SN2N results, and clean ground truth images. The \nsynthetic structures (16.25 nm pixel size) were convolved with a 150 nm size \nPSF and downsampled by 2, 3, 4, 5, 6, 8, and 16 times. SSIM and PSNR values of \nSN2N results are marked on the bottom right corner. b, SSIM (top) and PSNR \n(bottom) values from data under different downsampling rates. c, Ablation \ntests for individual components of SN2N. From left to right: Raw (top)/ground \ntruth (bottom), without downsampling process (Raw2Raw, the same noisy \nimage as both input and label), with downsampling only, including the Fourier \nup-sampling step, integrating the self-constrained learning process, and \nsupplementing our Patch2Patch augmentation. We found that the network \nwas not able to execute denoising without the downsampling step, and the \nupsampling step enforced the consistency of the predicted structural scale. \nThe self-constrained learning process strengthened the data-efficiency and \nperformance, and the addition of Patch2Patch further maximized the data \nefficiency. d, SSIM values of different components in SN2N. e, SN2N denoising \nresults under different interpolation methods. From left to right: Raw input, \nSN2N results using data without interpolation, with bilinear interpolation, and \nour Fourier interpolation as training sets, and ground truth image. f, SSIM values \nof SN2N under different interpolation strategies. g, SN2N results under different \nself-constrained regularization weights (values labeled on the top left corners). \nh, Average SSIM values under different self-constrained regularization weights \n(n = 10). i, SN2N denoising results under different photon levels (from several \nhundred to single photon, Methods). j, SSIM values under different photon \nlevels. In a, e, g, and i, the models were trained with 50 frames (full data case).  \nIn c, the models were trained with both 50 frames (full data case) and one \nimage (1/50 data case). In a, c, e, and g, the models were trained under noise \nlevel 1 conditions. Error bars: s.e.m. Experiments were repeated ten times \nindependently with similar results; scale bars, 1 µm.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 3 | See next page for caption.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 3 | Testing results of different noise levels and data \namounts. a, Denoising results of various methods under three different noise \nlevels (Level 3, Level 2, and Level 1, from top to bottom) using the full training set. \nFrom left to right: Raw input, denoising results of PURE, ACsN, N2V, supervised \nlearning ('Supervised'), SN2N without constraint ('SN2N w/o c'), and full SN2N. \nb, Quantitative comparisons of the results shown in a using PSNR (left), SSIM \n(middle), and RMSE (left) metrics (n = 10, mearsurements). c, Denoising results \nof learning-based methods using three different amounts (1/50, 1/5, and full \ndata, from top to bottom) of training data under Level 1 of noise. d, Quantitative \ncomparison of the results shown in c using PSNR (left), SSIM (middle), and RMSE \n(left) metrics (n = 10, mearsurements). k denotes the slope (red lines) of the \ncorresponding metric values along the data increment. e-f, Data uncertainty \n(e) and model uncertainty (f) of neural network models trained by different \ndata amounts. Average standard derivation (STD) values calculated from ten \npredictions of ten repetitively acquired data or ten repetitively trained models. \nCenterline, medians; limits, 75% and 25%; whiskers, maximum and minimum; \nerror bars, s.e.m. Experiments were repeated three times independently with \nsimilar results; scale bars, 1 µm.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 4 | See next page for caption.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 4 | Comparisons of SN2N versus RL-SN2N using SD-SIM \nand applying SN2N on SD-confocal microscopy. a, Zoomed-in views (top) and \nFWHM distribution plots (bottom, calculated by LuckyProfiler) of denoising \nresults by different methods (c.f., Fig. 2a). b, Comparison of SN2N and RL-SN2N. \nTop: SD-SIM (left) and its RL result (right); Bottom SN2N result (left) and RL-SN2N \nresult (right). c, LRQ values of results in b. d, SN2N denoising results (right) of SD \nconfocal image (left) recording the Argo-SIM slide. e, LRQ values of results in d. \nf, Comparisons of SN2N with 2D U-Net (left), SN2N with 3D U-Net (middle), and \nRL-SN2N with 3D U-Net (right) of volumetric data (c.f., Fig. 3b). g, Magnified views \nand their xz and yz cross-sections from white boxed regions in f. Experiments \nwere repeated three times independently with similar results; scale bars, 500 nm \n(a, b, d); 1 µm (f, g).\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 5 | See next page for caption.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 5 | SN2N-empowered automated subcellular segmentation \nand tracking. a, Workflow. Step 1, RL deconvolution; Step 2, RL-SN2N inference; \nStep 3, segmentation; Step 4, tracking. Step 5, extraction of motion features. Step \n6, classification. Step 7, topology graph construction. Step 8, specific downstream \nanalysis. b, A representative example for dual-color SR imaging of mitochondria \n(Mito, green) and ER (magenta) labeled with Tom20–mCherry and Sec61β-EGFP in \nlive COS-7 cells under raw SD-SIM (left) and RL-SN2N (right). c, The white box in b is \nenlarged and shown at seven time points under different configurations. From top \nto bottom: Raw SD-SIM, dual-color RL-SN2N, single-channel (Mito) RL-SN2N, RL SD-\nSIM segmentation, and RL-SN2N segmentation results. The yellow and white arrows \nindicate the mitochondrial fission and before fission, respectively. d, e, Results of \nMito (d) and ER (e) segmentations (first row) using the Otsu hard threshold (first \ncolumn) and Mitonet/ERnet (second column) and their skeletonizations (second \nrow) under SD-SIM (left) and RL-SN2N (right). f, Otsu segmentation results for  \nLys (red) and GA (blue) under SDSIM (left) and RL-SN2N (right). g, A representative \n4-color segmentation result under RL-SN2N. h, Spatial distribution of Lys assigned \nwith different motion behaviors. i, Distribution of estimated α values of Lys versus \ntheir temporal average of minimum distances to Mito (n = 46). j, Distribution of the \nLys-Mito MOCs’ standard deviation (S.D.) versus their mean values. k, Illustrations \nof the MSD curves for different motion behaviors of Lys. Curves are color-coded \nby the corresponding ER-Lys distances. Experiments were repeated three times \nindependently with similar results; scale bars, 2 μm (b, c, e).\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 6 | See next page for caption.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 6 | RL-SN2N can suppress noise in undersampled \ndata from EMCCD SD-SIM. a, 3D renderings of live COS-7 cells labeled with \nHoechst (green) and MitoTracker Deep Red (magenta) under raw SD-SIM \nequipped with an EMCCD camera (94 nm pixel size versus <150 nm resolution). \nb, Representative lateral slices from volume in a. c, RL-SN2N results of a with \nadditional 2× upsampling before RL deconvolution (47 nm pixel size).  \nd, Representative lateral slices from volume in c. e, Another RL-SN2N time point. \nf, Representative lateral slices from volume in e. Experiments were repeated \nthree times independently with similar results; scale bars, 5 μm.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 7 | See next page for caption.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 7 | Full data of SOFI-SN2N results and SN2N-assisted \nexpansion microscopy (ExM-SN2N). a, The whole field-of-views of the  \nwide-field (top left), 20-frames 2nd order SOFI (2nd SOFI 20f, top right), 2D-SIM \n(bottom left), and 20-frame 2nd order SOFI-SN2N (bottom right) images  \n(c.f., Fig. 6b). b, SN2N results of 2nd, 3rd, and 4th orders SOFI (from left to right) \nusing 20, 50, 100, 200, 500, 1000 frames (from top to bottom) (c.f., Fig. 6e). \nc-e, Average SSIM values of 2nd (c), 3rd (d), and 4th (e) SOFI-SN2N results (n = 5, \nmearsurements). (f) Comparison of temporal and spatial sampling methods. \nFrom left to right: SOFI reconstruction, SN2N result using temporal sampling \n(the first 20 frames vs. the second 20 frames), and SN2N result using spatial \nsampling. g, A 2 times-expanded (2×, top) and 4-times expanded (4×, bottom) \nCOS-7 cell was immunostained with a primary antibody against α-tubulin and a \nsecond antibody conjugated with Alexa Fluor 488 under wide-field microscopy \n(left) and its SN2N denoised result (right). Signal-to-background ratios (SBR) \nare labeled. h, Magnified views of the white boxed regions in g under ExM (top) \nand SN2N denoised results (bottom). i, Intensity profiles and multiple Gaussian \nfitting of the filaments indicated by the white arrows in h. Numbers represent the \ndistances between peaks; a.u., arbitrary units. j, A 4.5-times expanded (4.5×) COS-\n7 cell labeled with Sec61β–GFP under wide-field microscopy (left) and its SN2N \ndenoised result (right). k, Enlarged regions enclosed by the white box in j seen \nunder ExM-4.5× (left) and its SN2N result (right). Centerline, medians; limits, 75% \nand 25%; whiskers, maximum and minimum; error bars, s.e.m. Experiments were \nrepeated three times independently with similar results; scale bars, 2 µm (a),  \n1 µm (b, h, j, k), and 5 μm (g).\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 8 | See next page for caption.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 8 | SN2N removes random, non-continuous artifacts in \nlow-SNR SIM with two strategies. a, Pipeline of SIM-SN2N using raw image \nresampling strategy (Methods). The self-supervised data generation is applied on \nthe 9 raw images followed by the SIM reconstruction. b, d, f, clathrin-coated pits \n(CCPs, b), microtubules (d), and ER (f), recorded by SIM under low-SNR condition \n(SIM-L, left) and their SN2N results (right). c, e, g, SIM reconstructions (left) of \nCCPs (c), microtubules (e), and ER (g) under low (L, top), medium (M, middle), \nand high (H, bottom) SNR conditions and their SN2N results (right). SSIM values \nare labeled on the bottom right corners. h, The mitochondrial cristae structures \nin live COS-7 cells labeled with MitoTracker Green under 2D-SIM (bottom left \nboxed region) and SN2N-SIM imaging at the first time point. i, j, Representative \nmontages of 11 time points from yellow boxed region in h under 2D-SIM (i) and \nSIM-SN2N (j). k, Workflow of SIM-SN2N using SIM image resampling strategy \n(Methods). After SIM reconstruction, we apply a SIM-specfic self-supervised \ndata generation involving 3 × 3 pixels (1 + 3 + 7 + 9 versus 2 + 4 + 6 + 8) resampling \nfollowed by a 3× Fourier interpolation. l, n, p, CCPs (l), microtubules (n), and \nER (p), recorded by SIM under low-SNR condition (left) and their SN2N results \n(right). m, o, q, SIM reconstructions (left) of CCPs (m), microtubules (o), and ER \n(q) under low (top), medium (middle), and high (bottom) SNR conditions and \ntheir SN2N results (right). SSIM values are labeled on the bottom right corners.  \nr, A representative living COS-7 cell labeled with LifeAct–EGFP under ultrafast \nTIRF (left), TIRF-SIM (middle), and SIM-SN2N (right). s, t, Enlarged regions \nenclosed by the white box in r, under TIRF-SIM (s) and SIM-SN2N (t). Experiments \nwere repeated three times independently with similar results. Scale bars,  \n2 μm (b, d, f, h) and 1 μm (c, e, g, j, r, t).\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 9 | See next page for caption.\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 9 | SN2N maintains the linear response of Ca2+ transients \nobtained by the SD-SIM. a, A representative live COS-7 cell was transfected with \nGCaMP6s, stimulated with ATP (10 μM). One snapshot under the SD-SIM (left) \nand after the SN2N (right) were shown. b, Magnified views of regions enclosed \nby white boxes 1-4 in a. c, ATP stimulated calcium traces from corresponding \nmacrodomains in b. d, Denoising of fast calcium transients using published two-\nphoton microscopy dataset from ref. 22. From left to right: Low-SNR recording of \nsomatic signals, DeepCAD data, SN2N (spatial) denoising results, SN2N (temporal) \ndenoising results by replacing the spatial diagonal resampling with the temporally \ninterleaved resampling, and the high-SNR data (tenfold imaging SNR). Magnified \nview of white-boxed regions is shown at the bottom row. e, Fluorescence traces \nextracted from the yellow boxed regions in d. The trajectories’ Pearson correlation \ncoefficients (r) were labeled at the bottom right corners. Experiments were \nrepeated three times independently with similar results. Scale bars, 5 µm (a),  \n2 µm (b), and 50 µm (d).\n\n\nNature Methods\nArticle\nhttps://doi.org/10.1038/s41592-024-02400-9\nExtended Data Fig. 10 | Generalization of SN2N across different SNR \nconditions and pixel sizes (structural scales). a-f, Testing generalization across \ndifferent SNR conditions. Color-coded 3D distributions and their xz and yz cross-\nsections of all mitochondria (labeled with Tom20–mCherry) in a live COS-7 cell \n(c.f., Fig. 3b). a-c, SN2N results (trained with the first volume) of the first volume \n(0 min) (a) and the last volume (2.5 min) (b), and the SN2N prediction of SN2N \nperdition (SN2N2) from the last SD-SIM volume (c). d, e, SN2N results (trained with \nthe last SD-SIM volume) of the first (d) and the last (e) volume. f, Zoom-in views \nfrom white-boxed regions in a-e. First column: 0 min (top) and 2.5 min (bottom) \nSD-SIM; second column: SN2N and SN2N2 (bottom half of bottom) results (trained \nwith 0 min SD-SIM volume) of 0 min (top) and 2.5 min (bottom) SD-SIM; SN2N \nresults (trained with 2.5 min SD-SIM volume) of 0 min (top) and 2.5 min (bottom) \nSD-SIM. g-k, Testing generalization across different pixel sizes. g, Nuclear pores in \nHeLa cells were labeled with an anti-Mab414 primary antibody and the Alexa594 \nsecondary antibody and observed under STED and STED-SN2N configurations.  \nh, STED images (left) of 20.66 nm pixel size (top), 7.10 nm pixel size (middle),  \nand 20.66 nm pixel size subsampled from 7.10 nm (bottom), and their SN2N \nresults (right) from model trained by data of 20.66 nm pixel size. i, STED images \n(left) of 7.10 nm pixel size (top), 20.66 nm pixel size (middle), and 7.10 nm pixel  \nsize Fourier upsampled from 20.66 nm (bottom), and their SN2N results (right) \nfrom the model trained by data of 7.10 nm pixel size. j, k, Average FWHM values  \nof STED (gray) and SN2N results (yellow) from the model trained by data of  \n20.66 nm pixel size (j) and 7.10 nm pixel size (k) (n = 5, measurements). Centerline, \nmedians; limits, 75% and 25%; whiskers, maximum and minimum; error bars,  \ns.e.m. Experiments were repeated three times independently with similar results. \nScale bars, 5 µm (e), 1 µm (f-h).\n\n\n\n\nα\nα\n\n\n\n\nNature Computational Science | Volume 3 | December 2023 | 1067–1080\n1067\nnature computational science\nhttps://doi.org/10.1038/s43588-023-00568-2\nArticle\nSpatial redundancy transformer for self-\nsupervised fluorescence image denoising\nXinyang Li \n \n 1,2,3,9, Xiaowan Hu2,9, Xingye Chen1,3,4,9, Jiaqi Fan2,5, Zhifeng Zhao1,3, \nJiamin Wu \n \n 1,3,6,7 \n, Haoqian Wang \n \n 2,8 \n & Qionghai Dai \n \n 1,3,6,7 \nFluorescence imaging with high signal-to-noise ratios has become \nthe foundation of accurate visualization and analysis of biological \nphenomena. However, the inevitable noise poses a formidable challenge \nto imaging sensitivity. Here we provide the spatial redundancy denoising \ntransformer (SRDTrans) to remove noise from fluorescence images in \na self-supervised manner. First, a sampling strategy based on spatial \nredundancy is proposed to extract adjacent orthogonal training pairs, which \neliminates the dependence on high imaging speed. Second, we designed \na lightweight spatiotemporal transformer architecture to capture long-\nrange dependencies and high-resolution features at low computational \ncost. SRDTrans can restore high-frequency information without producing \noversmoothed structures and distorted fluorescence traces. Finally, we \ndemonstrate the state-of-the-art denoising performance of SRDTrans \non single-molecule localization microscopy and two-photon volumetric \ncalcium imaging. SRDTrans does not contain any assumptions about the \nimaging process and the sample, thus can be easily extended to various \nimaging modalities and biological applications.\nThe rapid development of intravital imaging techniques enables \nresearchers to observe biological structures and activities at microm-\neter and even nanometer scales1,2. As an imaging method with great \nprevalence, fluorescence microscopy has contributed to the discov-\nery of a series of new physiological and pathological mechanisms \ndue to its high spatiotemporal resolution and molecular specific-\nity3–5. The fundamental goal of fluorescence microscopy is to obtain \nclean and sharp images containing sufficient information about the \nsample, which can guarantee the accuracy of downstream analysis \nand support convincing conclusions. However, limited by multiple \nbiophysical and biochemical factors (for example, labeling concen-\ntration, fluorophore brightness, phototoxicity, photobleaching and \nso on), fluorescence imaging is conducted in photon-limited condi-\ntions and the inherent photon shot noise severely degrades the image \nsignal-to-noise ratio (SNR), especially in low-illumination and high- \nspeed observations6.\nVarious methods have been proposed to remove noise from fluo-\nrescence images. Conventional denoising algorithms based on numeri-\ncal filtering and mathematical optimization suffer from unsatisfactory \nperformance and limited applicability7,8. In the past few years, deep \nlearning has shown remarkable performance in image denoising9,10. \nAfter iterative training on a dataset with ground truth (GT), deep neu-\nral networks can learn the mapping between noisy images and their \nclean counterparts. Such a supervised manner depends heavily on \npaired GT images11–15. When observing the activity of living organ-\nisms, obtaining pixel-wise registered clean images is a great challenge \nbecause the sample often undergoes fast dynamics. To alleviate this \ncontradiction, some self-supervised methods have been proposed \nReceived: 14 June 2023\nAccepted: 7 November 2023\nPublished online: 11 December 2023\n Check for updates\n1Department of Automation, Tsinghua University, Beijing, China. 2Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, \nChina. 3Institute for Brain and Cognitive Sciences, Tsinghua University, Beijing, China. 4Research Institute for Frontier Science, Beihang University, \nBeijing, China. 5Department of Electronic Engineering, Tsinghua University, Beijing, China. 6Beijing Key Laboratory of Multi-dimension and Multi-scale \nComputational Photography (MMCP), Tsinghua University, Beijing, China. 7IDG/McGovern Institute for Brain Research, Tsinghua University, Beijing, \nChina. 8The Shenzhen Institute of Future Media Technology, Shenzhen, China. 9These authors contributed equally: Xinyang Li, Xiaowan Hu, Xingye Chen. \n e-mail: wujiamin@tsinghua.edu.cn; wanghaoqian@tsinghua.edu.cn; qhdai@tsinghua.edu.cn\n\n\nNature Computational Science | Volume 3 | December 2023 | 1067–1080\n1068\nArticle\nhttps://doi.org/10.1038/s43588-023-00568-2\nor vertically to exploit spatial correlations fully and isotropically. \n \nA simplified implementation of sampling is depicted in Fig. 1b. As the \nnoise of adjacent pixels is independent while the signals are closely \ncorrelated, the substack filled by the central pixels can be used as the \ntraining input and the other two spatially adjacent substacks are used as \ncorresponding targets to optimize the network parameters. Compared \nwith other methods19,20,22, our sampling strategy is more effective and \ncomprehensive in preserving both spatial and temporal information \n(Supplementary Figs. 1 and 2). In the inference stage, low-SNR stacks \nwill be fed into pre-trained SRDTrans models without spatial downsam-\npling. To overcome the locality of convolutional kernels, we constructed \na transformer network to capture endogenous non-local spatial fea-\ntures and long-range temporal dependencies using the self-attention \nmechanism (Fig. 1c). The restoration of each pixel can simultaneously \nintegrate the temporal information of all frames and the spatial infor-\nmation of all pixels, even if they are far from each other. Besides, the \nnetwork does not contain any spatial downsampling module, allowing \nmore high-frequency components to flow through the network and \navoiding the loss of spatial resolution. Furthermore, as the amount \nof data in fluorescence imaging is often very large, sometimes at the \npetabyte scale, the transformer network was designed to be as light-\nweight as possible to relieve the computational burden of large-scale \ndata processing. Compared with other transformer networks27–30, our \narchitecture can achieve the best performance with more than one \norder of magnitude fewer parameters (Supplementary Tables 1 and 2). \nThe lightweight architecture of SRDTrans also makes it easy to train a \ngood model even with a small amount of training data (for example, \n500 frames, 490 × 490 pixels each frame), relieving the pressure of \ncapturing large-scale datasets (Supplementary Fig. 3).\nTo demonstrate the predominance of our transformer network \nover CNNs, we generated simulated calcium imaging data (Methods \nand Supplementary Fig. 4) and used them to train a 3D U-Net31, as \nwell as the transformer of SRDTrans, using the same spatial redun-\ndancy sampling strategy. We term the former as spatial redundancy \ndenoising CNN (SRDCNN). The visualized feature maps of deep layers \nintuitively show the superiority of SRDTrans in revealing fine-grained \npatterns (Fig. 1d). As features flow through the network, the limited \nreceptive field of convolutional kernels makes CNN-based methods \nfocus on only rough features while our transformer architecture still \nhas a strong perception of sophisticated structures. We also compare \nthe denoised images of the two architectures (Fig. 1e). The result of \nSRDCNN is obviously oversmoothed, especially in regions with sharp \nedges, which is a manifestation of spectral bias that the low-frequency \ninformation is overfitted while the high-frequency information is \nhardly preserved (Supplementary Fig. 5). This deficiency is largely alle-\nviated by SRDTrans, and more subcellular structures such as dendritic \nfibers can be restored accurately. The intensity profile deconstructed \nfrom the SRDTrans denoised image is more consistent with the GT \n(Fig. 1f). Moreover, lacking the ability to capture long-range temporal \nfor more applicable and practical denoising in fluorescence imag-\ning6,16–23. Among them, the first kind of methods rely on the similarity \nbetween adjacent frames6,16–18. But when the sample changes very \nquickly or the imaging speed is too slow, the time-lapse data cannot \nprovide enough temporal redundancy. This is a common problem in \nvolumetric imaging as the volume rate decreases proportionally to \nthe number of imaging planes. The dissimilarity between adjacent \nframes will result in inferior performance and distorted structures and \nfluorescence kinetics. The other kind of methods learn to denoise only \nfrom spatially adjacent pixels in two-dimensional frames19–23. However, \nwithout utilizing endogenous temporal correlations, these methods \nperform poorly on time-lapse imaging. Therefore, to achieve better \ndenoising performance, the ability to simultaneously extract global \nspatial information and long-range temporal correlations is essential, \nwhich is lacking in convolutional neural networks (CNNs) because of \nthe locality of convolutional kernels24. Moreover, the inherent spectral \nbias makes CNNs tend to fit low-frequency features preferentially while \nignoring high-frequency features, inevitably producing oversmoothed \ndenoising results25.\nHere we present the spatial redundancy denoising transformer \n(SRDTrans) to address these dilemmas. On the one hand, a spatial \nredundancy sampling strategy is proposed to extract three-dimen-\nsional (3D) training pairs from the original time-lapse data in two \northogonal directions. This scheme has no dependence on the similar-\nity between two adjacent frames, so SRDTrans is applicable to very fast \nactivities and extremely low imaging speed, which is complementary \nto our previously proposed DeepCAD that leverages temporal redun-\ndancy6,18. On the other hand, we designed a lightweight spatiotemporal \ntransformer network to fully exploit long-range correlations. The \noptimized feature interaction mechanism allows our model to obtain \nhigh-resolution features with a small number of parameters. Compared \nwith classical CNNs, the proposed SRDTrans has stronger abilities for \nglobal perception and high-frequency maintenance, enabling the rev-\nelation of fine-grained spatiotemporal patterns that were previously \nindiscernible. We demonstrate the superior denoising performance of \nSRDTrans on two representative applications. The first one is single-\nmolecule localization microscopy (SMLM) with adjacent frames being \nrandom subsets of fluorophores26. The other one is two-photon calcium \nimaging of large 3D neuronal populations with a volumetric speed as \nlow as 0.3 Hz. Extensive qualitative and quantitative results indicate \nthat SRDTrans can serve as a fundamental denoising tool for fluores-\ncence imaging to observe various cellular and subcellular phenomena.\nResults\nPrinciple of SRDTrans\nThe self-supervised framework of SRDTrans is shown schematically \nin Fig. 1a. For spatial redundancy sampling, spatially adjacent sub-\nstacks are sampled by orthogonal masks from the original low-SNR \nimage stack. Each target is adjacent to the input stack horizontally \nFig. 1 | Principle of SRDTrans and performance evaluation. a, Self-supervised \ntraining strategy of SRDTrans. The original low-SNR stack of H × W × T pixels  \nis sampled by orthogonal masks, producing three downsampled substacks \n(input, target 1 and target 2) of H/2 × W/2 × T pixels. The ‘input’ substack is fed  \ninto the transformer network, and the corresponding output is compared with \nthe ‘target’ substacks to calculate the loss function for parameter optimization. \nb, Simplified schematic of spatial redundancy sampling (H = 4, W = 4, T = 1).  \nA 4 × 4 patch is split into four 2 × 2 blocks and three adjacent pixels are randomly \nselected in each block. The central pixel (labeled as ‘2’) is horizontally or vertically \nadjacent to the other two pixels (labeled as ‘1’ and ‘3’). c, The architecture of \nthe lightweight spatiotemporal transformer. It consists of a temporal encoder \nmodule, an STB and a temporal decoder module. Each temporal encoder \ncompresses the temporal scale (t) of the input by a factor of r (r = 4 in this work) \nusing convolution. In the STB module, the input is divided into small patches, \nand different feature maps of the same spatial position are stitched together \nin the patch flattening layer. The position embedding layer records the spatial \nposition of each patch so that it can be mapped back after the global interaction \nin the multi-head self-attention layer. The self-attention mechanism can calculate \nthe spatiotemporal correlation between all local patches. The output of the STB \nmodule will be uncompressed to the original temporal scale by the following \ntemporal decoder module. d, Visualizing the feature responses in SRDCNN (the \nlast layer of STB) and SRDTrans (the last layer of 3D U-Net). SRDCNN represents \nthe method that replaces the transformer network in SRDTrans with a 3D U-Net. \nScale bar, 60 μm. e, Comparing the denoising performance of SRDCNN and \nSRDTrans on simulated calcium imaging data (30 Hz). Scale bars, 40 μm for the \nwhole FOV and 10 μm for magnified views. f, Pixel intensity along the red dashed \nline in e. g, Evaluating the ability of SRDCNN and SRDTrans to capture long-range \ndependencies. Models were trained and validated on simulated calcium imaging \ndata (30 Hz) of different input temporal scales (T). All values are shown as \nmean ± s.d. (N = 6,000 independent frames).\n\n\nNature Computational Science | Volume 3 | December 2023 | 1067–1080\n1069\nArticle\nhttps://doi.org/10.1038/s43588-023-00568-2\ncorrelations also drags down the denoising performance of SRDCNN \non time-lapse imaging data. For SRDTrans, the output SNR continu-\nously grows as the input temporal scale (T) increases (Fig. 1g). A more \ncomprehensive evaluation of the influence of input temporal scale \nindicates that SRDTrans can make full use of the information offered \nby temporally distant pixels (Supplementary Figs. 6 and 7). We also \ninvestigated the generalization ability of the proposed method by \ncross-dataset and cross-modality validation, which shows that training \non data with the same SNR and imaging modality can obtain the best \ndenoising performance (Supplementary Fig. 8). To verify the practi-\ncality of SRDTrans at extremely low imaging speed, we compared the \nperformance of methods combining different networks and sampling \nschemes on simulated calcium imaging data sampled at 0.3 Hz (Sup-\nplementary Video 1). When the similarity between adjacent frames is \nlow, using spatial redundancy sampling is more reasonable (Supple-\nmentary Table 3). Succinctly, the synergy between spatial redundancy \nsampling and dedicated transformer architecture endows SRDTrans \nwith the ability to restore high-resolution structures and fast dynamics.\nHigh-performance SMLM with SRDTrans denoising\nGiven N detected fluorescence photons, the lower bound of the preci-\nsion of SMLM scales to 1/√N  (ref. 26), which is the mathematical \n \nc\nSTB\nc\nPatch flattening\nMulti-head self-attention\nPosition embedding\nSpatiotemporal transformer block (STB) \nTemporal encoder\n=\n=\nTemporal decoder\na \nb \nOutput\nLearning\nLearning\nTarget 2\nTraining pairs\nImages\nH\nW\nSpatial redundancy  sampling\nT\nT\nOrthogonal masks\nTarget 1\nInput\nT\nf \ng \nRaw data\nd \nSRDCNN\nSRDTrans\nMask 1 Mask 2 Mask 3\nT\nW\nH\nNetwork\n1\n1\n2\n1\n3\n3 2\n3\n2\n1\n1\n1\n2\n2\n2\n2\n3\n3\n3\n1\n1\n1\n1\n2 2\n2 2\n3 3\n3 3\nMask 1\nMask 2\nMask 3\n3\n4 × 4 patch\nOrthogonal mask\n2 × 2 patches\n1\n1\n2 3\n1\nRaw\nSRDCNN\nSRDTrans\nGT\nt\nt\nr\nt\nr\nt\nW\n2\nH\n2\n0\n5\n10\n15\n0.2\n0.4\n0.6\n0.8\n1.0\nNormalized intensity (a.u.)\nDistance (µm)\nMask ID\n*\n*\nRaw data\nGT\nSRDCNN\ne \nNormalized intensity\n1.0\n0\nOutput SNR (dB)\n200\nT (slices)\n600\n1,000\n24\n25\n27\n28\n26\nSRDTrans\n\n\nNature Computational Science | Volume 3 | December 2023 | 1067–1080\n1070\nArticle\nhttps://doi.org/10.1038/s43588-023-00568-2\nformula of the shot-noise limit or the well-known standard quantum \nlimit32,33. This indicates that the fundamental physical limit of localiza-\ntion precision is shot noise. To investigate the benefits that our denois-\ning method can bring to SMLM, we first applied SRDTrans to simulated \nstochastic optical reconstruction microscopy (STORM) data with GT \nfor quantitative evaluation34,35. The noise-free single-molecule-emis-\nsion images were synthesized by TestSTORM36 and corresponding \nnoisy images with different SNRs were then generated by applying \ndifferent levels of mixed Poisson–Gaussian noise (Methods and Sup-\nplementary Fig. 9). For the image stack of each SNR, we trained a spe-\ncific model for it and then processed it using the model to obtain the \ndenoised image stack. Quantitative comparisons of the visualized \nimages (Fig. 2a and Supplementary Video 2) and the extracted inten-\nsity profiles (Fig. 2b) show that the results of SRDTrans are highly \nconsistent with the GT. Over a wide range of imaging SNRs, including \nsome extremely low-SNR conditions, SRDTrans can substantially \nimprove the image quality evaluated by the SNR at the pixel-intensity \nlevel and structural similarity (SSIM) at the perception level (Fig. 2c). \nCompared with other self-supervised methods16–22, SRDTrans can \nbetter preserve the distribution and intensity of emitters owing to its \nstrong ability in exploiting high-resolution features and long-range \ndependencies (Supplementary Fig. 10).\nNext, we evaluate the improvement of single-molecule localiza-\ntion performance with the enhancement of the image SNR. To recon-\nstruct super-resolved images, we applied ThunderSTORM37, one of the \nmost frequently used localization software with an excellent balance \nbetween accuracy and execution runtime38,39. The image reconstructed \nfrom the original noisy data is contaminated by noise and contains \nmany misidentified molecules (Fig. 2d). By contrast, the reconstructed \nimage from SRDTrans denoised images reveals clear and continuous \ncytoskeletal filaments that are not previously recognizable because \nof suppressed localization error and improved resolution (Fig. 2e and \nSupplementary Fig. 11). For better quantitative analysis, we matched \nthe detected fluorescent molecules with the GT using the Hungar-\nian algorithm39. From raw images, few fluorescent molecules can be \ndetected and most of them are wrongly localized. After SRDTrans \ndenoising, almost all molecules can be detected in high agreement \nwith the GT (Fig. 2f and Supplementary Fig. 12). Using the Jaccard \nindex and root-mean-squared error (r.m.s.e.) as the metrics to quantify \nthe proportion of correctly detected molecules and the localization \naccuracy of those detected molecules, respectively, we found that the \nJaccard index was improved by ~6-fold (85.7 ± 3.51% versus 14.1 ± 2.76%, \nmean ± s.d.) and r.m.s.e. was reduced by ~3.4-fold (24.86 ± 3.24 nm \nversus 85.94 ± 7.76 nm) after denoising (Fig. 2g). From a more com-\nprehensive perspective, we further adopted the metric termed effi-\nciency that combines Jaccard index and r.m.s.e.39. The results show that \nSRDTrans improved the efficiency of single-molecule localization from \n−21.54 ± 5.71 to 71.33 ± 3.38 (Fig. 2h), fully demonstrating the benefits \nof SRDTrans on SMLM.\nWe further applied SRDTrans to experimental SMLM data of micro-\ntubules to validate its ability in revealing subcellular structures. Raw \nframes were captured with low excitation power and short exposure \ntime to reduce phototoxicity and emitter density. The experimentally \nobtained single-molecule-emission images and SRDTrans denoised \nimages are shown in Fig. 3a. Disturbed by the noise, the reconstruc-\ntion algorithm can hardly localize the fluorescent molecules in the \nraw frames. The reconstructed super-resolution image contains \nmany erroneous spots and cannot reveal any meaningful structures \n \n(Fig. 3b). In comparison, SRDTrans can effectively suppress the noise \nand remove localization artifacts in the reconstructed image, making \nthe distribution and extension directions of microtubules visible. We \ncomputed the Fourier-ring correlation (FRC)40,41 curve to quantify the \nresolution from the SMLM images. The image resolution is defined as \nthe inverse of the spatial frequency at the intersection of the FRC curve \nand the threshold line (~0.143). Benefitting from the removal of noise, \nthe resolution of SRDTrans denoised data was improved from 52.4 nm \nto 36.0 nm (Fig. 3c) and the localization uncertainty was reduced from \n8.0 ± 6.88 nm to 5.0 ± 1.34 nm (Fig. 3d). In addition to the data acquired \nby our instrument, we also used SRDTrans to denoise publicly avail-\nable SMLM data contributed by other laboratories42 (Fig. 3e). The \nreconstructed super-resolution images indicate that SRDTrans can \neliminate the artifacts and bring more complete organelle structures \n(Fig. 3f,g). We applied Gaussian fitting to the intensity profile perpen-\ndicular to the microtube filaments and measured the full-width at \nhalf-maximum (FWHM) to quantify the image resolution (Fig. 3h). The \nSRDTrans denoised data show improved resolution as the averaged \nFWHM dropped from 187.89 ± 22.22 nm to 60.96 ± 7.51 nm (Fig. 3i). As \nSMLM is heading towards live-cell imaging and long-term observation43 \nour denoising method promises to be a beneficial tool to reduce the \nlaser power by resolving fluorescent molecules from very-low-SNR \nframes. For 3D SMLM, as the axial positions of molecules are estimated \nthrough point-spread-function engineering26, SRDTrans is expected to \noffer great help by resolving single-molecule-emission patterns from \nlow-SNR images.\nApplying SRDTrans to two-photon volumetric calcium \nimaging\nIn multiphoton imaging, the volumetric imaging speed decreases \nlinearly as the number of scanning planes increases. Thus, the achiev-\nable sampling rate for observing neuronal populations with large axial \nranges is often quite low, making the denoising methods that rely on \nthe similarity between temporally adjacent frames infeasible6,16–18. How-\never, SRDTrans provides an opportunity to restore the highly degraded \nfluorescence signals in large-scale volumetric calcium imaging by \nutilizing the similarity between spatially adjacent pixels. To evaluate \nthe denoising performance of SRDTrans on calcium imaging with dif-\nferent sampling rates, we generated realistic calcium imaging data \nwith synchronized GT using neural anatomy and optical microscopy \n(NAOMi)44. We started from denoising high-sampling-rate (30 Hz) data \nand found that SRDTrans can effectively remove noise and recover \npreviously indiscernible structures such as soma, neurites and vascular \nshadows (Fig. 4a, Supplementary Fig. 13 and Supplementary Video 3). \n \nThe enhancement is manifested not only in the visual effect but also, \nmore importantly, in the accurate restoration of pixel intensities \n \nFig. 2 | Validation of SRDTrans on simulated SMLM data. a, Single-molecule-\nemission images before and after denoising. Left: raw data. Middle: SRDTrans \ndenoised data. Right: GT. Magnified views of boxed regions show the emission \npattern of a bunch of fluorescent molecules. Scale bars, 2 μm for the whole FOV \nand 0.5 μm for magnified views. The SNR value of the raw data and denoised data \nare noted. b, Intensity profiles along the white dashed lines in a. c, Quantitative \nevaluation of the denoising performance with SNR and SSIM. Left: image SNR \nbefore and after denoising. Right: image SSIM before and after denoising. Each \ndata point shows the statistical result of 24,000 frames. All values are shown \nas mean ± s.d. (N = 24,000 independent frames). d, Reconstructed super-\nresolution images of microtubules before and after denoising. Left: the image \nreconstructed from raw data. Middle: the image reconstructed from SRDTrans \ndenoised data. Right: GT. Scale bar, 5 μm. e, Merged image of the yellow boxed \nregion in d. Magenta, the image reconstructed from raw data; green, the image \nreconstructed from SRDTrans denoised data; red, GT. The overlapping positions \nof red and green appear yellow. Scale bar, 1 μm. f, Consistency analysis of the \nlocalized fluorescent molecules in raw images (left) and SRDTrans denoised \nimages (right) relative to the GT. A magnified view of the boxed region is shown \nat the bottom left of each panel. g, Tukey box-and-whisker plots (Methods) \nshowing the localization precision quantified with the Jaccard index (left, higher \nis better) and r.m.s.e. (right, lower is better) before and after SRDTrans denoising \n(N = 5,000 independent molecules). h, Evaluating the performance of single-\nmolecule localization before (blue) and after (orange) denoising with a more \ncomprehensive metric termed efficiency.\n\n\nNature Computational Science | Volume 3 | December 2023 | 1067–1080\n1071\nArticle\nhttps://doi.org/10.1038/s43588-023-00568-2\n(Fig. 4b). Visualization in the frequency domain (calculated by discrete \nFourier transform) shows that SRDTrans can restore most of the fre-\nquency components (Fig. 4c), especially the high-frequency compo-\nnents lost by DeepCAD6,18 and DeepInterpolation17, thereby leading to \nhigh performance in denoising calcium imaging data (Supplementary \n \nFig. 14). Such a remarkable denoising capability can be maintained over \na wide range of input SNRs (from −2.08 dB to 17.68 dB), and the average \nSNR improvement is about 22 ± 2.47 dB (Fig. 4d). We also verified the \nperformance of SRDTrans on experimentally obtained calcium imag-\ning data with a synchronized high-SNR (tenfold photons) reference6, \nwhich shows that the neuronal structures and dynamics swamped by \nnoise can be restored authentically (Supplementary Fig. 15).\nThen we investigate the denoising performance of SRDTrans on \ncalcium imaging data sampled at 0.3 Hz, which is 100 times lower than \nthe imaging speed demonstrated above. Bilateral assessments in both \nthe space domain and the frequency domain reveal that SRDTrans can \na\nGT\nRaw data (SNR = -0.05 dB)\nSRDTrans (SNR = 26.53 dB)\nc\nGT\nRaw data\nSRDTrans\nd\nGT\nRaw data\nSRDTrans\nMerged\ne\nf\nNormalized intensity (a.u.)\nDistance (µm)\nDistance (µm)\nb\nRaw data\nSRDTrans\nGT\nGT molecules\nDetected molecules\ng\nRaw data\nSRDTrans\nGT\nRaw data\nSRDTrans\nInput SNR (dB)\n20\nInput SSIM\nh\nEficiency = 80\nEficiency = 60\nEficiency = 40\nEficiency = 20\nEficiency = 0\nEficiency = –20\nr.m.s.e. (nm)\n25\n50\n75\n100\nJaccard (%)\n0\n20\n40\n60\n80\n100\nRaw data\nSRDTrans\nRaw data\nSRDTrans\nOutput SNR (dB)\n0\n5\n10\n15\n10\n20\n30\n40\n0.2\n0.4\n0.6\n0.8\nOutput SSIM\n0.4\n0.6\n0.8\n0.3\n0.9\n1.5\n0.2\n0.4\n0.6\n0.8\n1.0\n0.3\n0.9\n1.5\n0.2\n0.4\n0.6\n0.8\n1.0\n1.0\n100\nJaccard (%)\n20\n40\n60\n80\n40\n80\n120\n60\n100\n20\nr.m.s.e. (nm)\n0\nRaw data\nSRDTrans\n\n\nNature Computational Science | Volume 3 | December 2023 | 1067–1080\n1072\nArticle\nhttps://doi.org/10.1038/s43588-023-00568-2\nRaw data\nSRDTrans\ne\ng\nRaw data\nSRDTrans\n(i)\n(i)\n(ii)\n(ii)\n(iii)\n(iii)\na\nb\nc\ni\nRaw data\nSRDTrans\n50\n150\n200\n100\n250\nRaw data\nSRDTrans\nRaw data\nSRDTrans\nh\nNormalized intensity (a.u.)\nDistance (nm)\n110.32 nm\n54.69 nm\n(ii)\n(ii)\n0\n100\n200\n300\n0\n100\n400\n115.79 nm\n66.70 nm\n(i)\n(i)\nDistance (nm)\n0.4\n0.8\n0.4\n0.8\n0.4\n0\nRaw data\nSRDTrans\n0.8\n90.45 nm\nDistance (nm)\n100\n300\nGaussian fitting\n(iii)\n(iii)\n62.29 nm\n200\nFWHM (nm)\nSpatial frequency (nm–1)\nRaw data\nSRDTrans\nSmooth fitting\nThreshold\nd\n0\n0.01\n0.02\n0.03\n0.8\n1.0\n0.2\n0.4\n0.6\nCut-of = 52.4 nm\nCut-of = 36.0 nm\nUncertainty (nm)\nRaw data\nSRDTrans\n0\n10\n20\n30\nf\nNormalized FRC\nRaw data\nSRDTrans\n\n\nNature Computational Science | Volume 3 | December 2023 | 1067–1080\n1073\nArticle\nhttps://doi.org/10.1038/s43588-023-00568-2\naccurately retrieve the fluorescence signals from the original highly \ndegraded images without structural blurring and frequency deficiency \n(Fig. 4e and Supplementary Fig. 14). When the sampling rate is much \nlower than the fluorescence dynamics, the large discrepancy between the \nsignals in two adjacent frames cannot provide the temporal correlation \nrequired by DeepCAD and DeepInterpolation, so they are not accurate \nenough to be used in conditions of low imaging speed or fast activity \n(Supplementary Fig. 16). To figure out how SRDTrans works at differ-\nent imaging speeds, we performed an ablation study on different sam-\npling strategies and network architectures (Fig. 4f and Supplementary \n \nTable 3). The results indicate that spatial redundancy sampling performs \nbetter at low imaging speeds, whereas temporal redundancy sampling \nperforms better at high imaging speeds. Almost equally for all imaging \nspeeds, our transformer architecture offers an additional improvement \n(~2.05 ± 0.27 dB) over conventional CNNs. In general, the synergistic \ncombination of the spatial redundancy sampling and the transformer \narchitecture in SRDTrans provides better performance than DeepCAD at \nall imaging speeds (Supplementary Figs. 17 and 18). In the time domain, \nthe superior ability of SRDTrans can reveal high-fidelity calcium transi-\ntions without distorting fluorescence kinetics (Fig. 4g). Moreover, we \nalso simulated fast-moving objects to imitate migrating cells that are \nwidely existed in living organisms. The quantitative evaluation shows \nthat SRDTrans can preserve the structure of densely distributed objects \neven if they are moving at a speed of up to 9 pixels per frame (Supple-\nmentary Fig. 19 and Supplementary Video 4), alleviating the shortage \nof denoising methods for fast-moving cells and organelles.\nFinally, we went a step further in denoising calcium imaging data \nby applying SRDTrans to volumetric recordings, which is not achievable \nfor other self-supervised denoising methods6,17,18 because of their heavy \nreliance on high sampling rates. We used transgenic mice expressing \nthe GCaMP6f genetically encoded calcium indicator45 and imaged \na brain volume of 500 × 500 × 200 μm3 in the mouse cortex using a \ntwo-photon microscope. We scanned 100 planes with a frame rate of \n30 Hz, and thus the overall volume rate was 0.3 Hz. For the denoising \nof volumetric calcium imaging data, we extracted all the frames of \neach imaging plane and reorganized them into a separate time-lapse \n(xy–t) stack. The time-lapse stacks of all imaging planes were used for \nnetwork training. A 3D visualization of the neural volume shows that \nthe spatial profiles and firing states of the neurons can be revealed \nafter denoising, which otherwise would be swamped by severe shot \nnoise (Fig. 5a and Supplementary Video 5). For better comparison, \nwe present the snapshots of a certain imaging plane at two different \nmoments. With the enhancement of SRDTrans, the structure and dis-\ntribution of the neurons become clearly observable (Fig. 5b). We also \nextracted the fluorescence traces along the time axis and found that \na large number of calcium transients can be restored after denoising \n(Fig. 5c). The dramatically improved SNR would propel the decoding \nof underlying neural activity from fluorescence signals. As neural cir-\ncuits in the mammalian brain are spatially coordinated and temporally \norchestrated, deciphering the function of large neuronal ensembles \nrequires large-scale volumetric imaging with a high SNR. The superior \ndenoising performance of SRDTrans provides an opportunity to imple-\nment high-sensitivity volumetric calcium imaging for investigating \nfunctionally concerted neurons and recognizing circuit motifs, espe-\ncially those distributed across multiple cortical layers.\nDiscussion\nSRDTrans does not rely on any assumptions about the contrast mecha-\nnism, noise model, sample dynamics and imaging speed. Thus, it can be \nreadily extended to other biological samples and imaging modalities \n(Supplementary Fig. 20), such as membrane voltage imaging, single-\nprotein detection, light-sheet microscopy, confocal microscopy, light-\nfield microscopy and super-resolution microscopy46–51. The limitation \nof SRDTrans lies in the basic assumption that neighboring pixels should \nhave approximate structures. If the spatial sampling rate is too low to \nprovide enough redundancy, SRDTrans would fail. Another potential \nrisk is the generalization ability as the lightweight network architecture \nof SRDTrans is more suitable for specific tasks. We believe training \nspecific models for specific data is the most reliable way to use deep \nlearning for fluorescence image denoising. Therefore, a new model \nshould be trained to ensure optimal results when the imaging param-\neter, modality and sample change.\nFig. 3 | Applying SRDTrans to experimental SMLM data. a, Experimentally \nobtained single-molecule-emission images. Left: raw data. Right: SRDTrans \ndenoised data. The magnified views of two boxed regions are shown below \neach image. Scale bars, 2 μm for the whole FOV and 0.5 μm for magnified views. \nb, Reconstructed super-resolution images. The microtubules in fixed BSC-1 \ncells were labeled with Cy5. Scale bars, 2 μm for the whole FOV and 0.5 μm for \nmagnified views. c, FRC curves of the raw reconstructed image (blue) and the \nSRDTrans enhanced reconstructed image (orange). The estimated resolution \n(52.4 nm for raw image and 36.0 nm for SRDTrans denoised image) is the inverse \nof the spatial frequency where the FRC curve drops below the cut-off threshold \n(~0.143). d, Tukey box-and-whisker plots (Methods) showing the localization \nuncertainty before and after denoising (N = 1,048,575 detected molecules for \nraw data, N = 395,908 detected molecules for SRDTrans). The uncertainty was \ncalculated by the ThunderSTORM plugin. e, Single-molecule-emission images \nfrom the open-source platform ShareLoc51. Left: raw data. Right: SRDTrans \ndenoised data. Scale bar, 2 μm. f, Reconstructed super-resolution images of \nmicrotubules (immuno-labeled with Alexa 647). Scale bar, 2 μm. g, Magnified \nviews of boxed regions. Scale bar, 0.5 μm. h, Intensity profiles perpendicular \nto the microtubule filaments indicated in g. Blue line, raw data; orange \nlines, SRDTrans denoised data; dashed line, the Gaussian fitting result. The \ncorresponding FWHM value is quantified as 2.335σ, where σ denotes the standard \ndeviation of the Gaussian fitting result. i, Tukey box-and-whisker plots (Methods) \nshowing the FWHM of randomly selected filaments (blue, raw data; orange, \nSRDTrans denoised data; N = 80 independent filament segments).\nFig. 4 | Evaluating the performance of SRDTrans on simulated calcium \nimaging data. a, SRDTrans denoised calcium imaging data sampled at 30 Hz. \nMagnified views show the neural activity of the yellow boxed region in a short \nperiod (~2 s). Left: the original low-SNR data. Middle: SRDTrans denoised data. \nRight: GT. Scale bars, 60 μm for the whole FOV and 30 μm for magnified views. \nThe magenta arrowhead indicates a dendritic fiber and the yellow arrowhead \nindicates two neighboring somas. b, Pixel intensity along the yellow dashed line \nin a. Top left: raw data. Middle left: SRDTrans denoised data. Bottom left: GT. \nRight: plotting the three intensity profiles in one coordinate. The similarity with \nGT is quantified by Pearson correlation coefficients (R). c, Frequency spectrum \ncalculated by discrete Fourier transform before and after denoising. Magnified \nviews of the boxed regions show the frequency components within the optical \ntransfer function. The similarity in the frequency domain is quantified by LFD. \nd, The performance of SRDTrans at different SNR levels. All values are shown as \nmean ± s.d. (N = 6,000 independent frames). e, Comparing the performance of \nDeepCAD and SRDTrans on calcium imaging data sampled at 0.3 Hz. Magnified \nviews show the neural activity of yellow boxed regions in a 20 s time window. \nScale bars, 100 μm for the whole FOV and 40 μm for magnified views. The \nyellow and purple arrowheads point to a firing neuron and a resting neuron, \nrespectively. f, Ablation experiments to investigate the effects of different \nsampling strategies and network architectures. SRDTrans (orange) uses spatial \nredundancy sampling and a lightweight spatiotemporal transformer. DeepCAD \n(purple) combines temporal redundancy sampling and a CNN (3D U-Net). \nSRDCNN (green) is the method combining spatial redundancy sampling and \na CNN (3D U-Net). All values are shown as mean ± s.d. (N = 1,000 independent \nframes for each frame rate). g, Fluorescence traces (F) extracted from 50 \nrandomly selected neuronal pixels. The similarity with GT is quantified by \nPearson correlation coefficients (R). Top: traces extracted from raw data. Middle: \ntraces extracted from DeepCAD denoised data. Middle bottom: traces extracted \nfrom SRDTrans denoised data. Bottom: GT.\n\n\nNature Computational Science | Volume 3 | December 2023 | 1067–1080\n1074\nArticle\nhttps://doi.org/10.1038/s43588-023-00568-2\na\nSRDTrans (SNR = 26.51 dB)\nGT\nRaw data (30 Hz, SNR = 0.91 dB)\n5.20 s\n6.07 s\n4.03 s\n4.27 s\n4.43 s\n5.20 s\n6.07 s\n166.67 s\n156.67 s\n166.67 s\n156.67 s\n150\n300\n450\n600 (s)\n176.67 s\n176.67 s\nR = 0.568\nR = 0.782\nR = 0.984\nR = 1.000\n∆F/F (normalized)\n1.0\n0\n4.43 s\nNormalized intensity\n4.03 s\n4.27 s\n4.43 s\n5.20 s\n4.03 s\n4.27 s\n6.07 s\nb\nRaw data\nR = 0.024\nR = 0.996\nR = 1.000\nSRDTrans\nGT\n166.67 s\n0.5\n(normalized)\n20 s\n176.67 s\nRaw data\nSRDTrans\n0\nOutput SNR (dB)\n10\n20\n30\n40\nInput SNR (dB)\n0\n–5\n10\n5\n20\n15\nSNR= -0.85 dB\nSNR=19.88 dB\nSNR=13.32 dB\ne\nRaw data (0.3 Hz, SNR = –0.83 dB)\n1.0\n0\nSRDTrans (SNR = 19.62 dB)\nDeepCAD (SNR = 13.10 dB) \nGT\n0.1 (normalized)\n30 µm\nNormalized intensity\nd\nSRDTrans\nGT\nRaw data\nLFD = 28.43 dB\nLFD = 7.96 dB\nc\n0\n156.67 s\n166.67 s\n176.67 s\n156.67 s\nf\ng\nGT\nDeepCAD\nRaw data\nImaging speed (Hz)\nSRDTrans\nOutput SNR (dB)\n10\n–1\n10\n0\n10\n1\n10\n2\n14\n22\n26\n18\nSRDTrans\nSRDCNN\nDeepCAD\n\n\nNature Computational Science | Volume 3 | December 2023 | 1067–1080\n1075\nArticle\nhttps://doi.org/10.1038/s43588-023-00568-2\nAs the development of fluorescence indicators heads towards \nfaster kinetics to monitor biological dynamics at the millisecond \nscale52,53, the imaging speed required to record these fast activities is \ncontinuously growing. Obtaining adequate sampling rates is becoming \nmore and more challenging for denoising methods relying on temporal \nredundancy. Our rationale is to fill the niche by seeking to utilize spatial \na\nRaw data\nx\ny\nx\nz\ny\nz\nRaw data\nSRDTrans\nTime = 273.33 s\nc\nx\nz\ny\nz\nx\ny\nb\nSRDTrans\nRaw data\nx\nz\nTime = 243.33 s\n150\n100\n50\n150\n200\n250\n300\n350\n400\n450\n0\n50\n100\n150\n200\n250\n300\n350\n400\n450\n100\n50\n0\nPosition Z (µm)\nPosition X (µm)\nPosition Y (µm)\ny\nz\nx\ny\ny\nz\nx\ny\nx\nz\nSRDTrans\n1,200\n1,600\n400\n800\n0\n1,200\n1,600\n400\n800\n0\n∆F/F (normalized)\n1.0\n0\nRaw data\nTime (s)\nTime (s)\nSRDTrans\n0\n150\n100\n50\n150\n200\n250\n300\n350\n400\n450\n0\n50\n100\n150\n200\n250\n300\n350\n400\n450\n100\n50\nPosition Z (µm)\nPosition X (µm)\nPosition Y (µm)\nFig. 5 | High-sensitivity calcium imaging of large neural volumes. a, Three-\ndimensional visualization of the neural activity of a 510 × 510 × 200 μm3 volume \n(100 planes, 0.3 Hz volume rate) in the mouse cortex. Left: the original low-SNR \nvolume. Right: the same volume denoised with SRDTrans. Magnified views of \nyellow boxed regions are shown under each snapshot. Scale bars, 100 μm for the \nwhole FOV and 10 μm for magnified views. b, Raw frames and SRDTrans denoised \ncounterparts of a single imaging plane at two different moments. The x–z and \ny–z cross-sections of the volume are shown alongside each x–y plane. Magnified \nviews of yellow boxed regions are shown at the bottom right of the images. Scale \nbars, 70 μm for the whole FOV and 20 μm for magnified views. c, Fluorescence \ntraces (F) extracted from all pixels on the yellow dashed line in b. Left: traces \nextracted from raw data. Right: traces extracted from SRDTrans denoised data. \nYellow arrowheads point to some representative calcium transients.\n\n\nNature Computational Science | Volume 3 | December 2023 | 1067–1080\n1076\nArticle\nhttps://doi.org/10.1038/s43588-023-00568-2\nredundancy as an alternative to enable self-supervised denoising in \nmore imaging applications. Although the perfect case for spatial redun-\ndancy sampling is that the spatial sampling rate is two times higher \nthan the Nyquist sampling of the diffraction limit to ensure that two \nadjacent pixels have nearly identical optical signals, the endogenous \nsimilarity between two spatially downsampled substacks is sufficient \nto guide the training of the network in most cases. However, this does \nnot mean that the proposed spatial redundancy sampling strategy can \nfully replace temporal redundancy sampling, as the ablation study \n(Fig. 4f) shows that, if equipped with the same network architecture, \nthe temporal redundancy sampling can achieve better performance \nin high-speed imaging. The superiority of SRDTrans over DeepCAD at \nhigh imaging speeds is actually attributed to the transformer archi-\ntecture. In general, spatial redundancy and temporal redundancy are \ntwo complementary sampling strategies to enable self-supervised \ntraining of denoising networks for fluorescence time-lapse imaging. \nWhich sampling strategy is used depends on which kind of redundancy \nis more sufficient in the data. It is noteworthy that there are still many \ncases where neither redundancy is sufficient to support current sam-\npling strategies. Developing specific or more general self-supervised \ndenoising methods is of persistent value for fluorescence imaging.\nMethods\nThe spatial redundancy sampling strategy\nIn SRDTrans, we employed a spatial redundancy sampling strategy \nto produce training pairs. The detailed implementation for generat-\ning training pairs in SRDTrans is shown in Fig. 1b and Supplementary \nFig. 1d. For each image inside an input training stack with H × W × T \npixels (H, W and T are the height, width and length of the input image \nstack), we spatially divided it into many adjacent small patches with \n2 × 2 pixels. Next, we randomly selected three adjacent pixels from \neach patch. The central pixel was used to compose the input substack \nand the two spatially adjacent pixels were used to compose the two \ntarget substacks. After traversing all patches, we can finally obtain \nthree downsampled substacks with the size of H/2 × W/2 × T pixels. As \nthe fluorescence signals of spatially adjacent pixels are closely corre-\nlated, the input substack and each target substack can be considered \nas two independent samples of the same underlying pattern. Thus, \nthe input substack and the two target substacks can form two training \npairs, which can be used for the self-supervised training of denoising \nnetworks (Supplementary Note 1).\nNetwork architecture and loss function\nThe transformer architecture of SRDTrans is composed of three parts: \na temporal encoder module, a spatiotemporal transformer block (STB) \nand a temporal decoder module (Fig. 1c). The temporal encoder module \nis equipped with two temporal encoders. Each temporal encoder can \ncompress the temporal scale of the input substack by a factor of r (r is 4 \nin this work) using a convolutional layer with 3 × 3 kernels. In contrast to \nthe U-Net-type architectures with many upsampling and downsampling \noperations, SRDTrans does not reduce the size of feature maps in these \nencoders (Supplementary Fig. 21). Thus, for an input substack with a \nsize of H/2 × W/2 × T pixels, the output size of the temporal encoder \nmodule is H/2 × W/2 × T/r2. The output from the temporal encoder \nmodule will be fed into the STB to extract global information both in \nspace and time. The STB contains a temporal transformer block and \na spatial transformer block (Supplementary Fig. 22). Specifically, in \nthe temporal transformer block, the input is divided into patches \nwith the size of p × p × T/r2 (p is 7 in this work). These patches are first \nflattened into one-dimensional vectors and input into the position \nembedding layer, where spatial concatenation and linear transfor-\nmation are performed. Two multi-head self-attention layers are then \ncascaded to extract temporal correlations inside the data. In the spatial \ntransformer block, a Swin transformer block27 is adopted to capture \nfine-grained spatial features with high efficiency. Local features flow \nfully in multi-head self-attention layers and densely interact with long-\nrange global features. Finally, the output of the STB is remapped by \nthe temporal decoder module, and its temporal scale can be rescaled \nto T. This decompression operation is implemented by two cascaded \ntemporal decoders using convolutional layers with 3 × 3 kernels.\nWe used a linear combination of L1-norm loss and L2-norm loss as \nthe loss function to optimize the parameters of SRDTrans, which shows \nbetter performance than L1-norm loss and improved convergence \ncompared with L2-norm loss (Supplementary Fig. 23). We define the \ninput substack filled with central pixels as Sc, the target substack filled \nwith its vertically adjacent pixels as Sv and the target substack filled with \nits horizontally adjacent pixels as Sh. The total loss consists of two pairs \nof training losses, which is defined as:\nLver = ‖FSRDTrans(Si) −Sv‖\n2\n2 + |FSRDTrans(Si) −Sv|1,\nLhor = ‖FSRDTrans(Si) −Sh‖\n2\n2 + |FSRDTrans(Si) −Sh|1,\nLtotal = Lver + Lhor.\nwhere Lver and Lhor denote the loss of the vertically and horizontally \nadjacent substacks, respectively.\nTraining and inference\nTo achieve optimal performance, specific models were trained for \nstacks with different SNRs. One or more training stacks (xy–t or xy–z) \nwere divided into a specified number of 3D (xy–t) training pairs (6,000 \nby default). The batch size for all experiments was set to the number of \ngraphics processing units (NVIDIA GeForce RTX 3090 for most cases) \nbeing used and the patch size was set to be 128 × 128 × 128 pixels. All \nextracted training pairs were geometrically transformed by random \nflipping or rotation for eightfold data augmentation. The synergy of \nour lightweight architecture and data augmentation eliminates the risk \nof overfitting (Supplementary Fig. 24). The compression factor of each \ntemporal encoder was set to 4. In the STB, we set the internal patch size \nto 7, the number of heads in the multi-headed self-attention block to \n \n8 and the embedded feature channels to 128. For model optimization, \nwe used the Adam optimizer and the exponential decay rate for the first \nmoment was 0.9, the exponential decay rate for the second moment was \n0.999 and the learning rate was 0.00001. PyTorch was used to construct \nthe network and implement all operations. In the inference stage, the \nraw noisy data were not spatially subsampled and the model of the last \ntraining epoch was selected for final processing. The denoised result \nof each image stack was saved as a separate TIF file.\nData simulation\nQuantitative evaluations were performed on simulated data because \nnoise-free images (GT) are available. To synthesize noise-free two-\nphoton calcium imaging data, we used NAOMi44, which can generate \nrealistic calcium imaging data with high-fidelity tissue characteristics \nand fluorescence kinetics. Then we applied different levels of mixed \nPoisson–Gaussian noise to generate calcium imaging data of differ-\nent imaging SNRs6,18. We also simulated data containing only Poisson \nor Gaussian noise to show the comparable denoising performance of \nSRDTrans on these two types of noise (Supplementary Fig. 25). To gen-\nerate calcium imaging data of different sampling rate, we first synthe-\nsized images sampled at 30 Hz and 1 Hz, and the data of other sampling \nrates were obtained by extracting frames at different intervals. The \nimage size for all simulated calcium imaging data was 490 × 490 pixels \nand the pixel size was 1.02 μm.\nTo generate realistic SMLM data, we first acquired reconstructed \nsuper-resolution images from the ShareLoc.XYZ platform (https://\nshareloc.xyz/)42. These images were experimentally obtained on a \nNikon N-STORM microscope and contained densely distributed micro-\ntubules immuno-labeled with Alexa 64754. The tracks of all microtubules \nin a selected region of interest were extracted semi-automatically using \n\n\nNature Computational Science | Volume 3 | December 2023 | 1067–1080\n1077\nArticle\nhttps://doi.org/10.1038/s43588-023-00568-2\nthe JFilament plugin of ImageJ55. We then generated synthetic single-\nmolecule-emission image stacks (GT images without noise) using the \nTestSTORM36 simulator by loading the microtubule patterns from \nJFilament. All fluorescent molecules were linked on the microtubule \npattern with a radius of 12.5 nm. For imaging parameters, the numerical \naperture was 1.4 and the frame rate was 200 Hz (5 ms exposure time). \nThe image size was 328 × 328 pixels and the pixel size was 30 nm. Noisy \nstacks were generated by applying mixed Poisson–Gaussian noise post \nhoc with a customized MATLAB script6,18.\nWe synthesized moving objects with different moving speeds \nusing the Modified National Institute of Standards and Technology \n(MNIST) dataset56. Each frame was defined as an image of 512 × 512 \npixels with a black background and many bright moving digits. Each \ndigit was an image patch (28 × 28 pixels) randomly extracted from the \nMNIST dataset, moved in a random direction, and appeared or disap-\npeared only once. The total number of digits in the field of view (FOV) \nwas 500. We first generated the data with a moving speed of 0.5 pixels \nper frame. The data with higher moving speeds were then generated \nby extracting frames at different intervals. The total frame number for \nall moving speeds was 5,000. The final experiment was implemented \non 20 datasets with moving speeds from 0.5 to 10 pixels per frame. \nThe sampling interval of the moving speed was 0.5 pixels per frame.\nSMLM imaging\nThe imaging samples (including the buffer solution and the sam-\nple holder) for the SMLM experiments were purchased from \nStandard Imaging Company (https://www.standardimaging.cn/\nstandardsample?lang=en). The SMLM experiments were performed \non a commercial microscope (Nikon N-STORM) equipped with laser \nsources of 405 nm and 640 nm, which were used for activation and \nexcitation, respectively. A scientific complementary metal-oxide semi-\nconductor camera (Hamamatsu Flash 4.0) was placed at the image \nplane to capture the emission signals. To mimic living-cell imaging, we \nused low excitation power to reduce phototoxicity and short exposure \ntime to obtain images with low emitter density. The detailed settings \nare summarized in Supplementary Table 4.\nSMLM sample preparation\nThe Biologics Standards-Cercopithecus-1 (BSC-1) cell line purchased \nfrom Pricella Life Technology was used for SMLM imaging. BSC-1 \ncells were cultured in DMEM (Invitrogen, 11965-118) supplemented \nwith 10% fetal bovine serum (Gibco, 16010-159). To prevent bacterial \ncontamination, 100 μg ml−1 penicillin and streptomycin (Invitrogen, \n15140122) were added into the DMEM medium. Cells were grown under \nstandard cell culture conditions (5% CO2, humidified atmosphere at \n37 °C). BSC-1 cells were plated on 1.5 glass-bottom dishes over 48 h \nbefore sample preparation. For cell passage, cells were washed with \npre-warmed PBS (Life Technologies, 14190500BT) 3 times and digested \nwith 25% trypsin (Gibco, 25200-056) for 30 s. BSC-1 l cell lines were \ntested for potential mycoplasma contamination (MycoAlert, Lonza) \nand all tests showed negative results. For immunofluorescence stain-\ning, cells were grown on 35 mm, 1.5 glass coverslips. We pre-treated \nglass-bottom dishes with fibronectin (Invitrogen, 33016015) for 1 h at \n37 °C to increase cell adhesion. On the day of sample preparation, the \ncell density should be about 50–70%. Cells were fixed with 37 °C pre-\nwarmed fixation buffer for 10 min, containing 4% paraformaldehyde \n(EMS) and 0.1% glutaraldehyde in PBS. Then the sample was washed \nthree times with PBS. For quenching the background fluorescence, \nwe incubated the cells with 2 ml 0.1% NaBH4 solution in PBS for 7 min, \noptionally shaking on the shaker (<1 Hz). The sample was washed 3 \ntimes with 2 ml PBS and then incubated for 30 min in PBS containing 5% \nBSA (Jackson, 001-000-162) and 0.5% Triton X-100 (Fisher Scientific) \nat 37 °C. All antibodies were diluted in the 5% BSA + 0.5% triton solution \ndescribed above. Next, we incubated the sample for 40 min with the \nappropriate dilution of primary antibodies: mouse anti-beta-tubulin \n(E7, DSHB) at 25 °C. After primary antibodies incubation, the cells \nwere washed 3 times with 2 ml PBS for 5 min. Secondary antibodies \n(AffiniPure Donkey Anti-mouse IgG, 715-005-150, Jackson Immuno \nResearch) were incubated for 60 min with the appropriate dilutions \nof secondary antibodies (conjugated with Cy5) at 25 °C. After being \nwashed 3 times with PBS, cells were fixed with post-fixation buffer \nfor 10 min. The sample was stored at 4 °C in PBS and protected from \nlight. Before imaging, we used an imaging buffer (STIBa-031, Standard \nImaging Company) to replace PBS.\nSMLM reconstruction\nThe super-resolution SMLM images were reconstructed by the Thun-\nderSTORM37 Fiji plugin. For our experimentally obtained data, hard \nthresholding was performed to zero out those pixels with values \nsmaller than a manually adjusted threshold to suppress the patterned \nnoise of the camera. The detailed configuration is set as: the image filter \nwas the wavelet filter (B-spline) with an order of 3 and a scale of 2; the \nalgorithm for determining the approximate position of molecules was \nthe local maximum algorithm; the subpixel localization is performed \nby the integrated Gaussian point-spread-function model with a fitting \nradius of 3 pixels; a fitting method of ‘weighted least squares’, and an \ninitial sigma of 1.6 pixels. Both visualization images are generated \nby averaged shifted histogram with a magnification of 5. For better \nvisualization, the single-molecule-emission images and reconstructed \nsuper-resolution images were rendered with pseudo-color and their \ncontrast and brightness were manually adjusted.\nMouse preparation and calcium imaging\nAll experiments involving mice were performed in accordance with the \ninstitutional guidelines for animal welfare and have been approved by \nthe Animal Care and Use Committee of Tsinghua University. All mice \nwere aged 8–12 weeks and were housed in cages (24 °C, 50% humidity) \nin groups of 1–5 under a reverse light cycle. Transgenic mice hybridized \nbetween Rasgrf2-2A-dCre mice and Ai148 (TIT2L-GC6f-ICL-tTA2)-D \nmice expressing Cre-dependent GCaMP6f genetically encoded cal-\ncium indicator were used for calcium imaging of neural circuits. Both \nmale and female mice were used without randomization or blinding. \nCraniotomy surgeries were conducted to remove the skull and a cov-\nerslip was implanted on the craniotomy region for chronic imaging. \nTwo-photon volumetric imaging of the mouse cortex was performed \non head-fixed mice without anesthesia using a standard two-photon \nmicroscope controlled with ScanImage 5.7. The neural volume being \nrecorded was located at the primary visual cortex with a depth of about \n150–350 μm below the dura, and was scanned for 100 planes with an \naxial step of 2 μm. The whole imaging session lasted 30 min with a \nvolume rate of 0.3 Hz.\nThree-dimensional visualization of neural activity\nFor volumetric calcium imaging, we used Imaris 9.0 (Oxford Instru-\nments) to visualize the calcium activity of the neuronal population \nbefore and after denoising. Specifically, we imported the four-dimen-\nsional (xyz–t) data into Imaris, applied pseudo-color to the images, and \nthen performed 3D rendering using the maximum intensity projection \nmode. The contrast and brightness were adjusted to make structures \nin the volume as clear as possible. All values for gamma correction \nwere set to one.\nMethod comparison\nWe compared the performance of SRDTrans with six baseline self-\nsupervised methods: Noise2Noise16, Noise2Void19, Noise2Self20, Proba-\nbilistic Noise2Void21, Neighbor2Neighbor22, DeepInterpolation17 and \nDeepCAD6,18. These methods were all implemented by open-source \ncodes released by the relevant papers. The denoising model of each \nmethod was trained and tested on the same datasets. For the methods \ndesigned for two-dimensional images, we split the time-lapse (xy–t) \n\n\nNature Computational Science | Volume 3 | December 2023 | 1067–1080\n1078\nArticle\nhttps://doi.org/10.1038/s43588-023-00568-2\nimage stack into a series of two-dimensional frames to match the input \ndimension. Training and inference were performed frame by frame. \nWe followed the default training settings about network architec-\ntures and hyperparameters for all methods. Specifically, the model of \nDeepInterpolation was fine-tuned based on a public pre-trained model \n(pre-trained with 225,000 two-photon images of the Ai93 reporter \nline). Other methods were trained from scratch. The detailed settings \nof each method are listed in Supplementary Table 5.\nEvaluation metrics\nWe used several metrics to evaluate the performance of different \ndenoising methods. For an image (or an image stack) Sx and its GT Sy, \nthe metrics are defined as follows.\nSNR measures the pixel-level deviation between two images using \nthe logarithmic decibel scale, which is formulated as\nSNR = 10 log10\n‖\n‖Sy‖\n‖\n2\n2\n‖\n‖Sx −Sy‖\n‖\n2\n2\n.\nSSIM measures the similarity between two images on a perceptual \nlevel, including luminance, contrast and structure. The definition is\nSSIM =\n(2μxμy + c1)(2σxy + c2)\n( μ2\nx + μ2\ny + c1)(σ2\nx + σ2\ny + c2)\n,\nwhere {μx, μy} and {σx, σy} are the means and variances of Sx and Sy, \nrespectively. σxy is the covariance of Sx and Sy. The two constants c1 \nand c2 are defined as c1 = (k1L)2 and c2 = (k2L)2 with k1 = 0.01, k2 = 0.03 \nand L = 65,535.\nThe Jaccard index measures the proportion of correctly detected \nmolecules in SMLM. The correctly localized fluorescent molecules are \ntrue positives (TP). The incorrectly localized molecules are false posi-\ntives (FP) and the undetected molecules are false negatives (FN). The \nJaccard index is formulated as:\nJaccard = 100\nTP\nTP + FP + FN %.\nThe r.m.s.e. quantifies the mean difference between the local-\nized positions (Px) and GT positions (Py) of all detected fluorescent \nmolecules:\nr.m.s.e. =\n√\n√\n√\n1\nTP ∑\nTP\n‖\n‖Py −Px‖\n‖\n2\n2,\nEfficiency (E) is a comprehensive metric combining the Jaccard \nindex and r.m.s.e. to measure the performance of single-molecule \nlocalization39. It can simultaneously reflect the ability to detect mol-\necules from images (measured by Jaccard) and the ability to precisely \nlocate molecules (measured by r.m.s.e.), which is defined as:\nEfficiency = 100 −√(100 −Jaccard)\n2 + α2r.m.s.e.2,\nwhere α = 1 nm−1 controls the trade-off between Jaccard and r.m.s.e.\nThe Pearson correlation coefficient measures the similarity \nbetween a variable (images and fluorescence traces) and its GT, which \nis formulated as\nR =\nE[(Sx −μx)(Sy −μy)]\nσxσy\n,\nwhere E represents the arithmetic mean. {μx, μy} and {σx, σy} are the \nmeans and variances of Sx and Sy, respectively.\nLogarithmic frequency distance (LFD) quantifies the spectral \ndifference between two images in the frequency domain. For images \nwith a size of H × W pixels, LFD is formulated as:\nLFD = log10 [ 1\nHW (\nH−1\n∑\nu=0\nW−1\n∑\nv=0\n‖\n‖FSx(u, v) −FSy(u, v)‖\n‖\n2\n2) + 1] .\nFSx and FSy are the discrete Fourier transform of Sx and Sy, respec-\ntively. u and v are the pixel index in the frequency domain.\nStatistics and reproducibility\nTo ensure the reproducibility of the findings, we report the sample \nsize and statistics in the legend and text of each experiment. All box \nplots are drawn in the standard Tukey box-and-whisker format. The \nupper and lower quartiles are represented by box bounds, and the \nlines in the boxes indicate the median. The lower whisker represents \nthe minimum observed value, equal to the lower quartile minus 1.5× \nthe interquartile range. The upper whisker the maximum observed \nvalue, equal to the upper quartile plus 1.5× the interquartile range. \nResults obtained through experimental or observational studies or \nstatistical analysis of datasets can be reproduced with high reliability \nwhen the study is repeated. Representative images are shown in figures \nand similar results are achieved on all test samples. Experiments in Figs. \n1d,e and 4a were repeated with 6,000 frames. Experiments in Figs. 2a \nand 3a,e were repeated with 24,000, 180,000 and 60,000 frames, \nrespectively. Experiments in Figs. 4e and 5b were repeated with 1,000 \nand 548 frames, respectively.\nReporting summary\nFurther information on research design is available in the Nature Port-\nfolio Reporting Summary linked to this article.\nData availability\nBoth simulated and experimentally obtained data of two-photon cal-\ncium imaging and single-molecule localization microscopy used in this \nwork can be found at https://github.com/cabooster/SRDTrans/tree/\nmain/datasets (refs. 57–60). Source data are provided with this paper.\nCode availability\nThe open-source Python code of SRDTrans is available at the Zenodo \nrepository61 and on GitHub (https://github.com/cabooster/SRDTrans).\nReferences\n1.\t\nRoyer, L. A. et al. Adaptive light-sheet microscopy for long-term, \nhigh-resolution imaging in living organisms. Nat. Biotechnol. 34, \n1267–1278 (2016).\n2.\t\nFan, J. et al. Video-rate imaging of biological dynamics at \ncentimetre scale and micrometre resolution. Nat. Photon. 13, \n809–816 (2019).\n3.\t\nBalzarotti, F. et al. Nanometer resolution imaging and tracking of \nfluorescent molecules with minimal photon fluxes. Science 355, \n606–612 (2017).\n4.\t\nWu, J. et al. Iterative tomography with digital adaptive optics \npermits hour-long intravital observation of 3D subcellular \ndynamics at millisecond scale. Cell 184, 3318–3332 (2021).\n5.\t\nVerweij, F. J. et al. The power of imaging to understand \nextracellular vesicle biology in vivo. Nat. Methods 18,  \n1013–1026 (2021).\n6.\t\nLi, X. et al. Real-time denoising enables high-sensitivity \nfluorescence time-lapse imaging beyond the shot-noise limit.  \nNat. Biotechnol. 41, 282–292 (2023).\n7.\t\nMeiniel, W., Olivo-Marin, J. C. & Angelini, E. D. Denoising of \nmicroscopy images: a review of the state-of-the-art, and a  \nnew sparsity-based method. IEEE Trans. Image Process. 27, \n3842–3856 (2018).\n\n\nNature Computational Science | Volume 3 | December 2023 | 1067–1080\n1079\nArticle\nhttps://doi.org/10.1038/s43588-023-00568-2\n8.\t\nDabov, K., Foi, A., Katkovnik, V. & Egiazarian, K. Image denoising \nby sparse 3-D transform-domain collaborative filtering. IEEE \nTrans. Image Process. 16, 2080–2095 (2007).\n9.\t\nZhang, K. et al. Beyond a Gaussian denoiser: residual learning of \ndeep CNN for image denoising. IEEE Trans. Image Process. 26, \n3142–3155 (2017).\n10.\t Tai, Y., Yang, J., Liu, X. & Xu, C. MemNet: a persistent memory \nnetwork for image restoration. In Proc. IEEE/CVF Conference  \non Computer Vision and Pattern Recognition 4539–4547  \n(IEEE, 2017).\n11.\t\nWeigert, M. et al. Content-aware image restoration: pushing the \nlimits of fluorescence microscopy. Nat. Methods 15, 1090–1097 \n(2018).\n12.\t Belthangady, C. & Royer, L. A. Applications, promises, and  \npitfalls of deep learning for fluorescence image reconstruction. \nNat. Methods 16, 1215–1225 (2019).\n13.\t Chen, J. et al. Three-dimensional residual channel attention \nnetworks denoise and sharpen fluorescence microscopy image \nvolumes. Nat. Methods 18, 678–687 (2021).\n14.\t Chaudhary, S., Moon, S. & Lu, H. Fast, efficient, and accurate \nneuro-imaging denoising via supervised deep learning.  \nNat. Commun. 13, 5165 (2022).\n15.\t Wang, Z., Xie, Y. & Ji, S. Global voxel transformer networks for \naugmented microscopy. Nat. Mach. Intell. 3, 161–171 (2021).\n16.\t Lehtinen, J. et al. Noise2Noise: learning image restoration \nwithout clean data. In Proc. 35th International Conference  \non Machine Learning (eds Dy, J. & Krause, A.) 2965–2974  \n(PMLR, 2018).\n17.\t Lecoq, J. et al. Removing independent noise in systems \nneuroscience data using DeepInterpolation. Nat. Methods 18, \n1401–1408 (2021).\n18.\t Li, X. et al. Reinforcing neuron extraction and spike inference \nin calcium imaging using deep self-supervised denoising. Nat. \nMethods 18, 1395–1400 (2021).\n19.\t Krull, A., Buchholz, T.-O. & Jug, F. Noise2Void—learning denoising \nfrom single noisy images. In Proc. IEEE/CVF Conference on \nComputer Vision and Pattern Recognition 2129–2137 (IEEE, 2019).\n20.\t Batson, J. & Royer, L. Noise2Self: blind denoising by self-\nsupervision. In Proc. 36th International Conference on Machine \nLearning 524–533 (PMLR, 2019).\n21.\t Krull, A., Vičar, T., Prakash, M., Lalit, M. & Jug, F. Probabilistic \nnoise2void: unsupervised content-aware denoising. Front. \nComput. Sci. https://doi.org/10.3389/fcomp.2020.00005 (2020).\n22.\t Huang, T. et al. Neighbor2Neighbor: self-supervised denoising \nfrom single noisy images. In Proc. IEEE/CVF Conference  \non Computer Vision and Pattern Recognition 14781–14790  \n(IEEE, 2021).\n23.\t Lequyer, J. et al. A fast blind zero-shot denoiser. Nat. Mach. Intell. \n4, 953–963 (2022).\n24.\t Luo, W. et al. Understanding the effective receptive field in deep \nconvolutional neural networks. Adv. Neural Inf. Process. Syst. 29, \n4905–4913 (2016).\n25.\t Rahaman N. et al. On the spectral bias of neural networks.  \nIn International Conference on Machine Learning 5301–5310 \n(PMLR, 2019).\n26.\t Lelek, M. et al. Single-molecule localization microscopy. Nat. Rev. \nMethods Prim. 1, 39 (2021).\n27.\t Liu, Z. et al. Swin transformer: hierarchical vision transformer \nusing shifted windows. In Proc. IEEE/CVF International Conference \non Computer Vision 10012–10022 (IEEE, 2021).\n28.\t Zhou H. et al. nnFormer: interleaved transformer for volumetric \nsegmentation. Preprint at https://arxiv.org/abs/2109.03201 (2021).\n29.\t Hatamizadeh, A. et al. UNETR: transformers for 3D medical \nimage segmentation. In Proc. IEEE/CVF Winter Conference on \nApplications of Computer Vision 574–584 (IEEE, 2022).\n30.\t Hatamizadeh, A. et al. Swin UNETR: Swin transformers for \nsemantic segmentation of brain tumors in MRI images. In \nInternational MICCAI Brainlesion Workshop (eds Crimi, A. et al.) \n272–284 (Springer, 2021).\n31.\t Çiçek, Ö. et al. 3D U-Net: learning dense volumetric \nsegmentation from sparse annotation. In Medical Image \nComputing and Computer-Assisted Intervention—MICCAI 2016 \n(eds Ourselin, S. et al.) 424–432 (Springer, 2016).\n32.\t Taylor, M. A. & Bowen, W. P. Quantum metrology and its \napplication in biology. Phys. Rep. 615, 1–59 (2016).\n33.\t Nagata, T. et al. Beating the standard quantum limit with four-\nentangled photons. Science 316, 726–729 (2007).\n34.\t Rust, M., Bates, M. & Zhuang, X. Sub-diffraction-limit imaging \nby stochastic optical reconstruction microscopy (STORM). Nat. \nMethods 3, 793–796 (2006).\n35.\t Nehme, E., Weiss, L. E., Michaeli, T. & Shechtman, Y. Deep-STORM: \nsuper-resolution single-molecule microscopy by deep learning. \nOptica 5, 458–464 (2018).\n36.\t Sinkó, J. et al. TestSTORM: simulator for optimizing sample \nlabeling and image acquisition in localization based super-\nresolution microscopy. Biomed. Opt. Express 5, 778–787 (2014).\n37.\t Ovesný, M. et al. ThunderSTORM: a comprehensive ImageJ \nplug-in for PALM and STORM data analysis and super-resolution \nimaging. Bioinformatics 30, 2389–2390 (2014).\n38.\t Sage, D. et al. Quantitative evaluation of software packages  \nfor singlemolecule localization microscopy. Nat. Methods 12, \n717–724 (2015).\n39.\t Sage, D. et al. Super-resolution fight club: assessment of 2D \nand 3D single-molecule localization microscopy software. Nat. \nMethods 16, 387–395 (2019).\n40.\t Nieuwenhuizen, R. et al. Measuring image resolution in optical \nnanoscopy. Nat. Methods 10, 557–562 (2013).\n41.\t Descloux, A., Grußmayer, K. S. & Radenovic, A. Parameter-free \nimage resolution estimation based on decorrelation analysis.  \nNat. Methods 16, 918–924 (2019).\n42.\t Ouyang, W. et al. ShareLoc—an open platform for sharing \nlocalization microscopy data. Nat. Methods 19, 1331–1333 \n(2022).\n43.\t Jones, S. et al. Fast, three-dimensional super-resolution imaging \nof live cells. Nat. Methods 8, 499–505 (2011).\n44.\t Song, A., Gauthier, J. L., Pillow, J. W., Tank, D. W. & Charles, A. S. \nNeural anatomy and optical microscopy (NAOMi) simulation for \nevaluating calcium imaging methods. J. Neurosci. Methods 358, \n109173 (2021).\n45.\t Chen, T. W. et al. Ultrasensitive fluorescent proteins for imaging \nneuronal activity. Nature 499, 295–300 (2013).\n46.\t Zhao, Z. et al. Two-photon synthetic aperture microscopy for \nminimally invasive fast 3D imaging of native subcellular behaviors \nin deep tissue. Cell 186, 2475–2491 (2023).\n47.\t Platisa, J. et al. High-speed low-light in vivo two-photon voltage \nimaging of large neuronal populations. Nat. Methods 20,  \n1095–1103 (2023).\n48.\t Zhao, W. et al. Sparse deconvolution improves the resolution \nof live-cell super-resolution fluorescence microscopcy. Nat. \nBiotechnol. 40, 606–617 (2022).\n49.\t Dahmardeh, M. et al. Self-supervised machine learning pushes \nthe sensitivity limit in label-free detection of single proteins below \n10 kDa. Nat. Methods 20, 442–447 (2023).\n50.\t Li, X. et al. Unsupervised content-preserving transformation for \noptical microscopy. Light. Sci. Appl. 10, 44 (2021).\n51.\t Qiao, C. et al. Rationalized deep learning super-resolution \nmicroscopy for sustained live imaging of rapid subcellular \nprocesses. Nat. Biotechnol. 41, 367–377 (2023).\n52.\t Zhang, Y. et al. Fast and sensitive GCaMP calcium indicators for \nimaging neural populations. Nature 615, 884–891 (2023).\n\n\nNature Computational Science | Volume 3 | December 2023 | 1067–1080\n1080\nArticle\nhttps://doi.org/10.1038/s43588-023-00568-2\n53.\t Liu, Z. et al. Sustained deep-tissue voltage recording using  \na fast indicator evolved for two-photon microscopy. Cell 185, \n3408–3425 (2022).\n54.\t Jimenez, A., Friedl, K. & Leterrier, C. About samples, giving \nexamples: optimized single molecule localization microscopy. \nMethods 174, 100–114 (2020).\n55.\t Smith, M. B. et al. Segmentation and tracking of cytoskeletal \nfilaments using open active contours. Cytoskeleton 67,  \n693–705 (2010).\n56.\t LeCun, Y., Bottou, L., Bengio, Y. & Haffner, P. Gradient-based \nlearning applied to document recognition. Proc. IEEE 86, \n2278–2324 (1998).\n57.\t Li, X. et al. SRDTrans dataset: simulated calcium imaging data \nsampled at 30 Hz under different SNRs. Zenodo https://doi.org/ \n10.5281/zenodo.8332083 (2023).\n58.\t Li, X. et al. SRDTrans dataset: simulated calcium imaging data \nat different imaging speeds. Zenodo https://doi.org/10.5281/\nzenodo.7812544 (2023).\n59.\t Li, X. et al. SRDTrans dataset: simulated SMLM data under \ndifferent SNRs. Zenodo https://doi.org/10.5281/zenodo.7812589 \n(2023).\n60.\t Li, X. et al. SRDTrans dataset: SRDTrans dataset: experimentally \nobtained SMLM data Zenodo https://doi.org/10.5281/zenodo. \n7813184 (2023).\n61.\t Li, X. et al. Code for SRDTrans. Zenodo https://doi.org/10.5281/\nzenodo.10023889 (2023).\nAcknowledgements\nThis research was supported by the National Natural Science \nFoundation of China (62088102, 62222508, 62071272) and \nNational Key Research and Development Program of China \n(Project No. 2022YFB36066), in part by the Shenzhen Science and \nTechnology Project under Grant (CJGJZD20200617102601004, \nJCYJ20220818101001004). This work was also supported by the \nChinese Postdoctoral Foundation (BX2021159) and Shuimu Tsinghua \nScholar Program. We thank H. Hao from Standard Imaging (Beijing) \nBiotechnology Co., Ltd for providing information about SMLM \nsample preparation.\nAuthor contributions\nQ.D., H.W. and J.W. supervised this research. Q.D., H.W., J.W.  \nand X.L. conceived and initiated this project. X.L., X.H. and X.C. \ndesigned detailed implementations and performed imaging \nexperiments. X.H. and X.L. developed the Python code, performed \nsimulations and processed relevant imaging data. J.F. prepared \nsamples and provided models animals. Z.Z. gave critical support on \nthe two-photon imaging system and imaging procedures. X.H., X.L. \nand X.C. analyzed the data, prepared figures and videos. X.L., X.H., \nX.C., J.F., Z.Z. and J.W. participated in discussions about the results \nand gave valuable advice. All authors participated in the drafting of \nthe paper.","difficulty":"hard","domain":"Multi-Document QA","length":"short","question":"When it comes to large-scale volumetric calcium imaging, which statement is true?","sub_domain":"Academic"}

Source: https://huggingface.co/datasets/zai-org/LongBench-v2

initial import

Posting: /agents

GET /api/v1/write?intent=publish&task_id=063060d1-c3b0-592c-a263-5380965c5233&body={url_encoded_text}&agent_name={optional_name}&nonce={optional_random_id}
