# LongBench v2 / 66ec0874821e116aacb194a4

task_id: b0a2722f-c7f1-5506-be8c-365a7b70e03a
task_key: train--66ec0874821e116aacb194a4
task_revision_id: 3

{"choice_A":"2D Gaussian Splatting improves projection accuracy by flattening the 3D Gaussian volume into a 2D disk, which allows for multi-scale affine transformations to align the splats with surface geometry. While 3D Gaussian Splatting relies on covariance matrix factorization to adjust Gaussian scaling, 2DGS avoids the need for this factorization by using homogeneous transformations that directly operate on 2D tangent planes. This enables consistent projections across views, as the surface normals are inherently preserved by aligning the splat’s principal directions to the surface.","choice_B":"2D Gaussian Splatting resolves depth inconsistencies by applying a perspective-accurate ray-splat intersection technique, which uses a depth-weighted variance reduction process to ensure Gaussian splats align more tightly with the object’s surface. Unlike 3D Gaussian Splatting, which uses volumetric Gaussian splatting with affine projection matrices, 2DGS eliminates the distortions caused by depth gradients through the use of conic surface projections. This ensures that the splats are concentrated along the surface with minimal distortion, improving surface reconstruction accuracy across multiple views.","choice_C":"2D Gaussian Splatting handles multi-view inconsistencies by employing a depth distortion loss that compresses the variance of the splats along the ray-splat intersection points, ensuring more accurate alignment with surface geometry. 3D Gaussian Splatting relies on a volumetric projection technique, where the Gaussian’s contribution is spread over multiple depths, leading to inconsistent surface reconstruction. By utilizing a projective splat accumulation process, 2DGS ensures that the projected splats remain aligned across views, avoiding the distortions caused by the varying depth intersections in 3DGS.","choice_D":"2D Gaussian Splatting surpasses 3D Gaussian Splatting by eliminating the need for depth gradient-based normal inference. In 3DGS, normals are derived from depth gradients, which vary depending on the intersection of the Gaussian with the camera ray. 2DGS introduces principal curvature alignment, where the splat’s normal is computed directly from the Gaussian’s tangent vectors, ensuring consistent alignment with the surface normals across views. This prevents the angular discrepancies caused by varying projection planes in 3DGS, which leads to inaccurate surface reconstructions.","context":"3D Gaussian Splatting for Real-Time Radiance Field Rendering\nGround Truth\nInstantNGP (9.2  fps) \nPlenoxels (8.2 fps) \nTrain: 7min, PSNR: 22.1\nTrain: 26min, PSNR: 21.9\nMip-NeRF360 (0.071 fps) \nTrain: 48 h, PSNR: 24.3\nOurs (135  fps) \nTrain: 6 min, PSNR: 23.6\nOurs (93  fps) \nTrain: 51min, PSNR: 25.2\nFig. 1. Our method achieves real-time rendering of radiance fields with quality that equals the previous method with the best quality [Barron et al. 2022],\nwhile only requiring optimization times competitive with the fastest previous methods [Fridovich-Keil and Yu et al. 2022; Müller et al. 2022]. Key to this\nperformance is a novel 3D Gaussian scene representation coupled with a real-time differentiable renderer, which offers significant speedup to both scene\noptimization and novel view synthesis. Note that for comparable training times to InstantNGP [Müller et al. 2022], we achieve similar quality to theirs; while\nthis is the maximum quality they reach, by training for 51min we achieve state-of-the-art quality, even slightly better than Mip-NeRF360 [Barron et al. 2022].\nRadiance Field methods have recently revolutionized novel-view synthesis\nof scenes captured with multiple photos or videos. However, achieving high\nvisual quality still requires neural networks that are costly to train and ren-\nder, while recent faster methods inevitably trade off speed for quality. For\nunbounded and complete scenes (rather than isolated objects) and 1080p\nresolution rendering, no current method can achieve real-time display rates.\nWe introduce three key elements that allow us to achieve state-of-the-art\nvisual quality while maintaining competitive training times and importantly\nallow high-quality real-time (≥30 fps) novel-view synthesis at 1080p resolu-\ntion. First, starting from sparse points produced during camera calibration,\nwe represent the scene with 3D Gaussians that preserve desirable proper-\nties of continuous volumetric radiance fields for scene optimization while\navoiding unnecessary computation in empty space; Second, we perform\ninterleaved optimization/density control of the 3D Gaussians, notably opti-\nmizing anisotropic covariance to achieve an accurate representation of the\nscene; Third, we develop a fast visibility-aware rendering algorithm that\nsupports anisotropic splatting and both accelerates training and allows real-\ntime rendering. We demonstrate state-of-the-art visual quality and real-time\nrendering on several established datasets.\nCCS Concepts: • Computing methodologies →Rendering; Point-based\nmodels; Rasterization; Machine learning approaches.\n∗Both authors contributed equally to the paper.\nAuthors’ addresses: Bernhard Kerbl, bernhard.kerbl@inria.fr, Inria, Université Côte\nd’Azur, France; Georgios Kopanas, georgios.kopanas@inria.fr, Inria, Université Côte\nd’Azur, France; Thomas Leimkühler, thomas.leimkuehler@mpi-inf.mpg.de, Max-\nPlanck-Institut für Informatik, Germany; George Drettakis, george.drettakis@inria.fr,\nInria, Université Côte d’Azur, France.\nPublication rights licensed to ACM. ACM acknowledges that this contribution was\nauthored or co-authored by an employee, contractor or affiliate of a national govern-\nment. As such, the Government retains a nonexclusive, royalty-free right to publish or\nreproduce this article, or to allow others to do so, for Government purposes only.\n© 2018 Copyright held by the owner/author(s). Publication rights licensed to ACM.\n0730-0301/2018/0-ART0 $15.00\nhttps://doi.org/XXXXXXX.XXXXXXX\nAdditional Key Words and Phrases: novel view synthesis, radiance fields, 3D\ngaussians, real-time rendering\nACM Reference Format:\nBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Dret-\ntakis. 2018. 3D Gaussian Splatting for Real-Time Radiance Field Rendering.\nACM Trans. Graph. 0, 0, Article 0 ( 2018), 14 pages. https://doi.org/XXXXXXX.\nXXXXXXX\n1\nINTRODUCTION\nMeshes and points are the most common 3D scene representations\nbecause they are explicit and are a good fit for fast GPU/CUDA-based\nrasterization. In contrast, recent Neural Radiance Field (NeRF) meth-\nods build on continuous scene representations, typically optimizing\na Multi-Layer Perceptron (MLP) using volumetric ray-marching for\nnovel-view synthesis of captured scenes. Similarly, the most efficient\nradiance field solutions to date build on continuous representations\nby interpolating values stored in, e.g., voxel [Fridovich-Keil and Yu\net al. 2022] or hash [Müller et al. 2022] grids or points [Xu et al. 2022].\nWhile the continuous nature of these methods helps optimization,\nthe stochastic sampling required for rendering is costly and can\nresult in noise. We introduce a new approach that combines the best\nof both worlds: our 3D Gaussian representation allows optimization\nwith state-of-the-art (SOTA) visual quality and competitive training\ntimes, while our tile-based splatting solution ensures real-time ren-\ndering at SOTA quality for 1080p resolution on several previously\npublished datasets [Barron et al. 2022; Hedman et al. 2018; Knapitsch\net al. 2017] (see Fig. 1).\nOur goal is to allow real-time rendering for scenes captured with\nmultiple photos, and create the representations with optimization\ntimes as fast as the most efficient previous methods for typical\nreal scenes. Recent methods achieve fast training [Fridovich-Keil\nACM Trans. Graph., Vol. 42, No. 4, Article . Publication date: August 2023.\n\n\n2\n•\nBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis\nand Yu et al. 2022; Müller et al. 2022], but struggle to achieve the\nvisual quality obtained by the current SOTA NeRF methods, i.e.,\nMip-NeRF360 [Barron et al. 2022], which requires up to 48 hours of\ntraining time. The fast – but lower-quality – radiance field methods\ncan achieve interactive rendering times depending on the scene\n(10-15 frames per second), but fall short of real-time rendering at\nhigh resolution.\nOur solution builds on three main components. We first intro-\nduce 3D Gaussians as a flexible and expressive scene representation.\nWe start with the same input as previous NeRF-like methods, i.e.,\ncameras calibrated with Structure-from-Motion (SfM) [Snavely et al.\n2006] and initialize the set of 3D Gaussians with the sparse point\ncloud produced for free as part of the SfM process. In contrast to\nmost point-based solutions that require Multi-View Stereo (MVS)\ndata [Aliev et al. 2020; Kopanas et al. 2021; Rückert et al. 2022], we\nachieve high-quality results with only SfM points as input. Note\nthat for the NeRF-synthetic dataset, our method achieves high qual-\nity even with random initialization. We show that 3D Gaussians\nare an excellent choice, since they are a differentiable volumetric\nrepresentation, but they can also be rasterized very efficiently by\nprojecting them to 2D, and applying standard 𝛼-blending, using an\nequivalent image formation model as NeRF. The second component\nof our method is optimization of the properties of the 3D Gaussians\n– 3D position, opacity 𝛼, anisotropic covariance, and spherical har-\nmonic (SH) coefficients – interleaved with adaptive density control\nsteps, where we add and occasionally remove 3D Gaussians during\noptimization. The optimization procedure produces a reasonably\ncompact, unstructured, and precise representation of the scene (1-5\nmillion Gaussians for all scenes tested). The third and final element\nof our method is our real-time rendering solution that uses fast GPU\nsorting algorithms and is inspired by tile-based rasterization, fol-\nlowing recent work [Lassner and Zollhofer 2021]. However, thanks\nto our 3D Gaussian representation, we can perform anisotropic\nsplatting that respects visibility ordering – thanks to sorting and 𝛼-\nblending – and enable a fast and accurate backward pass by tracking\nthe traversal of as many sorted splats as required.\nTo summarize, we provide the following contributions:\n• The introduction of anisotropic 3D Gaussians as a high-quality,\nunstructured representation of radiance fields.\n• An optimization method of 3D Gaussian properties, inter-\nleaved with adaptive density control that creates high-quality\nrepresentations for captured scenes.\n• A fast, differentiable rendering approach for the GPU, which\nis visibility-aware, allows anisotropic splatting and fast back-\npropagation to achieve high-quality novel view synthesis.\nOur results on previously published datasets show that we can opti-\nmize our 3D Gaussians from multi-view captures and achieve equal\nor better quality than the best quality previous implicit radiance\nfield approaches. We also can achieve training speeds and quality\nsimilar to the fastest methods and importantly provide the first\nreal-time rendering with high quality for novel-view synthesis.\n2\nRELATED WORK\nWe first briefly overview traditional reconstruction, then discuss\npoint-based rendering and radiance field work, discussing their\nsimilarity; radiance fields are a vast area, so we focus only on directly\nrelated work. For complete coverage of the field, please see the\nexcellent recent surveys [Tewari et al. 2022; Xie et al. 2022].\n2.1\nTraditional Scene Reconstruction and Rendering\nThe first novel-view synthesis approaches were based on light fields,\nfirst densely sampled [Gortler et al. 1996; Levoy and Hanrahan 1996]\nthen allowing unstructured capture [Buehler et al. 2001]. The advent\nof Structure-from-Motion (SfM) [Snavely et al. 2006] enabled an\nentire new domain where a collection of photos could be used to\nsynthesize novel views. SfM estimates a sparse point cloud during\ncamera calibration, that was initially used for simple visualization\nof 3D space. Subsequent multi-view stereo (MVS) produced im-\npressive full 3D reconstruction algorithms over the years [Goesele\net al. 2007], enabling the development of several view synthesis\nalgorithms [Chaurasia et al. 2013; Eisemann et al. 2008; Hedman\net al. 2018; Kopanas et al. 2021]. All these methods re-project and\nblend the input images into the novel view camera, and use the\ngeometry to guide this re-projection. These methods produced ex-\ncellent results in many cases, but typically cannot completely re-\ncover from unreconstructed regions, or from “over-reconstruction”,\nwhen MVS generates inexistent geometry. Recent neural render-\ning algorithms [Tewari et al. 2022] vastly reduce such artifacts and\navoid the overwhelming cost of storing all input images on the GPU,\noutperforming these methods on most fronts.\n2.2\nNeural Rendering and Radiance Fields\nDeep learning techniques were adopted early for novel-view synthe-\nsis [Flynn et al. 2016; Zhou et al. 2016]; CNNs were used to estimate\nblending weights [Hedman et al. 2018], or for texture-space solutions\n[Riegler and Koltun 2020; Thies et al. 2019]. The use of MVS-based\ngeometry is a major drawback of most of these methods; in addition,\nthe use of CNNs for final rendering frequently results in temporal\nflickering.\nVolumetric representations for novel-view synthesis were ini-\ntiated by Soft3D [Penner and Zhang 2017]; deep-learning tech-\nniques coupled with volumetric ray-marching were subsequently\nproposed [Henzler et al. 2019; Sitzmann et al. 2019] building on a con-\ntinuous differentiable density field to represent geometry. Rendering\nusing volumetric ray-marching has a significant cost due to the large\nnumber of samples required to query the volume. Neural Radiance\nFields (NeRFs) [Mildenhall et al. 2020] introduced importance sam-\npling and positional encoding to improve quality, but used a large\nMulti-Layer Perceptron negatively affecting speed. The success of\nNeRF has resulted in an explosion of follow-up methods that address\nquality and speed, often by introducing regularization strategies; the\ncurrent state-of-the-art in image quality for novel-view synthesis is\nMip-NeRF360 [Barron et al. 2022]. While the rendering quality is\noutstanding, training and rendering times remain extremely high;\nwe are able to equal or in some cases surpass this quality while\nproviding fast training and real-time rendering.\nThe most recent methods have focused on faster training and/or\nrendering mostly by exploiting three design choices: the use of spa-\ntial data structures to store (neural) features that are subsequently\ninterpolated during volumetric ray-marching, different encodings,\nACM Trans. Graph., Vol. 42, No. 4, Article . Publication date: August 2023.\n\n\n3D Gaussian Splatting for Real-Time Radiance Field Rendering\n•\n3\nand MLP capacity. Such methods include different variants of space\ndiscretization [Chen et al. 2022b,a; Fridovich-Keil and Yu et al. 2022;\nGarbin et al. 2021; Hedman et al. 2021; Reiser et al. 2021; Takikawa\net al. 2021; Wu et al. 2022; Yu et al. 2021], codebooks [Takikawa\net al. 2022], and encodings such as hash tables [Müller et al. 2022],\nallowing the use of a smaller MLP or foregoing neural networks\ncompletely [Fridovich-Keil and Yu et al. 2022; Sun et al. 2022].\nMost notable of these methods are InstantNGP [Müller et al. 2022]\nwhich uses a hash grid and an occupancy grid to accelerate compu-\ntation and a smaller MLP to represent density and appearance; and\nPlenoxels [Fridovich-Keil and Yu et al. 2022] that use a sparse voxel\ngrid to interpolate a continuous density field, and are able to forgo\nneural networks altogether. Both rely on Spherical Harmonics: the\nformer to represent directional effects directly, the latter to encode\nits inputs to the color network. While both provide outstanding\nresults, these methods can still struggle to represent empty space\neffectively, depending in part on the scene/capture type. In addition,\nimage quality is limited in large part by the choice of the structured\ngrids used for acceleration, and rendering speed is hindered by the\nneed to query many samples for a given ray-marching step. The un-\nstructured, explicit GPU-friendly 3D Gaussians we use achieve faster\nrendering speed and better quality without neural components.\n2.3\nPoint-Based Rendering and Radiance Fields\nPoint-based methods efficiently render disconnected and unstruc-\ntured geometry samples (i.e., point clouds) [Gross and Pfister 2011].\nIn its simplest form, point sample rendering [Grossman and Dally\n1998] rasterizes an unstructured set of points with a fixed size, for\nwhich it may exploit natively supported point types of graphics APIs\n[Sainz and Pajarola 2004] or parallel software rasterization on the\nGPU [Laine and Karras 2011; Schütz et al. 2022]. While true to the\nunderlying data, point sample rendering suffers from holes, causes\naliasing, and is strictly discontinuous. Seminal work on high-quality\npoint-based rendering addresses these issues by “splatting” point\nprimitives with an extent larger than a pixel, e.g., circular or elliptic\ndiscs, ellipsoids, or surfels [Botsch et al. 2005; Pfister et al. 2000; Ren\net al. 2002; Zwicker et al. 2001b].\nThere has been recent interest in differentiable point-based render-\ning techniques [Wiles et al. 2020; Yifan et al. 2019]. Points have been\naugmented with neural features and rendered using a CNN [Aliev\net al. 2020; Rückert et al. 2022] resulting in fast or even real-time\nview synthesis; however they still depend on MVS for the initial\ngeometry, and as such inherit its artifacts, most notably over- or\nunder-reconstruction in hard cases such as featureless/shiny areas\nor thin structures.\nPoint-based 𝛼-blending and NeRF-style volumetric rendering\nshare essentially the same image formation model. Specifically, the\ncolor 𝐶is given by volumetric rendering along a ray:\n𝐶=\n𝑁\n∑︁\n𝑖=1\n𝑇𝑖(1 −exp(−𝜎𝑖𝛿𝑖))c𝑖\nwith 𝑇𝑖= exp ©\n­\n«\n−\n𝑖−1\n∑︁\n𝑗=1\n𝜎𝑗𝛿𝑗ª\n®\n¬\n,\n(1)\nwhere samples of density 𝜎, transmittance 𝑇, and color c are taken\nalong the ray with intervals 𝛿𝑖. This can be re-written as\n𝐶=\n𝑁\n∑︁\n𝑖=1\n𝑇𝑖𝛼𝑖c𝑖,\n(2)\nwith\n𝛼𝑖= (1 −exp(−𝜎𝑖𝛿𝑖)) and 𝑇𝑖=\n𝑖−1\nÖ\n𝑗=1\n(1 −𝛼𝑖).\nA typical neural point-based approach (e.g., [Kopanas et al. 2022,\n2021]) computes the color 𝐶of a pixel by blending N ordered points\noverlapping the pixel:\n𝐶=\n∑︁\n𝑖∈N\n𝑐𝑖𝛼𝑖\n𝑖−1\nÖ\n𝑗=1\n(1 −𝛼𝑗),\n(3)\nwhere c𝑖is the color of each point and 𝛼𝑖is given by evaluating a\n2D Gaussian with covariance Σ [Yifan et al. 2019] multiplied with a\nlearned per-point opacity.\nFrom Eq. 2 and Eq. 3, we can clearly see that the image formation\nmodel is the same. However, the rendering algorithm is very differ-\nent. NeRFs are a continuous representation implicitly representing\nempty/occupied space; expensive random sampling is required to\nfind the samples in Eq. 2 with consequent noise and computational\nexpense. In contrast, points are an unstructured, discrete represen-\ntation that is flexible enough to allow creation, destruction, and\ndisplacement of geometry similar to NeRF. This is achieved by opti-\nmizing opacity and positions, as shown by previous work [Kopanas\net al. 2021], while avoiding the shortcomings of a full volumetric\nrepresentation.\nPulsar [Lassner and Zollhofer 2021] achieves fast sphere rasteri-\nzation which inspired our tile-based and sorting renderer. However,\ngiven the analysis above, we want to maintain (approximate) con-\nventional 𝛼-blending on sorted splats to have the advantages of vol-\numetric representations: Our rasterization respects visibility order\nin contrast to their order-independent method. In addition, we back-\npropagate gradients on all splats in a pixel and rasterize anisotropic\nsplats. These elements all contribute to the high visual quality of\nour results (see Sec. 7.3). In addition, previous methods mentioned\nabove also use CNNs for rendering, which results in temporal in-\nstability. Nonetheless, the rendering speed of Pulsar [Lassner and\nZollhofer 2021] and ADOP [Rückert et al. 2022] served as motivation\nto develop our fast rendering solution.\nWhile focusing on specular effects, the diffuse point-based ren-\ndering track of Neural Point Catacaustics [Kopanas et al. 2022]\novercomes this temporal instability by using an MLP, but still re-\nquired MVS geometry as input. The most recent method [Zhang\net al. 2022] in this category does not require MVS, and also uses\nSH for directions; however, it can only handle scenes of one object\nand needs masks for initialization. While fast for small resolutions\nand low point counts, it is unclear how it can scale to scenes of\ntypical datasets [Barron et al. 2022; Hedman et al. 2018; Knapitsch\net al. 2017]. We use 3D Gaussians for a more flexible scene rep-\nresentation, avoiding the need for MVS geometry and achieving\nreal-time rendering thanks to our tile-based rendering algorithm\nfor the projected Gaussians.\nACM Trans. Graph., Vol. 42, No. 4, Article . Publication date: August 2023.\n\n\n4\n•\nBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis\nA recent approach [Xu et al. 2022] uses points to represent a\nradiance field with a radial basis function approach. They employ\npoint pruning and densification techniques during optimization, but\nuse volumetric ray-marching and cannot achieve real-time display\nrates.\nIn the domain of human performance capture, 3D Gaussians have\nbeen used to represent captured human bodies [Rhodin et al. 2015;\nStoll et al. 2011]; more recently they have been used with volumetric\nray-marching for vision tasks [Wang et al. 2023]. Neural volumetric\nprimitives have been proposed in a similar context [Lombardi et al.\n2021]. While these methods inspired the choice of 3D Gaussians as\nour scene representation, they focus on the specific case of recon-\nstructing and rendering a single isolated object (a human body or\nface), resulting in scenes with small depth complexity. In contrast,\nour optimization of anisotropic covariance, our interleaved optimiza-\ntion/density control, and efficient depth sorting for rendering allow\nus to handle complete, complex scenes including background, both\nindoors and outdoors and with large depth complexity.\n3\nOVERVIEW\nThe input to our method is a set of images of a static scene, together\nwith the corresponding cameras calibrated by SfM [Schönberger\nand Frahm 2016] which produces a sparse point cloud as a side-\neffect. From these points we create a set of 3D Gaussians (Sec. 4),\ndefined by a position (mean), covariance matrix and opacity 𝛼, that\nallows a very flexible optimization regime. This results in a reason-\nably compact representation of the 3D scene, in part because highly\nanisotropic volumetric splats can be used to represent fine structures\ncompactly. The directional appearance component (color) of the\nradiance field is represented via spherical harmonics (SH), following\nstandard practice [Fridovich-Keil and Yu et al. 2022; Müller et al.\n2022]. Our algorithm proceeds to create the radiance field represen-\ntation (Sec. 5) via a sequence of optimization steps of 3D Gaussian\nparameters, i.e., position, covariance, 𝛼and SH coefficients inter-\nleaved with operations for adaptive control of the Gaussian density.\nThe key to the efficiency of our method is our tile-based rasterizer\n(Sec. 6) that allows 𝛼-blending of anisotropic splats, respecting visi-\nbility order thanks to fast sorting. Out fast rasterizer also includes\na fast backward pass by tracking accumulated 𝛼values, without a\nlimit on the number of Gaussians that can receive gradients. The\noverview of our method is illustrated in Fig. 2.\n4\nDIFFERENTIABLE 3D GAUSSIAN SPLATTING\nOur goal is to optimize a scene representation that allows high-\nquality novel view synthesis, starting from a sparse set of (SfM)\npoints without normals. To do this, we need a primitive that inherits\nthe properties of differentiable volumetric representations, while\nat the same time being unstructured and explicit to allow very fast\nrendering. We choose 3D Gaussians, which are differentiable and\ncan be easily projected to 2D splats allowing fast 𝛼-blending for\nrendering.\nOur representation has similarities to previous methods that use\n2D points [Kopanas et al. 2021; Yifan et al. 2019] and assume each\npoint is a small planar circle with a normal. Given the extreme\nsparsity of SfM points it is very hard to estimate normals. Similarly,\noptimizing very noisy normals from such an estimation would be\nvery challenging. Instead, we model the geometry as a set of 3D\nGaussians that do not require normals. Our Gaussians are defined\nby a full 3D covariance matrix Σ defined in world space [Zwicker\net al. 2001a] centered at point (mean) 𝜇:\n𝐺(𝑥) = 𝑒−1\n2 (𝑥)𝑇Σ−1(𝑥)\n(4)\n. This Gaussian is multiplied by 𝛼in our blending process.\nHowever, we need to project our 3D Gaussians to 2D for rendering.\nZwicker et al. [2001a] demonstrate how to do this projection to\nimage space. Given a viewing transformation 𝑊the covariance\nmatrix Σ′ in camera coordinates is given as follows:\nΣ′ = 𝐽𝑊Σ 𝑊𝑇𝐽𝑇\n(5)\nwhere 𝐽is the Jacobian of the affine approximation of the projective\ntransformation. Zwicker et al. [2001a] also show that if we skip the\nthird row and column of Σ′, we obtain a 2×2 variance matrix with\nthe same structure and properties as if we would start from planar\npoints with normals, as in previous work [Kopanas et al. 2021].\nAn obvious approach would be to directly optimize the covariance\nmatrix Σ to obtain 3D Gaussians that represent the radiance field.\nHowever, covariance matrices have physical meaning only when\nthey are positive semi-definite. For our optimization of all our pa-\nrameters, we use gradient descent that cannot be easily constrained\nto produce such valid matrices, and update steps and gradients can\nvery easily create invalid covariance matrices.\nAs a result, we opted for a more intuitive, yet equivalently ex-\npressive representation for optimization. The covariance matrix Σ\nof a 3D Gaussian is analogous to describing the configuration of an\nellipsoid. Given a scaling matrix 𝑆and rotation matrix 𝑅, we can\nfind the corresponding Σ:\nΣ = 𝑅𝑆𝑆𝑇𝑅𝑇\n(6)\nTo allow independent optimization of both factors, we store them\nseparately: a 3D vector 𝑠for scaling and a quaternion 𝑞to represent\nrotation. These can be trivially converted to their respective matrices\nand combined, making sure to normalize 𝑞to obtain a valid unit\nquaternion.\nTo avoid significant overhead due to automatic differentiation\nduring training, we derive the gradients for all parameters explicitly.\nDetails of the exact derivative computations are in appendix A.\nThis representation of anisotropic covariance – suitable for op-\ntimization – allows us to optimize 3D Gaussians to adapt to the\ngeometry of different shapes in captured scenes, resulting in a fairly\ncompact representation. Fig. 3 illustrates such cases.\n5\nOPTIMIZATION WITH ADAPTIVE DENSITY\nCONTROL OF 3D GAUSSIANS\nThe core of our approach is the optimization step, which creates\na dense set of 3D Gaussians accurately representing the scene for\nfree-view synthesis. In addition to positions 𝑝, 𝛼, and covariance\nΣ, we also optimize SH coefficients representing color 𝑐of each\nGaussian to correctly capture the view-dependent appearance of\nthe scene. The optimization of these parameters is interleaved with\nsteps that control the density of the Gaussians to better represent\nthe scene.\nACM Trans. Graph., Vol. 42, No. 4, Article . Publication date: August 2023.\n\n\n3D Gaussian Splatting for Real-Time Radiance Field Rendering\n•\n5\nDiﬀerentiable\nTile Rasterizer\nAdaptive\nDensity Control\nProjection\nInitialization\nSfM Points\n3D Gaussians\nImage\nCamera\nGradient Flow\nOperation Flow\nFig. 2. Optimization starts with the sparse SfM point cloud and creates a set of 3D Gaussians. We then optimize and adaptively control the density of this set\nof Gaussians. During optimization we use our fast tile-based renderer, allowing competitive training times compared to SOTA fast radiance field methods.\nOnce trained, our renderer allows real-time navigation for a wide variety of scenes.\nOriginal\nShrunken\nGaussians\nFig. 3. We visualize the 3D Gaussians after optimization by shrinking them\n60% (far right). This clearly shows the anisotropic shapes of the 3D Gaussians\nthat compactly represent complex geometry after optimization. Left the\nactual rendered image.\n5.1\nOptimization\nThe optimization is based on successive iterations of rendering and\ncomparing the resulting image to the training views in the captured\ndataset. Inevitably, geometry may be incorrectly placed due to the\nambiguities of 3D to 2D projection. Our optimization thus needs to\nbe able to create geometry and also destroy or move geometry if it\nhas been incorrectly positioned. The quality of the parameters of the\ncovariances of the 3D Gaussians is critical for the compactness of\nthe representation since large homogeneous areas can be captured\nwith a small number of large anisotropic Gaussians.\nWe use Stochastic Gradient Descent techniques for optimization,\ntaking full advantage of standard GPU-accelerated frameworks,\nand the ability to add custom CUDA kernels for some operations,\nfollowing recent best practice [Fridovich-Keil and Yu et al. 2022;\nSun et al. 2022]. In particular, our fast rasterization (see Sec. 6) is\ncritical in the efficiency of our optimization, since it is the main\ncomputational bottleneck of the optimization.\nWe use a sigmoid activation function for 𝛼to constrain it in\nthe [0 −1) range and obtain smooth gradients, and an exponential\nactivation function for the scale of the covariance for similar reasons.\nWe estimate the initial covariance matrix as an isotropic Gaussian\nwith axes equal to the mean of the distance to the closest three points.\nWe use a standard exponential decay scheduling technique similar\nto Plenoxels [Fridovich-Keil and Yu et al. 2022], but for positions\nonly. The loss function is L1 combined with a D-SSIM term:\nL = (1 −𝜆)L1 + 𝜆LD-SSIM\n(7)\nWe use 𝜆= 0.2 in all our tests. We provide details of the learning\nschedule and other elements in Sec. 7.1.\n5.2\nAdaptive Control of Gaussians\nWe start with the initial set of sparse points from SfM and then apply\nour method to adaptively control the number of Gaussians and their\ndensity over unit volume1, allowing us to go from an initial sparse\nset of Gaussians to a denser set that better represents the scene, and\nwith correct parameters. After optimization warm-up (see Sec. 7.1),\nwe densify every 100 iterations and remove any Gaussians that are\nessentially transparent, i.e., with 𝛼less than a threshold 𝜖𝛼.\nOur adaptive control of the Gaussians needs to populate empty\nareas. It focuses on regions with missing geometric features (“under-\nreconstruction”), but also in regions where Gaussians cover large\nareas in the scene (which often correspond to “over-reconstruction”).\nWe observe that both have large view-space positional gradients.\nIntuitively, this is likely because they correspond to regions that are\nnot yet well reconstructed, and the optimization tries to move the\nGaussians to correct this.\nSince both cases are good candidates for densification, we den-\nsify Gaussians with an average magnitude of view-space position\ngradients above a threshold 𝜏pos, which we set to 0.0002 in our tests.\nWe next present details of this process, illustrated in Fig. 4.\nFor small Gaussians that are in under-reconstructed regions, we\nneed to cover the new geometry that must be created. For this, it is\npreferable to clone the Gaussians, by simply creating a copy of the\nsame size, and moving it in the direction of the positional gradient.\nOn the other hand, large Gaussians in regions with high variance\nneed to be split into smaller Gaussians. We replace such Gaussians\nby two new ones, and divide their scale by a factor of 𝜙= 1.6 which\nwe determined experimentally. We also initialize their position by\nusing the original 3D Gaussian as a PDF for sampling.\nIn the first case we detect and treat the need for increasing both\nthe total volume of the system and the number of Gaussians, while\nin the second case we conserve total volume but increase the num-\nber of Gaussians. Similar to other volumetric representations, our\noptimization can get stuck with floaters close to the input cameras;\nin our case this may result in an unjustified increase in the Gaussian\ndensity. An effective way to moderate the increase in the number\nof Gaussians is to set the 𝛼value close to zero every 𝑁= 3000\n1Density of Gaussians should not be confused of course with density 𝜎in the NeRF\nliterature.\nACM Trans. Graph., Vol. 42, No. 4, Article . Publication date: August 2023.\n\n\n6\n•\nBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis\nUnder-\nReconstruction\nClone\nSplit\nOptimization\nContinues\n…\n…\nOptimization\nContinues\nOver-\nReconstruction\nFig. 4.\nOur adaptive Gaussian densification scheme. Top row (under-\nreconstruction): When small-scale geometry (black outline) is insufficiently\ncovered, we clone the respective Gaussian. Bottom row (over-reconstruction):\nIf small-scale geometry is represented by one large splat, we split it in two.\niterations. The optimization then increases the 𝛼for the Gaussians\nwhere this is needed while allowing our culling approach to remove\nGaussians with 𝛼less than 𝜖𝛼as described above. Gaussians may\nshrink or grow and considerably overlap with others, but we peri-\nodically remove Gaussians that are very large in worldspace and\nthose that have a big footprint in viewspace. This strategy results\nin overall good control over the total number of Gaussians. The\nGaussians in our model remain primitives in Euclidean space at all\ntimes; unlike other methods [Barron et al. 2022; Fridovich-Keil and\nYu et al. 2022], we do not require space compaction, warping or\nprojection strategies for distant or large Gaussians.\n6\nFAST DIFFERENTIABLE RASTERIZER FOR GAUSSIANS\nOur goals are to have fast overall rendering and fast sorting to allow\napproximate 𝛼-blending – including for anisotropic splats – and to\navoid hard limits on the number of splats that can receive gradients\nthat exist in previous work [Lassner and Zollhofer 2021].\nTo achieve these goals, we design a tile-based rasterizer for Gauss-\nian splats inspired by recent software rasterization approaches [Lass-\nner and Zollhofer 2021] to pre-sort primitives for an entire image\nat a time, avoiding the expense of sorting per pixel that hindered\nprevious 𝛼-blending solutions [Kopanas et al. 2022, 2021]. Our fast\nrasterizer allows efficient backpropagation over an arbitrary num-\nber of blended Gaussians with low additional memory consump-\ntion, requiring only a constant overhead per pixel. Our rasterization\npipeline is fully differentiable, and given the projection to 2D (Sec. 4)\ncan rasterize anisotropic splats similar to previous 2D splatting\nmethods [Kopanas et al. 2021].\nOur method starts by splitting the screen into 16×16 tiles, and\nthen proceeds to cull 3D Gaussians against the view frustum and\neach tile. Specifically, we only keep Gaussians with a 99% confi-\ndence interval intersecting the view frustum. Additionally, we use a\nguard band to trivially reject Gaussians at extreme positions (i.e.,\nthose with means close to the near plane and far outside the view\nfrustum), since computing their projected 2D covariance would\nbe unstable. We then instantiate each Gaussian according to the\nnumber of tiles they overlap and assign each instance a key that\ncombines view space depth and tile ID. We then sort Gaussians\nbased on these keys using a single fast GPU Radix sort [Merrill\nand Grimshaw 2010]. Note that there is no additional per-pixel or-\ndering of points, and blending is performed based on this initial\nsorting. As a consequence, our 𝛼-blending can be approximate in\nsome configurations. However, these approximations become negli-\ngible as splats approach the size of individual pixels. We found that\nthis choice greatly enhances training and rendering performance\nwithout producing visible artifacts in converged scenes.\nAfter sorting Gaussians, we produce a list for each tile by iden-\ntifying the first and last depth-sorted entry that splats to a given\ntile. For rasterization, we launch one thread block for each tile. Each\nblock first collaboratively loads packets of Gaussians into shared\nmemory and then, for a given pixel, accumulates color and 𝛼values\nby traversing the lists front-to-back, thus maximizing the gain in\nparallelism both for data loading/sharing and processing. When we\nreach a target saturation of 𝛼in a pixel, the corresponding thread\nstops. At regular intervals, threads in a tile are queried and the pro-\ncessing of the entire tile terminates when all pixels have saturated\n(i.e., 𝛼goes to 1). Details of sorting and a high-level overview of the\noverall rasterization approach are given in Appendix C.\nDuring rasterization, the saturation of 𝛼is the only stopping cri-\nterion. In contrast to previous work, we do not limit the number\nof blended primitives that receive gradient updates. We enforce\nthis property to allow our approach to handle scenes with an arbi-\ntrary, varying depth complexity and accurately learn them, without\nhaving to resort to scene-specific hyperparameter tuning. During\nthe backward pass, we must therefore recover the full sequence of\nblended points per-pixel in the forward pass. One solution would\nbe to store arbitrarily long lists of blended points per-pixel in global\nmemory [Kopanas et al. 2021]. To avoid the implied dynamic mem-\nory management overhead, we instead choose to traverse the per-\ntile lists again; we can reuse the sorted array of Gaussians and tile\nranges from the forward pass. To facilitate gradient computation,\nwe now traverse them back-to-front.\nThe traversal starts from the last point that affected any pixel in\nthe tile, and loading of points into shared memory again happens\ncollaboratively. Additionally, each pixel will only start (expensive)\noverlap testing and processing of points if their depth is lower than\nor equal to the depth of the last point that contributed to its color\nduring the forward pass. Computation of the gradients described in\nSec. 4 requires the accumulated opacity values at each step during\nthe original blending process. Rather than trasversing an explicit\nlist of progressively shrinking opacities in the backward pass, we\ncan recover these intermediate opacities by storing only the total\naccumulated opacity at the end of the forward pass. Specifically, each\npoint stores the final accumulated opacity 𝛼in the forward process;\nwe divide this by each point’s 𝛼in our back-to-front traversal to\nobtain the required coefficients for gradient computation.\n7\nIMPLEMENTATION, RESULTS AND EVALUATION\nWe next discuss some details of implementation, present results and\nthe evaluation of our algorithm compared to previous work and\nablation studies.\nACM Trans. Graph., Vol. 42, No. 4, Article . Publication date: August 2023.\n\n\n3D Gaussian Splatting for Real-Time Radiance Field Rendering\n•\n7\nGround Truth\nOurs\nMip-NeRF360\nInstantNGP\nPlenoxels\nFig. 5. We show comparisons of ours to previous methods and the corresponding ground truth images from held-out test views. The scenes are, from the top\ndown: Bicycle, Garden, Stump, Counter and Room from the Mip-NeRF360 dataset; Playroom, DrJohnson from the Deep Blending dataset [Hedman et al.\n2018] and Truck and Train from Tanks&Temples. Non-obvious differences in quality highlighted by arrows/insets.\nACM Trans. Graph., Vol. 42, No. 4, Article . Publication date: August 2023.\n\n\n8\n•\nBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis\nTable 1. Quantitative evaluation of our method compared to previous work, computed over three datasets. Results marked with dagger † have been directly\nadopted from the original paper, all others were obtained in our own experiments.\nDataset\nMip-NeRF360\nTanks&Temples\nDeep Blending\nMethod|Metric\n𝑆𝑆𝐼𝑀↑\n𝑃𝑆𝑁𝑅↑\n𝐿𝑃𝐼𝑃𝑆↓\nTrain\nFPS\nMem\n𝑆𝑆𝐼𝑀↑\n𝑃𝑆𝑁𝑅↑\n𝐿𝑃𝐼𝑃𝑆↓\nTrain\nFPS\nMem\n𝑆𝑆𝐼𝑀↑\n𝑃𝑆𝑁𝑅↑\n𝐿𝑃𝐼𝑃𝑆↓\nTrain\nFPS\nMem\nPlenoxels\n0.626\n23.08\n0.463\n25m49s\n6.79\n2.1GB\n0.719\n21.08\n0.379\n25m5s\n13.0\n2.3GB\n0.795\n23.06\n0.510\n27m49s\n11.2\n2.7GB\nINGP-Base\n0.671\n25.30\n0.371\n5m37s\n11.7\n13MB\n0.723\n21.72\n0.330\n5m26s\n17.1\n13MB\n0.797\n23.62\n0.423\n6m31s\n3.26\n13MB\nINGP-Big\n0.699\n25.59\n0.331\n7m30s\n9.43\n48MB\n0.745\n21.92\n0.305\n6m59s\n14.4\n48MB\n0.817\n24.96\n0.390\n8m\n2.79\n48MB\nM-NeRF360\n0.792†\n27.69†\n0.237†\n48h\n0.06\n8.6MB\n0.759\n22.22\n0.257\n48h\n0.14\n8.6MB\n0.901\n29.40\n0.245\n48h\n0.09\n8.6MB\nOurs-7K\n0.770\n25.60\n0.279\n6m25s\n160\n523MB\n0.767\n21.20\n0.280\n6m55s\n197\n270MB\n0.875\n27.78\n0.317\n4m35s\n172\n386MB\nOurs-30K\n0.815\n27.21\n0.214\n41m33s\n134\n734MB\n0.841\n23.14\n0.183\n26m54s\n154\n411MB\n0.903\n29.41\n0.243\n36m2s\n137\n676MB\n7K iterations\n7K iterations\n30K iterations\n30K iterations\nFig. 6. For some scenes (above) we can see that even at 7K iterations (∼5min\nfor this scene), our method has captured the train quite well. At 30K itera-\ntions (∼35min) the background artifacts have been reduced significantly. For\nother scenes (below), the difference is barely visible; 7K iterations (∼8min)\nis already very high quality.\nTable 2. PSNR scores for Synthetic NeRF, we start with 100K randomly\ninitialized points. Competing metrics extracted from respective papers.\nMic\nChair\nShip\nMaterials\nLego\nDrums\nFicus\nHotdog\nAvg.\nPlenoxels\n33.26\n33.98\n29.62\n29.14\n34.10\n25.35\n31.83\n36.81\n31.76\nINGP-Base\n36.22\n35.00\n31.10\n29.78\n36.39\n26.02\n33.51\n37.40\n33.18\nMip-NeRF\n36.51\n35.14\n30.41\n30.71\n35.70\n25.48\n33.29\n37.48\n33.09\nPoint-NeRF\n35.95\n35.40\n30.97\n29.61\n35.04\n26.06\n36.13\n37.30\n33.30\nOurs-30K\n35.36\n35.83\n30.80\n30.00\n35.78\n26.15\n34.87\n37.72\n33.32\n7.1\nImplementation\nWe implemented our method in Python using the PyTorch frame-\nwork and wrote custom CUDA kernels for rasterization that are\nextended versions of previous methods [Kopanas et al. 2021], and\nuse the NVIDIA CUB sorting routines for the fast Radix sort [Mer-\nrill and Grimshaw 2010]. We also built an interactive viewer using\nthe open-source SIBR [Bonopera et al. 2020], used for interactive\nviewing. We used this implementation to measure our achieved\nframe rates. The source code and all our data are available at:\nhttps://repo-sam.inria.fr/fungraph/3d-gaussian-splatting/\nOptimization Details. For stability, we “warm-up” the computa-\ntion in lower resolution. Specifically, we start the optimization using\n4 times smaller image resolution and we upsample twice after 250\nand 500 iterations.\nSH coefficient optimization is sensitive to the lack of angular\ninformation. For typical “NeRF-like” captures where a central object\nis observed by photos taken in the entire hemisphere around it, the\noptimization works well. However, if the capture has angular regions\nmissing (e.g., when capturing the corner of a scene, or performing\nan “inside-out” [Hedman et al. 2016] capture) completely incorrect\nvalues for the zero-order component of the SH (i.e., the base or\ndiffuse color) can be produced by the optimization. To overcome\nthis problem we start by optimizing only the zero-order component,\nand then introduce one band of the SH after every 1000 iterations\nuntil all 4 bands of SH are represented.\n7.2\nResults and Evaluation\nResults. We tested our algorithm on a total of 13 real scenes\ntaken from previously published datasets and the synthetic Blender\ndataset [Mildenhall et al. 2020]. In particular, we tested our ap-\nproach on the full set of scenes presented in Mip-Nerf360 [Barron\net al. 2022], which is the current state of the art in NeRF rendering\nquality, two scenes from the Tanks&Temples dataset [2017] and\ntwo scenes provided by Hedman et al. [Hedman et al. 2018]. The\nscenes we chose have very different capture styles, and cover both\nbounded indoor scenes and large unbounded outdoor environments.\nWe use the same hyperparameter configuration for all experiments\nin our evaluation. All results are reported running on an A6000 GPU,\nexcept for the Mip-NeRF360 method (see below).\nIn supplemental, we show a rendered video path for a selection\nof scenes that contain views far from the input photos.\nReal-World Scenes. In terms of quality, the current state-of-the-\nart is Mip-Nerf360 [Barron et al. 2021]. We compare against this\nmethod as a quality benchmark. We also compare against two of\nthe most recent fast NeRF methods: InstantNGP [Müller et al. 2022]\nand Plenoxels [Fridovich-Keil and Yu et al. 2022].\nWe use a train/test split for datasets, using the methodology\nsuggested by Mip-NeRF360, taking every 8th photo for test, for con-\nsistent and meaningful comparisons to generate the error metrics,\nusing the standard PSNR, L-PIPS, and SSIM metrics used most fre-\nquently in the literature; please see Table 1. All numbers in the table\nare from our own runs of the author’s code for all previous meth-\nods, except for those of Mip-NeRF360 on their dataset, in which we\ncopied the numbers from the original publication to avoid confusion\nabout the current SOTA. For the images in our figures, we used our\nown run of Mip-NeRF360: the numbers for these runs are in Appen-\ndix D. We also show the average training time, rendering speed, and\nmemory used to store optimized parameters. We report results for a\nbasic configuration of InstantNGP (Base) that run for 35K iterations\nas well as a slightly larger network suggested by the authors (Big),\nand two configurations, 7K and 30K iterations for ours. We show\nACM Trans. Graph., Vol. 42, No. 4, Article . Publication date: August 2023.\n\n\n3D Gaussian Splatting for Real-Time Radiance Field Rendering\n•\n9\nTable 3. PSNR Score for ablation runs. For this experiment, we manually downsampled high-resolution versions of each scene’s input images to the established\nrendering resolution of our other experiments. Doing so reduces random artifacts (e.g., due to JPEG compression in the pre-downscaled Mip-NeRF360 inputs).\nTruck-5K\nGarden-5K\nBicycle-5K\nTruck-30K\nGarden-30K\nBicycle-30K\nAverage-5K\nAverage-30K\nLimited-BW\n14.66\n22.07\n20.77\n13.84\n22.88\n20.87\n19.16\n19.19\nRandom Init\n16.75\n20.90\n19.86\n18.02\n22.19\n21.05\n19.17\n20.42\nNo-Split\n18.31\n23.98\n22.21\n20.59\n26.11\n25.02\n21.50\n23.90\nNo-SH\n22.36\n25.22\n22.88\n24.39\n26.59\n25.08\n23.48\n25.35\nNo-Clone\n22.29\n25.61\n22.15\n24.82\n27.47\n25.46\n23.35\n25.91\nIsotropic\n22.40\n25.49\n22.81\n23.89\n27.00\n24.81\n23.56\n25.23\nFull\n22.71\n25.82\n23.18\n24.81\n27.70\n25.65\n23.90\n26.05\nthe difference in visual quality for our two configurations in Fig. 6.\nIn many cases, quality at 7K iterations is already quite good.\nThe training times vary over datasets and we report them sepa-\nrately. Note that image resolutions also vary over datasets. In the\nproject website, we provide all the renders of test views we used to\ncompute the statistics for all the methods (ours and previous work)\non all scenes. Note that we kept the native input resolution for all\nrenders.\nThe table shows that our fully converged model achieves qual-\nity that is on par and sometimes slightly better than the SOTA\nMip-NeRF360 method; note that on the same hardware, their aver-\nage training time was 48 hours2, compared to our 35-45min, and\ntheir rendering time is 10s/frame. We achieve comparable quality\nto InstantNGP and Plenoxels after 5-10m of training, but additional\ntraining time allows us to achieve SOTA quality which is not the\ncase for the other fast methods. For Tanks & Temples, we achieve\nsimilar quality as the basic InstantNGP at a similar training time\n(∼7min in our case).\nWe also show visual results of this comparison for a left-out\ntest view for ours and the previous rendering methods selected\nfor comparison in Fig. 5; the results of our method are for 30K\niterations of training. We see that in some cases even Mip-NeRF360\nhas remaining artifacts that our method avoids (e.g., blurriness in\nvegetation – in Bicycle, Stump – or on the walls in Room). In the\nsupplemental video and web page we provide comparisons of paths\nfrom a distance. Our method tends to preserve visual detail of well-\ncovered regions even from far away, which is not always the case\nfor previous methods.\nSynthetic Bounded Scenes. In addition to realistic scenes, we also\nevaluate our approach on the synthetic Blender dataset [Mildenhall\net al. 2020]. The scenes in question provide an exhaustive set of\nviews, are limited in size, and provide exact camera parameters. In\nsuch scenarios, we can achieve state-of-the-art results even with\nrandom initialization: we start training from 100K uniformly random\nGaussians inside a volume that encloses the scene bounds. Our\napproach quickly and automatically prunes them to about 6–10K\nmeaningful Gaussians. The final size of the trained model after 30K\niterations reaches about 200–500K Gaussians per scene. We report\nand compare our achieved PSNR scores with previous methods in\nTable 2 using a white background for compatibility. Examples can\n2We trained Mip-NeRF360 on a 4-GPU A100 node for 12 hours, equivalent to 48 hours\non a single GPU. Note that A100’s are faster than A6000 GPUs.\nbe seen in Fig. 10 (second image from the left) and in supplemental\nmaterial. The trained synthetic scenes rendered at 180–300 FPS.\nCompactness. In comparison to previous explicit scene representa-\ntions, the anisotropic Gaussians used in our optimization are capable\nof modelling complex shapes with a lower number of parameters.\nWe showcase this by evaluating our approach against the highly\ncompact, point-based models obtained by [Zhang et al. 2022]. We\nstart from their initial point cloud which is obtained by space carving\nwith foreground masks and optimize until we break even with their\nreported PSNR scores. This usually happens within 2–4 minutes.\nWe surpass their reported metrics using approximately one-fourth\nof their point count, resulting in an average model size of 3.8 MB,\nas opposed to their 9 MB. We note that for this experiment, we only\nused two degrees of our spherical harmonics, similar to theirs.\n7.3\nAblations\nWe isolated the different contributions and algorithmic choices\nwe made and constructed a set of experiments to measure their\neffect. Specifically we test the following aspects of our algorithm:\ninitialization from SfM, our densification strategies, anisotropic\ncovariance, the fact that we allow an unlimited number of splats\nto have gradients and use of spherical harmonics. The quantitative\neffect of each choice is summarized in Table 3.\nInitialization from SfM. We also assess the importance of initializ-\ning the 3D Gaussians from the SfM point cloud. For this ablation, we\nuniformly sample a cube with a size equal to three times the extent\nof the input camera’s bounding box. We observe that our method\nperforms relatively well, avoiding complete failure even without the\nSfM points. Instead, it degrades mainly in the background, see Fig. 7.\nAlso in areas not well covered from training views, the random\ninitialization method appears to have more floaters that cannot be\nremoved by optimization. On the other hand, the synthetic NeRF\ndataset does not have this behavior because it has no background\nand is well constrained by the input cameras (see discussion above).\nDensification. We next evaluate our two densification methods,\nmore specifically the clone and split strategy described in Sec. 5.\nWe disable each method separately and optimize using the rest of\nthe method unchanged. Results show that splitting big Gaussians\nis important to allow good reconstruction of the background as\nseen in Fig. 8, while cloning the small Gaussians instead of splitting\nthem allows for a better and faster convergence especially when\nthin structures appear in the scene.\nACM Trans. Graph., Vol. 42, No. 4, Article . Publication date: August 2023.\n\n\n10\n•\nBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis\nSfM\nRandom\nFig. 7.\nInitialization with SfM points helps. Above: initialization with a\nrandom point cloud. Below: initialization using SfM points.\nFull-5k \nNo Clone-5k\nNo Split-5k\nFig. 8.\nAblation of densification strategy for the two cases \"clone\" and\n\"split\" (Sec. 5).\nUnlimited depth complexity of splats with gradients. We evaluate\nif skipping the gradient computation after the 𝑁front-most points\nFig. 9. If we limit the number of points that receive gradients, the effect on\nvisual quality is significant. Left: limit of 10 Gaussians that receive gradients.\nRight: our full method.\nwill give us speed without sacrificing quality, as suggested in Pul-\nsar [Lassner and Zollhofer 2021]. In this test, we choose N=10, which\nis two times higher than the default value in Pulsar, but it led to\nunstable optimization because of the severe approximation in the\ngradient computation. For the Truck scene, quality degraded by\n11dB in PSNR (see Table 3, Limited-BW), and the visual outcome is\nshown in Fig. 9 for Garden.\nAnisotropic Covariance. An important algorithmic choice in our\nmethod is the optimization of the full covariance matrix for the 3D\nGaussians. To demonstrate the effect of this choice, we perform an\nablation where we remove anisotropy by optimizing a single scalar\nvalue that controls the radius of the 3D Gaussian on all three axes.\nThe results of this optimization are presented visually in Fig. 10.\nWe observe that the anisotropy significantly improves the quality\nof the 3D Gaussian’s ability to align with surfaces, which in turn\nallows for much higher rendering quality while maintaining the\nsame number of points.\nSpherical Harmonics. Finally, the use of spherical harmonics im-\nproves our overall PSNR scores since they compensate for the view-\ndependent effects (Table 3).\n7.4\nLimitations\nOur method is not without limitations. In regions where the scene\nis not well observed we have artifacts; in such regions, other meth-\nods also struggle (e.g., Mip-NeRF360 in Fig. 11). Even though the\nanisotropic Gaussians have many advantages as described above,\nour method can create elongated artifacts or “splotchy” Gaussians\n(see Fig. 12); again previous methods also struggle in these cases.\nWe also occasionally have popping artifacts when our optimiza-\ntion creates large Gaussians; this tends to happen in regions with\nview-dependent appearance. One reason for these popping artifacts\nis the trivial rejection of Gaussians via a guard band in the rasterizer.\nA more principled culling approach would alleviate these artifacts.\nAnother factor is our simple visibility algorithm, which can lead to\nGaussians suddenly switching depth/blending order. This could be\naddressed by antialiasing, which we leave as future work. Also, we\ncurrently do not apply any regularization to our optimization; doing\nso would help with both the unseen region and popping artifacts.\nWhile we used the same hyperparameters for our full evaluation,\nearly experiments show that reducing the position learning rate can\nbe necessary to converge in very large scenes (e.g., urban datasets).\nACM Trans. Graph., Vol. 42, No. 4, Article . Publication date: August 2023.\n\n\n3D Gaussian Splatting for Real-Time Radiance Field Rendering\n•\n11\nGround \nTruth\nFull\nIsotropic\nGround \nTruth\nFull\nIsotropic\nGround \nTruth\nFull\nIsotropic\nFig. 10. We train scenes with Gaussian anisotropy disabled and enabled. The use of anisotropic volumetric splats enables modelling of fine structures and has\na significant impact on visual quality. Note that for illustrative purposes, we restricted Ficus to use no more than 5k Gaussians in both configurations.\nEven though we are very compact compared to previous point-\nbased approaches, our memory consumption is significantly higher\nthan NeRF-based solutions. During training of large scenes, peak\nGPU memory consumption can exceed 20 GB in our unoptimized\nprototype. However, this figure could be significantly reduced by a\ncareful low-level implementation of the optimization logic (similar\nto InstantNGP). Rendering the trained scene requires sufficient GPU\nmemory to store the full model (several hundred megabytes for\nlarge-scale scenes) and an additional 30–500 MB for the rasterizer,\ndepending on scene size and image resolution. We note that there\nare many opportunities to further reduce memory consumption\nof our method. Compression techniques for point clouds is a well-\nstudied field [De Queiroz and Chou 2016]; it would be interesting to\nsee how such approaches could be adapted to our representation.\nFig. 11. Comparison of failure artifacts: Mip-NeRF360 has “floaters” and\ngrainy appearance (left, foreground), while our method produces coarse,\nanisoptropic Gaussians resulting in low-detail visuals (right, background).\nTrain scene.\nFig. 12. In views that have little overlap with those seen during training,\nour method may produce artifacts (right). Again, Mip-NeRF360 also has\nartifacts in these cases (left). DrJohnson scene.\n8\nDISCUSSION AND CONCLUSIONS\nWe have presented the first approach that truly allows real-time,\nhigh-quality radiance field rendering, in a wide variety of scenes\nand capture styles, while requiring training times competitive with\nthe fastest previous methods.\nOur choice of a 3D Gaussian primitive preserves properties of\nvolumetric rendering for optimization while directly allowing fast\nsplat-based rasterization. Our work demonstrates that – contrary to\nwidely accepted opinion – a continuous representation is not strictly\nnecessary to allow fast and high-quality radiance field training.\nThe majority (∼80%) of our training time is spent in Python code,\nsince we built our solution in PyTorch to allow our method to be\neasily used by others. Only the rasterization routine is implemented\nas optimized CUDA kernels. We expect that porting the remaining\noptimization entirely to CUDA, as e.g., done in InstantNGP [Müller\net al. 2022], could enable significant further speedup for applications\nwhere performance is essential.\nWe also demonstrated the importance of building on real-time\nrendering principles, exploiting the power of the GPU and speed of\nsoftware rasterization pipeline architecture. These design choices\nare the key to performance both for training and real-time render-\ning, providing a competitive edge in performance over previous\nvolumetric ray-marching.\nIt would be interesting to see if our Gaussians can be used to per-\nform mesh reconstructions of the captured scene. Aside from prac-\ntical implications given the widespread use of meshes, this would\nallow us to better understand where our method stands exactly in\nthe continuum between volumetric and surface representations.\nIn conclusion, we have presented the first real-time rendering\nsolution for radiance fields, with rendering quality that matches the\nbest expensive previous methods, with training times competitive\nwith the fastest existing solutions.\nACKNOWLEDGMENTS\nThis research was funded by the ERC Advanced grant FUNGRAPH\nNo 788065 http://fungraph.inria.fr. The authors are grateful to Adobe\nfor generous donations, the OPAL infrastructure from Université\nCôte d’Azur and for the HPC resources from GENCI–IDRIS (Grant\n2022-AD011013409). The authors thank the anonymous reviewers\nfor their valuable feedback, P. Hedman and A. Tewari for proof-\nreading earlier drafts also T. Müller, A. Yu and S. Fridovich-Keil for\nhelping with the comparisons.\nACM Trans. Graph., Vol. 42, No. 4, Article . Publication date: August 2023.\n\n\n12\n•\nBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis\nREFERENCES\nKara-Ali Aliev, Artem Sevastopolsky, Maria Kolos, Dmitry Ulyanov, and Victor Lem-\npitsky. 2020. Neural Point-Based Graphics. In Computer Vision – ECCV 2020: 16th\nEuropean Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXII. 696–\n712.\nJonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-\nBrualla, and Pratul P Srinivasan. 2021. Mip-nerf: A multiscale representation for\nanti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF International\nConference on Computer Vision. 5855–5864.\nJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman.\n2022. Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields. CVPR (2022).\nSebastien Bonopera, Jerome Esnault, Siddhant Prakash, Simon Rodriguez, Theo Thonat,\nMehdi Benadel, Gaurav Chaurasia, Julien Philip, and George Drettakis. 2020. sibr:\nA System for Image Based Rendering. https://gitlab.inria.fr/sibr/sibr_core\nMario Botsch, Alexander Hornung, Matthias Zwicker, and Leif Kobbelt. 2005. High-\nQuality Surface Splatting on Today’s GPUs. In Proceedings of the Second Eurographics\n/ IEEE VGTC Conference on Point-Based Graphics (New York, USA) (SPBG’05). Euro-\ngraphics Association, Goslar, DEU, 17–24.\nChris Buehler, Michael Bosse, Leonard McMillan, Steven Gortler, and Michael Cohen.\n2001. Unstructured lumigraph rendering. In Proc. SIGGRAPH.\nGaurav Chaurasia, Sylvain Duchene, Olga Sorkine-Hornung, and George Drettakis.\n2013. Depth synthesis and local warps for plausible image-based navigation. ACM\nTransactions on Graphics (TOG) 32, 3 (2013), 1–12.\nAnpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. 2022b. TensoRF:\nTensorial Radiance Fields. In European Conference on Computer Vision (ECCV).\nZhiqin Chen, Thomas Funkhouser, Peter Hedman, and Andrea Tagliasacchi. 2022a.\nMobileNeRF: Exploiting the Polygon Rasterization Pipeline for Efficient Neural Field\nRendering on Mobile Architectures. arXiv preprint arXiv:2208.00277 (2022).\nRicardo L De Queiroz and Philip A Chou. 2016. Compression of 3D point clouds using\na region-adaptive hierarchical transform. IEEE Transactions on Image Processing 25,\n8 (2016), 3947–3956.\nMartin Eisemann, Bert De Decker, Marcus Magnor, Philippe Bekaert, Edilson De Aguiar,\nNaveed Ahmed, Christian Theobalt, and Anita Sellent. 2008. Floating textures. In\nComputer graphics forum, Vol. 27. Wiley Online Library, 409–418.\nJohn Flynn, Ivan Neulander, James Philbin, and Noah Snavely. 2016. Deepstereo:\nLearning to predict new views from the world’s imagery. In CVPR.\nFridovich-Keil and Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo\nKanazawa. 2022. Plenoxels: Radiance Fields without Neural Networks. In CVPR.\nStephan J. Garbin, Marek Kowalski, Matthew Johnson, Jamie Shotton, and Julien\nValentin. 2021. FastNeRF: High-Fidelity Neural Rendering at 200FPS. In Proceedings\nof the IEEE/CVF International Conference on Computer Vision (ICCV). 14346–14355.\nMichael Goesele, Noah Snavely, Brian Curless, Hugues Hoppe, and Steven M Seitz.\n2007. Multi-view stereo for community photo collections. In ICCV.\nSteven J Gortler, Radek Grzeszczuk, Richard Szeliski, and Michael F Cohen. 1996. The\nlumigraph. In Proceedings of the 23rd annual conference on Computer graphics and\ninteractive techniques. 43–54.\nMarkus Gross and Hanspeter (Eds) Pfister. 2011. Point-based graphics. Elsevier.\nJeff P. Grossman and William J. Dally. 1998. Point Sample Rendering. In Rendering\nTechniques.\nPeter Hedman, Julien Philip, True Price, Jan-Michael Frahm, George Drettakis, and\nGabriel Brostow. 2018. Deep blending for free-viewpoint image-based rendering.\nACM Trans. on Graphics (TOG) 37, 6 (2018).\nPeter Hedman, Tobias Ritschel, George Drettakis, and Gabriel Brostow. 2016. Scalable\nInside-Out Image-Based Rendering. ACM Transactions on Graphics (SIGGRAPH\nAsia Conference Proceedings) 35, 6 (December 2016). http://www-sop.inria.fr/reves/\nBasilic/2016/HRDB16\nPeter Hedman, Pratul P. Srinivasan, Ben Mildenhall, Jonathan T. Barron, and Paul\nDebevec. 2021. Baking Neural Radiance Fields for Real-Time View Synthesis. ICCV\n(2021).\nPhilipp Henzler, Niloy J Mitra, and Tobias Ritschel. 2019. Escaping plato’s cave: 3d shape\nfrom adversarial rendering. In Proceedings of the IEEE/CVF International Conference\non Computer Vision. 9984–9993.\nArno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. 2017. Tanks and\ntemples: Benchmarking large-scale scene reconstruction. ACM Transactions on\nGraphics (ToG) 36, 4 (2017), 1–13.\nGeorgios Kopanas, Thomas Leimkühler, Gilles Rainer, Clément Jambon, and George\nDrettakis. 2022. Neural Point Catacaustics for Novel-View Synthesis of Reflections.\nACM Transactions on Graphics (SIGGRAPH Asia Conference Proceedings) 41, 6 (2022),\n201. http://www-sop.inria.fr/reves/Basilic/2022/KLRJD22\nGeorgios Kopanas, Julien Philip, Thomas Leimkühler, and George Drettakis. 2021. Point-\nBased Neural Rendering with Per-View Optimization. Computer Graphics Forum 40,\n4 (2021), 29–43. https://doi.org/10.1111/cgf.14339\nSamuli Laine and Tero Karras. 2011. High-performance software rasterization on GPUs.\nIn Proceedings of the ACM SIGGRAPH Symposium on High Performance Graphics.\n79–88.\nChristoph Lassner and Michael Zollhofer. 2021. Pulsar: Efficient Sphere-Based Neural\nRendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern\nRecognition (CVPR). 1440–1449.\nMarc Levoy and Pat Hanrahan. 1996. Light field rendering. In Proceedings of the 23rd\nannual conference on Computer graphics and interactive techniques. 31–42.\nStephen Lombardi, Tomas Simon, Gabriel Schwartz, Michael Zollhoefer, Yaser Sheikh,\nand Jason Saragih. 2021. Mixture of volumetric primitives for efficient neural\nrendering. ACM Transactions on Graphics (TOG) 40, 4 (2021), 1–13.\nDuane G Merrill and Andrew S Grimshaw. 2010. Revisiting sorting for GPGPU stream\narchitectures. In Proceedings of the 19th international conference on Parallel architec-\ntures and compilation techniques. 545–546.\nBen Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ra-\nmamoorthi, and Ren Ng. 2020. NeRF: Representing Scenes as Neural Radiance Fields\nfor View Synthesis. In ECCV.\nThomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. 2022. Instant\nNeural Graphics Primitives with a Multiresolution Hash Encoding. ACM Trans.\nGraph. 41, 4, Article 102 (July 2022), 15 pages.\nhttps://doi.org/10.1145/3528223.\n3530127\nEric Penner and Li Zhang. 2017. Soft 3D reconstruction for view synthesis. ACM\nTransactions on Graphics (TOG) 36, 6 (2017), 1–11.\nHanspeter Pfister, Matthias Zwicker, Jeroen van Baar, and Markus Gross. 2000. Surfels:\nSurface Elements as Rendering Primitives. In Proceedings of the 27th Annual Con-\nference on Computer Graphics and Interactive Techniques (SIGGRAPH ’00). ACM\nPress/Addison-Wesley Publishing Co., USA, 335–342.\nhttps://doi.org/10.1145/\n344779.344936\nChristian Reiser, Songyou Peng, Yiyi Liao, and Andreas Geiger. 2021. KiloNeRF: Speed-\ning up Neural Radiance Fields with Thousands of Tiny MLPs. In International\nConference on Computer Vision (ICCV).\nLiu Ren, Hanspeter Pfister, and Matthias Zwicker. 2002. Object Space EWA Surface\nSplatting: A Hardware Accelerated Approach to High Quality Point Rendering.\nComputer Graphics Forum 21 (2002).\nHelge Rhodin, Nadia Robertini, Christian Richardt, Hans-Peter Seidel, and Christian\nTheobalt. 2015. A versatile scene model with differentiable visibility applied to\ngenerative pose estimation. In Proceedings of the IEEE International Conference on\nComputer Vision. 765–773.\nGernot Riegler and Vladlen Koltun. 2020. Free view synthesis. In European Conference\non Computer Vision. Springer, 623–640.\nDarius Rückert, Linus Franke, and Marc Stamminger. 2022. ADOP: Approximate\nDifferentiable One-Pixel Point Rendering. ACM Trans. Graph. 41, 4, Article 99 (jul\n2022), 14 pages. https://doi.org/10.1145/3528223.3530122\nMiguel Sainz and Renato Pajarola. 2004. Point-based rendering techniques. Computers\nand Graphics 28, 6 (2004), 869–879. https://doi.org/10.1016/j.cag.2004.08.014\nJohannes Lutz Schönberger and Jan-Michael Frahm. 2016. Structure-from-Motion\nRevisited. In Conference on Computer Vision and Pattern Recognition (CVPR).\nMarkus Schütz, Bernhard Kerbl, and Michael Wimmer. 2022. Software Rasterization of\n2 Billion Points in Real Time. Proc. ACM Comput. Graph. Interact. Tech. 5, 3, Article\n24 (jul 2022), 17 pages. https://doi.org/10.1145/3543863\nVincent Sitzmann, Justus Thies, Felix Heide, Matthias Nießner, Gordon Wetzstein, and\nMichael Zollhofer. 2019. Deepvoxels: Learning persistent 3d feature embeddings. In\nProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.\n2437–2446.\nNoah Snavely, Steven M Seitz, and Richard Szeliski. 2006. Photo tourism: exploring\nphoto collections in 3D. In Proc. SIGGRAPH.\nCarsten Stoll, Nils Hasler, Juergen Gall, Hans-Peter Seidel, and Christian Theobalt. 2011.\nFast articulated motion tracking using a sums of gaussians body model. In 2011\nInternational Conference on Computer Vision. IEEE, 951–958.\nCheng Sun, Min Sun, and Hwann-Tzong Chen. 2022. Direct Voxel Grid Optimization:\nSuper-fast Convergence for Radiance Fields Reconstruction. In CVPR.\nTowaki Takikawa, Alex Evans, Jonathan Tremblay, Thomas Müller, Morgan McGuire,\nAlec Jacobson, and Sanja Fidler. 2022. Variable bitrate neural fields. In ACM SIG-\nGRAPH 2022 Conference Proceedings. 1–9.\nTowaki Takikawa, Joey Litalien, Kangxue Yin, Karsten Kreis, Charles Loop, Derek\nNowrouzezahrai, Alec Jacobson, Morgan McGuire, and Sanja Fidler. 2021. Neural\nGeometric Level of Detail: Real-time Rendering with Implicit 3D Shapes. (2021).\nAyush Tewari, Justus Thies, Ben Mildenhall, Pratul Srinivasan, Edgar Tretschk, W Yifan,\nChristoph Lassner, Vincent Sitzmann, Ricardo Martin-Brualla, Stephen Lombardi,\net al. 2022. Advances in neural rendering. In Computer Graphics Forum, Vol. 41.\nWiley Online Library, 703–735.\nJustus Thies, Michael Zollhöfer, and Matthias Nießner. 2019. Deferred neural rendering:\nImage synthesis using neural textures. ACM Transactions on Graphics (TOG) 38, 4\n(2019), 1–12.\nAngtian Wang, Peng Wang, Jian Sun, Adam Kortylewski, and Alan Yuille. 2023. VoGE: A\nDifferentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis.\nIn The Eleventh International Conference on Learning Representations.\nhttps://\nopenreview.net/forum?id=AdPJb9cud_Y\nACM Trans. Graph., Vol. 42, No. 4, Article . Publication date: August 2023.\n\n\n3D Gaussian Splatting for Real-Time Radiance Field Rendering\n•\n13\nOlivia Wiles, Georgia Gkioxari, Richard Szeliski, and Justin Johnson. 2020. Synsin:\nEnd-to-end view synthesis from a single image. In Proceedings of the IEEE/CVF\nConference on Computer Vision and Pattern Recognition. 7467–7477.\nXiuchao Wu, Jiamin Xu, Zihan Zhu, Hujun Bao, Qixing Huang, James Tompkin, and\nWeiwei Xu. 2022. Scalable Neural Indoor Scene Rendering. ACM Transactions on\nGraphics (TOG) (2022).\nYiheng Xie, Towaki Takikawa, Shunsuke Saito, Or Litany, Shiqin Yan, Numair Khan,\nFederico Tombari, James Tompkin, Vincent Sitzmann, and Srinath Sridhar. 2022.\nNeural fields in visual computing and beyond. In Computer Graphics Forum, Vol. 41.\nWiley Online Library, 641–676.\nQiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi, Zhixin Shu, Kalyan Sunkavalli, and\nUlrich Neumann. 2022. Point-nerf: Point-based neural radiance fields. In Proceedings\nof the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5438–5448.\nWang Yifan, Felice Serena, Shihao Wu, Cengiz Öztireli, and Olga Sorkine-Hornung.\n2019. Differentiable surface splatting for point-based geometry processing. ACM\nTransactions on Graphics (TOG) 38, 6 (2019), 1–14.\nAlex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and Angjoo Kanazawa. 2021.\nPlenOctrees for Real-time Rendering of Neural Radiance Fields. In ICCV.\nQiang Zhang, Seung-Hwan Baek, Szymon Rusinkiewicz, and Felix Heide. 2022. Dif-\nferentiable Point-Based Radiance Fields for Efficient View Synthesis. In SIGGRAPH\nAsia 2022 Conference Papers (Daegu, Republic of Korea) (SA ’22). Association for\nComputing Machinery, New York, NY, USA, Article 7, 12 pages. https://doi.org/10.\n1145/3550469.3555413\nTinghui Zhou, Shubham Tulsiani, Weilun Sun, Jitendra Malik, and Alexei A Efros.\n2016. View synthesis by appearance flow. In European conference on computer vision.\nSpringer, 286–301.\nMatthias Zwicker, Hanspeter Pfister, Jeroen Van Baar, and Markus Gross. 2001a. EWA\nvolume splatting. In Proceedings Visualization, 2001. VIS’01. IEEE, 29–538.\nMatthias Zwicker, Hanspeter Pfister, Jeroen van Baar, and Markus Gross. 2001b. Surface\nSplatting. In Proceedings of the 28th Annual Conference on Computer Graphics and\nInteractive Techniques (SIGGRAPH ’01). Association for Computing Machinery, New\nYork, NY, USA, 371–378. https://doi.org/10.1145/383259.383300\nA\nDETAILS OF GRADIENT COMPUTATION\nRecall that Σ/Σ′ are the world/view space covariance matrices of\nthe Gaussian, 𝑞is the rotation, and 𝑠the scaling, 𝑊is the viewing\ntransformation and 𝐽the Jacobian of the affine approximation of\nthe projective transformation. We can apply the chain rule to find\nthe derivatives w.r.t. scaling and rotation:\n𝑑Σ′\n𝑑𝑠= 𝑑Σ′\n𝑑Σ\n𝑑Σ\n𝑑𝑠\n(8)\nand\n𝑑Σ′\n𝑑𝑞= 𝑑Σ′\n𝑑Σ\n𝑑Σ\n𝑑𝑞\n(9)\nSimplifying Eq. 5 using𝑈= 𝐽𝑊and Σ′ being the (symmetric) upper\nleft 2×2 matrix of𝑈Σ𝑈𝑇, denoting matrix elements with subscripts,\nwe can find the partial derivatives 𝜕Σ′\n𝜕Σ𝑖𝑗=\n\u0010 𝑈1,𝑖𝑈1,𝑗𝑈1,𝑖𝑈2,𝑗\n𝑈1,𝑗𝑈2,𝑖𝑈2,𝑖𝑈2,𝑗\n\u0011\n.\nNext, we seek the derivatives 𝑑Σ\n𝑑𝑠and 𝑑Σ\n𝑑𝑞. Since Σ = 𝑅𝑆𝑆𝑇𝑅𝑇,\nwe can compute 𝑀= 𝑅𝑆and rewrite Σ = 𝑀𝑀𝑇. Thus, we can\nwrite 𝑑Σ\n𝑑𝑠=\n𝑑Σ\n𝑑𝑀\n𝑑𝑀\n𝑑𝑠and 𝑑Σ\n𝑑𝑞=\n𝑑Σ\n𝑑𝑀\n𝑑𝑀\n𝑑𝑞. Since the covariance ma-\ntrix Σ (and its gradient) is symmetric, the shared first part is com-\npactly found by 𝑑Σ\n𝑑𝑀= 2𝑀𝑇. For scaling, we further have 𝜕𝑀𝑖,𝑗\n𝜕𝑠𝑘\n=\n\u001a 𝑅𝑖,𝑘\nif j = k\n0\notherwise\n\u001b\n. To derive gradients for rotation, we recall the\nconversion from a unit quaternion 𝑞with real part 𝑞𝑟and imaginary\nparts 𝑞𝑖,𝑞𝑗,𝑞𝑘to a rotation matrix 𝑅:\n𝑅(𝑞) = 2\n©\n­\n­\n«\n1\n2 −(𝑞2\n𝑗+ 𝑞2\n𝑘)\n(𝑞𝑖𝑞𝑗−𝑞𝑟𝑞𝑘)\n(𝑞𝑖𝑞𝑘+ 𝑞𝑟𝑞𝑗)\n(𝑞𝑖𝑞𝑗+ 𝑞𝑟𝑞𝑘)\n1\n2 −(𝑞2\n𝑖+ 𝑞2\n𝑘)\n(𝑞𝑗𝑞𝑘−𝑞𝑟𝑞𝑖)\n(𝑞𝑖𝑞𝑘−𝑞𝑟𝑞𝑗)\n(𝑞𝑗𝑞𝑘+ 𝑞𝑟𝑞𝑖)\n1\n2 −(𝑞2\n𝑖+ 𝑞2\n𝑗)\nª\n®\n®\n¬\n(10)\nAs a result, we find the following gradients for the components of 𝑞:\n𝜕𝑀\n𝜕𝑞𝑟\n= 2\n\u0012\n0\n−𝑠𝑦𝑞𝑘𝑠𝑧𝑞𝑗\n𝑠𝑥𝑞𝑘\n0\n−𝑠𝑧𝑞𝑖\n−𝑠𝑥𝑞𝑗\n𝑠𝑦𝑞𝑖\n0\n\u0013\n,\n𝜕𝑀\n𝜕𝑞𝑖\n= 2\n\u0012\n0\n𝑠𝑦𝑞𝑗\n𝑠𝑧𝑞𝑘\n𝑠𝑥𝑞𝑗−2𝑠𝑦𝑞𝑖−𝑠𝑧𝑞𝑟\n𝑠𝑥𝑞𝑘\n𝑠𝑦𝑞𝑟\n−2𝑠𝑧𝑞𝑖\n\u0013\n𝜕𝑀\n𝜕𝑞𝑗\n= 2\n\u0012 −2𝑠𝑥𝑞𝑗𝑠𝑦𝑞𝑖\n𝑠𝑧𝑞𝑟\n𝑠𝑥𝑞𝑖\n0\n𝑠𝑧𝑞𝑘\n−𝑠𝑥𝑞𝑟𝑠𝑦𝑞𝑘−2𝑠𝑧𝑞𝑗\n\u0013\n,\n𝜕𝑀\n𝜕𝑞𝑘\n= 2\n\u0012 −2𝑠𝑥𝑞𝑘−𝑠𝑦𝑞𝑟𝑠𝑧𝑞𝑖\n𝑠𝑥𝑞𝑟\n−2𝑠𝑦𝑞𝑘𝑠𝑧𝑞𝑗\n𝑠𝑥𝑞𝑖\n𝑠𝑦𝑞𝑗\n0\n\u0013\n(11)\nDeriving gradients for quaternion normalization is straightforward.\nB\nOPTIMIZATION AND DENSIFICATION ALGORITHM\nOur optimization and densification algorithms are summarized in\nAlgorithm 1.\nAlgorithm 1 Optimization and Densification\n𝑤, ℎ: width and height of the training images\n𝑀←SfM Points\n⊲Positions\n𝑆,𝐶,𝐴←InitAttributes()\n⊲Covariances, Colors, Opacities\n𝑖←0\n⊲Iteration Count\nwhile not converged do\n𝑉, ˆ\n𝐼←SampleTrainingView()\n⊲Camera 𝑉and Image\n𝐼←Rasterize(𝑀, 𝑆, 𝐶, 𝐴, 𝑉)\n⊲Alg. 2\n𝐿←𝐿𝑜𝑠𝑠(𝐼, ˆ\n𝐼)\n⊲Loss\n𝑀, 𝑆, 𝐶, 𝐴←Adam(∇𝐿)\n⊲Backprop & Step\nif IsRefinementIteration(𝑖) then\nfor all Gaussians (𝜇, Σ,𝑐, 𝛼) in (𝑀,𝑆,𝐶,𝐴) do\nif 𝛼< 𝜖or IsTooLarge(𝜇, Σ) then\n⊲Pruning\nRemoveGaussian()\nend if\nif ∇𝑝𝐿> 𝜏𝑝then\n⊲Densification\nif ∥𝑆∥> 𝜏𝑆then\n⊲Over-reconstruction\nSplitGaussian(𝜇, Σ,𝑐, 𝛼)\nelse\n⊲Under-reconstruction\nCloneGaussian(𝜇, Σ,𝑐, 𝛼)\nend if\nend if\nend for\nend if\n𝑖←𝑖+ 1\nend while\nC\nDETAILS OF THE RASTERIZER\nSorting. Our design is based on the assumption of a high load\nof small splats, and we optimize for this by sorting splats once for\neach frame using radix sort at the beginning. We split the screen\ninto 16x16 pixel tiles (or bins). We create a list of splats per tile by\ninstantiating each splat in each 16×16 tile it overlaps. This results\nin a moderate increase in Gaussians to process which however is\namortized by simpler control flow and high parallelism of optimized\nGPU Radix sort [Merrill and Grimshaw 2010]. We assign a key for\neach splats instance with up to 64 bits where the lower 32 bits\nencode its projected depth and the higher bits encode the index of\nthe overlapped tile. The exact size of the index depends on how\nmany tiles fit the current resolution. Depth ordering is thus directly\nresolved for all splats in parallel with a single radix sort. After\nACM Trans. Graph., Vol. 42, No. 4, Article . Publication date: August 2023.\n\n\n14\n•\nBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis\nsorting, we can efficiently produce per-tile lists of Gaussians to\nprocess by identifying the start and end of ranges in the sorted\narray with the same tile ID. This is done in parallel, launching\none thread per 64-bit array element to compare its higher 32 bits\nwith its two neighbors. Compared to [Lassner and Zollhofer 2021],\nour rasterization thus completely eliminates sequential primitive\nprocessing steps and produces more compact per-tile lists to traverse\nduring the forward pass. We show a high-level overview of the\nrasterization approach in Algorithm 2.\nAlgorithm 2 GPU software rasterization of 3D Gaussians\n𝑤, ℎ: width and height of the image to rasterize\n𝑀, 𝑆: Gaussian means and covariances in world space\n𝐶, 𝐴: Gaussian colors and opacities\n𝑉: view configuration of current camera\nfunction Rasterize(𝑤, ℎ, 𝑀, 𝑆, 𝐶, 𝐴, 𝑉)\nCullGaussian(𝑝, 𝑉)\n⊲Frustum Culling\n𝑀′,𝑆′ ←ScreenspaceGaussians(𝑀, 𝑆, 𝑉)\n⊲Transform\n𝑇←CreateTiles(𝑤, ℎ)\n𝐿, 𝐾←DuplicateWithKeys(𝑀′, 𝑇)\n⊲Indices and Keys\nSortByKeys(𝐾, 𝐿)\n⊲Globally Sort\n𝑅←IdentifyTileRanges(𝑇, 𝐾)\n𝐼←0\n⊲Init Canvas\nfor all Tiles 𝑡in 𝐼do\nfor all Pixels 𝑖in 𝑡do\n𝑟←GetTileRange(𝑅, 𝑡)\n𝐼[𝑖] ←BlendInOrder(𝑖, 𝐿, 𝑟, 𝐾, 𝑀′, 𝑆′, 𝐶, 𝐴)\nend for\nend for\nreturn 𝐼\nend function\nNumerical stability. During the backward pass, we reconstruct\nthe intermediate opacity values needed for gradient computation by\nrepeatedly dividing the accumulated opacity from the forward pass\nby each Gaussian’s 𝛼. Implemented naïvely, this process is prone to\nnumerical instabilities (e.g., division by 0). To address this, both in\nthe forward and backward pass, we skip any blending updates with\n𝛼< 𝜖(we choose 𝜖as\n1\n255) and also clamp 𝛼with 0.99 from above.\nFinally, before a Gaussian is included in the forward rasterization\npass, we compute the accumulated opacity if we were to include it\nand stop front-to-back blending before it can exceed 0.9999.\nD\nPER-SCENE ERROR METRICS\nTables 4–9 list the various collected error metrics for our evaluation\nover all considered techniques and real-world scenes. We list both\nthe copied Mip-NeRF360 numbers and those of our runs used to\ngenerate the images in the paper; averages for these over the full\nMip-NeRF360 dataset are PSNR 27.58, SSIM 0.790, and LPIPS 0.240.\nTable 4. SSIM scores for Mip-NeRF360 scenes. † copied from original paper.\nbicycle\nflowers\ngarden\nstump\ntreehill\nroom\ncounter\nkitchen\nbonsai\nPlenoxels\n0.496\n0.431\n0.6063\n0.523\n0.509\n0.8417\n0.759\n0.648\n0.814\nINGP-Base\n0.491\n0.450\n0.649\n0.574\n0.518\n0.855\n0.798\n0.818\n0.890\nINGP-Big\n0.512\n0.486\n0.701\n0.594\n0.542\n0.871\n0.817\n0.858\n0.906\nMip-NeRF360†\n0.685\n0.583\n0.813\n0.744\n0.632\n0.913\n0.894\n0.920\n0.941\nMip-NeRF360\n0.685\n0.584\n0.809\n0.745\n0.631\n0.910\n0.892\n0.917\n0.938\nOurs-7k\n0.675\n0.525\n0.836\n0.728\n0.598\n0.884\n0.873\n0.900\n0.910\nOurs-30k\n0.771\n0.605\n0.868\n0.775\n0.638\n0.914\n0.905\n0.922\n0.938\nTable 5. PSNR scores for Mip-NeRF360 scenes. † copied from original paper.\nbicycle\nflowers\ngarden\nstump\ntreehill\nroom\ncounter\nkitchen\nbonsai\nPlenoxels\n21.912\n20.097\n23.4947\n20.661\n22.248\n27.594\n23.624\n23.420\n24.669\nINGP-Base\n22.193\n20.348\n24.599\n23.626\n22.364\n29.269\n26.439\n28.548\n30.337\nINGP-Big\n22.171\n20.652\n25.069\n23.466\n22.373\n29.690\n26.691\n29.479\n30.685\nMip-NeRF360†\n24.37\n21.73\n26.98\n26.40\n22.87\n31.63\n29.55\n32.23\n33.46\nMip-NeRF360\n24.305\n21.649\n26.875\n26.175\n22.929\n31.467\n29.447\n31.989\n33.397\nOurs-7k\n23.604\n20.515\n26.245\n25.709\n22.085\n28.139\n26.705\n28.546\n28.850\nOurs-30k\n25.246\n21.520\n27.410\n26.550\n22.490\n30.632\n28.700\n30.317\n31.980\nTable 6. LPIPS scores for Mip-NeRF360 scenes. † copied from original paper.\nbicycle\nflowers\ngarden\nstump\ntreehill\nroom\ncounter\nkitchen\nbonsai\nPlenoxels\n0.506\n0.521\n0.3864\n0.503\n0.540\n0.4186\n0.441\n0.447\n0.398\nINGP-Base\n0.487\n0.481\n0.312\n0.450\n0.489\n0.301\n0.342\n0.254\n0.227\nINGP-Big\n0.446\n0.441\n0.257\n0.421\n0.450\n0.261\n0.306\n0.195\n0.205\nMip-NeRF360†\n0.301\n0.344\n0.170\n0.261\n0.339\n0.211\n0.204\n0.127\n0.176\nMip-NeRF360\n0.305\n0.346\n0.171\n0.265\n0.347\n0.213\n0.207\n0.128\n0.179\nOurs-7k\n0.318\n0.417\n0.153\n0.287\n0.404\n0.272\n0.254\n0.161\n0.244\nOurs-30k\n0.205\n0.336\n0.103\n0.210\n0.317\n0.220\n0.204\n0.129\n0.205\nTable 7. SSIM scores for Tanks&Temples and Deep Blending scenes.\nTruck\nTrain\nDr Johnson\nPlayroom\nPlenoxels\n0.774\n0.663\n0.787\n0.802\nINGP-Base\n0.779\n0.666\n0.839\n0.754\nINGP-Big\n0.800\n0.689\n0.854\n0.779\nMip-NeRF360\n0.857\n0.660\n0.901\n0.900\nOurs-7k\n0.840\n0.694\n0.853\n0.896\nOurs-30k\n0.879\n0.802\n0.899\n0.906\nTable 8. PSNR scores for Tanks&Temples and Deep Blending scenes.\nTruck\nTrain\nDr Johnson\nPlayroom\nPlenoxels\n23.221\n18.927\n23.142\n22.980\nINGP-Base\n23.260\n20.170\n27.750\n19.483\nINGP-Big\n23.383\n20.456\n28.257\n21.665\nMip-NeRF360\n24.912\n19.523\n29.140\n29.657\nOurs-7k\n23.506\n18.892\n26.306\n29.245\nOurs-30k\n25.187\n21.097\n28.766\n30.044\nTable 9. LPIPS scores for Tanks&Temples and Deep Blending scenes.\nTruck\nTrain\nDr Johnson\nPlayroom\nPlenoxels\n0.335\n0.422\n0.521\n0.499\nINGP-Base\n0.274\n0.386\n0.381\n0.465\nINGP-Big\n0.249\n0.360\n0.352\n0.428\nMip-NeRF360\n0.159\n0.354\n0.237\n0.252\nOurs-7k\n0.209\n0.350\n0.343\n0.291\nOurs-30k\n0.148\n0.218\n0.244\n0.241\nACM Trans. Graph., Vol. 42, No. 4, Article . Publication date: August 2023.\n\n\n2D Gaussian Splatting for Geometrically Accurate Radiance Fields\nBINBIN HUANG, ShanghaiTech University, China\nZEHAO YU, University of Tübingen, Tübingen AI Center, Germany\nANPEI CHEN, University of Tübingen, Tübingen AI Center, Germany\nANDREAS GEIGER, University of Tübingen, Tübingen AI Center, Germany\nSHENGHUA GAO, ShanghaiTech University, China\nhttps://surfsplatting.github.io\nMesh\nRadiance field\nDisk (color)\nDisk (normal)\nSurface normal\n(a) 2D disks as surface elements\n(b) 2D Gaussian splatting\n(c) Meshing\nFig. 1. Our method, 2DGS, (a) optimizes a set of 2D oriented disks to represent and reconstruct a complex real-world scene from multi-view RGB images. These\noptimized 2D disks are tightly aligned to the surfaces. (b) With 2D Gaussian splatting, we allow real-time rendering of high quality novel view images with\nview consistent normals and depth maps. (c) Finally, our method provides detailed and noise-free triangle mesh reconstruction from the optimized 2D disks.\n3D Gaussian Splatting (3DGS) has recently revolutionized radiance field\nreconstruction, achieving high quality novel view synthesis and fast render-\ning speed. However, 3DGS fails to accurately represent surfaces due to the\nmulti-view inconsistent nature of 3D Gaussians. We present 2D Gaussian\nSplatting (2DGS), a novel approach to model and reconstruct geometrically\naccurate radiance fields from multi-view images. Our key idea is to collapse\nthe 3D volume into a set of 2D oriented planar Gaussian disks. Unlike 3D\nGaussians, 2D Gaussians provide view-consistent geometry while modeling\nsurfaces intrinsically. To accurately recover thin surfaces and achieve stable\noptimization, we introduce a perspective-accurate 2D splatting process uti-\nlizing ray-splat intersection and rasterization. Additionally, we incorporate\ndepth distortion and normal consistency terms to further enhance the quality\nof the reconstructions. We demonstrate that our differentiable renderer al-\nlows for noise-free and detailed geometry reconstruction while maintaining\ncompetitive appearance quality, fast training speed, and real-time rendering.\nOur code will be made publicly available.\nCCS Concepts: • Computing methodologies →Reconstruction; Render-\ning; Machine learning approaches.\nAdditional Key Words and Phrases: novel view synthesis, radiance fields,\nsurface splatting, surface reconstruction\nAuthors’ addresses: Binbin Huang, ShanghaiTech University, China; Zehao Yu, Univer-\nsity of Tübingen, Tübingen AI Center, Germany; Anpei Chen, University of Tübingen,\nTübingen AI Center, Germany; Andreas Geiger, University of Tübingen, Tübingen AI\nCenter, Germany; Shenghua Gao, ShanghaiTech University, China.\n1\nINTRODUCTION\nPhotorealistic novel view synthesis (NVS) and accurate geometry\nreconstruction stand as pivotal long-term objectives in computer\ngraphics and vision. Recently, 3D Gaussian Splatting (3DGS) [Kerbl\net al. 2023] has emerged as an appealing alternative to implicit [Bar-\nron et al. 2022a; Mildenhall et al. 2020] and feature grid-based rep-\nresentations [Barron et al. 2023; Müller et al. 2022] in NVS, due to\nits real-time photorealistic NVS results at high resolutions. Rapidly\nevolving, 3DGS has been quickly extended with respect to multiple\ndomains, including anti-aliasing rendering [Yu et al. 2024], material\nmodeling [Jiang et al. 2023; Shi et al. 2023], dynamic scene recon-\nstruction [Yan et al. 2023], and animatable avatar creation [Qian\net al. 2023; Zielonka et al. 2023]. Nevertheless, it falls short in cap-\nturing intricate geometry since the volumetric 3D Gaussian, which\nmodels the complete angular radiance, conflicts with the thin nature\nof surfaces.\nOn the other hand, earlier works [Pfister et al. 2000; Zwicker et al.\n2001a] have shown surfels (surface elements) to be an effective rep-\nresentation for complex geometry. Surfels approximate the object\nsurface locally with shape and shade attributes, and can be derived\nfrom known geometry. They are widely used in SLAM [Whelan et al.\narXiv:2403.17888v1  [cs.CV]  26 Mar 2024\n\n\n2\n•\nBinbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao\n2016] and other robotics tasks [Schöps et al. 2019] as an efficient ge-\nometry representation. Subsequent advancements [Yifan et al. 2019]\nhave incorporated surfels into a differentiable framework. However,\nthese methods typically require ground truth (GT) geometry, depth\nsensor data, or operate under constrained scenarios with known\nlighting.\nInspired by these works, we propose 2D Gaussian Splatting for\n3D scene reconstruction and novel view synthesis that combines\nthe benefits of both worlds, while overcoming their limitations. Un-\nlike 3DGS, our approach represents a 3D scene with 2D Gaussian\nprimitives, each defining an oriented elliptical disk. The significant\nadvantage of 2D Gaussian over its 3D counterpart lies in the accu-\nrate geometry representation during rendering. Specifically, 3DGS\nevaluates a Gaussian’s value at the intersection between a pixel ray\nand a 3D Gaussian [Keselman and Hebert 2022, 2023], which leads\nto inconsistency depth when rendered from different viewpoints.\nIn contrast, our method utilizes explicit ray-splat intersection, re-\nsulting in a perspective accuracy splatting, as illustrated in Figure 2,\nwhich in turn significantly improves reconstruction quality. Fur-\nthermore, the inherent surface normals in 2D Gaussian primitives\nenable direct surface regularization through normal constraints. In\ncontrast with surfels-based models [Pfister et al. 2000; Yifan et al.\n2019; Zwicker et al. 2001a], our 2D Gaussians can be recovered from\nunknown geometry with gradient-based optimization.\nWhile our 2D Gaussian approach excels in geometric model-\ning, optimizing solely with photometric losses can lead to noisy\nreconstructions, due to the inherently unconstrained nature of 3D\nreconstruction tasks, as noted in [Barron et al. 2022b; Yu et al. 2022b;\nZhang et al. 2020]. To enhance reconstructions and achieve smoother\nsurfaces, we introduce two regularization terms: depth distortion\nand normal consistency. The depth distortion term concentrates 2D\nprimitives distributed within a tight range along the ray, address-\ning the rendering process’s limitation where the distance between\nGaussians is ignored. The normal consistency term minimizes dis-\ncrepancies between the rendered normal map and the gradient of\nthe rendered depth, ensuring alignment between the geometries\ndefined by depth and normals. Employing these regularizations\nin combination with our 2D Gaussian model enables us to extract\nhighly accurate surface meshes, as demonstrated in Figure 1.\nIn summary, we make the following contributions:\n• We present a highly efficient differentiable 2D Gaussian ren-\nderer, enabling perspective-accurate splatting by leveraging\n2D surface modeling, ray-splat intersection, and volumetric\nintegration.\n• We introduce two regularization losses for improved and\nnoise-free surface reconstruction.\n• Our approach achieves state-of-the-art geometry reconstruc-\ntion and NVS results compared to other explicit representa-\ntions.\n2\nRELATED WORK\n2.1\nNovel view synthesis\nSignificant advancements have been achieved in NVS, particularly\nsince the introduction of Neural Radiance Fields (NeRF) [Milden-\nhall et al. 2021]. NeRF employs a multi-layer perceptron (MLP) to\nFig. 2. Comparison of 3DGS and 2DGS. 3DGS utilizes different intersec-\ntion planes for value evaluation when viewing from different viewpoints,\nresulting in inconsistency. Our 2DGS provides multi-view consistent value\nevaluations.\nrepresent geometry and view-dependent appearance, optimized via\nvolume rendering to deliver exceptional rendering quality. Post-\nNeRF developments have further enhanced its capabilities. For in-\nstance, Mip-NeRF [Barron et al. 2021] and subsequent works [Barron\net al. 2022a, 2023; Hu et al. 2023] tackle NeRF’s aliasing issues. Ad-\nditionally, the rendering efficiency of NeRF has seen substantial\nimprovements through techniques such as distillation [Reiser et al.\n2021; Yu et al. 2021] and baking [Chen et al. 2023a; Hedman et al.\n2021; Reiser et al. 2023; Yariv et al. 2023]. Moreover, the training\nand representational power of NeRF have been enhanced using\nfeature-grid based scene representations [Chen et al. 2022, 2023c;\nFridovich-Keil et al. 2022; Liu et al. 2020; Müller et al. 2022; Sun et al.\n2022a].\nRecently, 3D Gaussian Splatting (3DGS) [Kerbl et al. 2023] has\nemerged, demonstrating impressive real-time NVS results. This\nmethod has been quickly extended to multiple domains [Xie et al.\n2023; Yan et al. 2023; Yu et al. 2024; Zielonka et al. 2023]. In this work,\nwe propose to “flatten” 3D Gaussians to 2D Gaussian primitives to\nbetter align their shape with the object surface. Combined with\ntwo novel regularization losses, our approach reconstructs surfaces\nmore accurately than 3DGS while preserving its high-quality and\nreal-time rendering capabilities.\n2.2\n3D reconstruction\n3D Reconstruction from multi-view images has been a long-standing\ngoal in computer vision. Multi-view stereo based methods [Schön-\nberger et al. 2016; Yao et al. 2018; Yu and Gao 2020] rely on a modular\npipeline that involves feature matching, depth prediction, and fu-\nsion. In contrast, recent neural approaches [Niemeyer et al. 2020;\nYariv et al. 2020] represent surface implicilty via an MLP [Mescheder\net al. 2019; Park et al. 2019] , extracting surfaces post-training via\nthe Marching Cube algorithm. Further advancements [Oechsle et al.\n2021; Wang et al. 2021; Yariv et al. 2021] integrated implicit surfaces\nwith volume rendering, achieving detailed surface reconstructions\nfrom RGB images. These methods have been extended to large-scale\nreconstructions via additional regularization [Li et al. 2023; Yu et al.\n2022a,b], and efficient reconstruction for objects [Wang et al. 2023].\nDespite these impressive developments, efficient large-scale scene\nreconstruction remains a challenge. For instance, Neuralangelo [Li\net al. 2023] requires 128 GPU hours for reconstructing a single scene\nfrom the Tanks and Temples Dataset [Knapitsch et al. 2017]. In this\nwork, we introduce 2D Gaussian splatting, a method that signifi-\ncantly accelerates the reconstruction process. It achieves similar or\n\n\n2D Gaussian Splatting for Geometrically Accurate Radiance Fields\n•\n3\nslightly better results compared to previous implicit neural surface\nrepresentations, while being an order of magnitude faster.\n2.3\nDifferentiable Point-based Graphics\nDifferentiable point-based rendering [Aliev et al. 2020; Insafutdinov\nand Dosovitskiy 2018; Rückert et al. 2022; Wiles et al. 2020; Yifan\net al. 2019]has been explored extensively due to its efficiency and\nflexibility in representing intricate structures. Notably, NPBG [Aliev\net al. 2020] rasterizes point cloud features onto an image plane,\nsubsequently utilizing a convolutional neural network for RGB\nimage prediction. DSS [Yifan et al. 2019] focuses on optimizing ori-\nented point clouds from multi-view images under known lighting\nconditions. Pulsar [Lassner and Zollhofer 2021] introduces a tile-\nbased acceleration structure for more efficient rasterization. More\nrecently, 3DGS [Kerbl et al. 2023] optimizes anisotropic 3D Gauss-\nian primitives, demonstrating real-time photorealistic NVS results.\nDespite these advances, using point-based representations from un-\nconstrained multi-view images remains challenging. In this paper,\nwe demonstrate detailed surface reconstruction using 2D Gaussian\nprimitives. We also highlight the critical role of additional regular-\nization losses in optimization, showcasing their significant impact\non the quality of the reconstruction.\n2.4\nConcurrent work\nSince 3DGS [Kerbl et al. 2023] was introduced, it has been rapidly\nadapted across multiple domains. We now review the closest work\nin inverse rendering. These work [Gao et al. 2023; Jiang et al. 2023;\nLiang et al. 2023; Shi et al. 2023] extend 3DGS by modeling normals\nas additional attributes of 3D Gaussian primitives. Our approach,\nin contrast, inherently defines normals by representing the tangent\nspace of the 3D surface using 2D Gaussian primitives, aligning them\nmore closely with the underlying geometry. Additionally, the afore-\nmentioned works predominantly focus on estimating the material\nproperties of the scene and evaluating their results for relighting\ntasks. Notably, none of these works specifically target surface re-\nconstruction, the primary focus of our work.\nWe also highlight the distinctions between our method and con-\ncurrent works SuGaR [Guédon and Lepetit 2023] and NeuSG [Chen\net al. 2023b]. Unlike SuGaR, which approximates 2D Gaussians with\n3D Gaussians, our method directly employs 2D Gaussians, simpli-\nfying the process and enhancing the resulting geometry without\nadditional mesh refinement. NeuSG optimizes 3D Gaussian primi-\ntives and an implicit SDF network jointly and extracts the surface\nfrom the SDF network, while our approach leverages 2D Gaussian\nprimitives for surface approximation, offering a faster and concep-\ntually simpler solution.\n3\n3D GAUSSIAN SPLATTING\nKerbl et al. [Kerbl et al. 2023] propose to represent 3D scenes with\n3D Gaussian primitives and render images using differentiable vol-\nume splatting. Specifically, 3DGS explicitly parameterizes Gaussian\nprimitives via 3D covariance matrix 𝚺and their location p𝑘:\nG(p) = exp(−1\n2 (p −p𝑘)⊤𝚺−1(p −p𝑘))\n(1)\nwhere the covariance matrix 𝚺= RSS⊤R⊤is factorized into a scal-\ning matrix S and a rotation matrix R. To render an image, the 3D\nGaussian is transformed into the camera coordinates with world-\nto-camera transform matrix W and projected to image plane via a\nlocal affine transformation J [Zwicker et al. 2001a]:\n𝚺\n′ = JW𝚺W⊤J⊤\n(2)\nBy skipping the third row and column of 𝚺\n′, we obtain a 2D Gaussian\nG2𝐷with covariance matrix 𝚺2𝐷. Next, 3DGS [Kerbl et al. 2023]\nemploys volumetric alpha blending to integrate alpha-weighted\nappearance from front to back:\nc(x) =\n𝐾\n∑︁\n𝑘=1\nc𝑘𝛼𝑘G2𝐷\n𝑘\n(x)\n𝑘−1\nÖ\n𝑗=1\n(1 −𝛼𝑗G2𝐷\n𝑗\n(x))\n(3)\nwhere 𝑘is the index of the Gaussian primitives, 𝛼𝑘denotes the alpha\nvalues and c𝑘is the view-dependent appearance. The attributes of\n3D Gaussian primitives are optimized using a photometric loss.\nChallenges in Surface Reconstruction: Reconstructing surfaces\nusing 3D Gaussian modeling and splatting faces several challenges.\nFirst, the volumetric radiance representation of 3D Gaussians con-\nflicts with the thin nature of surfaces. Second, 3DGS does not na-\ntively model surface normals, essential for high-quality surface\nreconstruction. Third, the rasterization process in 3DGS lacks multi-\nview consistency, leading to varied 2D intersection planes for dif-\nferent viewpoints [Keselman and Hebert 2023], as illustrated in\nFigure 2 (a). Additionally, using an affine matrix for transforming a\n3D Gaussian into ray space only yields accurate projections near the\ncenter, compromising on perspective accuracy around surrounding\nregions [Zwicker et al. 2004]. Therefore, it often results in noisy\nreconstructions, as shown in Figure 5.\n4\n2D GAUSSIAN SPLATTING\nTo accurately reconstruct geometry while maintaining high-quality\nnovel view synthesis, we present differentiable 2D Gaussian splat-\nting (2DGS).\n4.1\nModeling\nUnlike 3DGS [Kerbl et al. 2023], which models the entire angular\nradiance in a blob, we simplify the 3-dimensional modeling by adopt-\ning “flat” 2D Gaussians embedded in 3D space. With 2D Gaussian\nmodeling, the primitive distributes densities within a planar disk,\ndefining the normal as the direction of steepest change of density.\nThis feature enables better alignment with thin surfaces. While pre-\nvious methods [Kopanas et al. 2021; Yifan et al. 2019] also utilize 2D\nGaussians for geometry reconstruction, they require a dense point\ncloud or ground-truth normals as input. By contrast, we simultane-\nously reconstruct the appearance and geometry given only a sparse\ncalibration point cloud and photometric supervision.\nAs illustrated in Figure 3, our 2D splat is characterized by its\ncentral point p𝑘, two principal tangential vectors t𝑢and t𝑣, and\na scaling vector S = (𝑠𝑢,𝑠𝑣) that controls the variances of the 2D\nGaussian. Notice that the primitive normal is defined by two orthog-\nonal tangential vectors t𝑤= t𝑢× t𝑣. We can arrange the orientation\ninto a 3 × 3 rotation matrix R = [t𝑢, t𝑣, t𝑤] and the scaling factors\ninto a 3 × 3 diagonal matrix S whose last entry is zero.\n\n\n4\n•\nBinbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao\n𝑠!𝐭!\n𝑠\"𝐭\"\n2D Gaussian Splat \nin object space\n2D Gaussian Splat \nin image space\nTangent frame (u,v)\nImage frame (x,y)\n𝐩!\n𝑠\"𝐭\"\n𝑠#𝐭#\nFig. 3. Illustration of 2D Gaussian Splatting. 2D Gaussian Splats are ellip-\ntical disks characterized by a center point p𝑘, tangential vectors t𝑢and\nt𝑣, and two scaling factors (𝑠𝑢and 𝑠𝑣) control the variance. Their elliptical\nprojections are sampled through the ray-splat intersection ( Section 4.2) and\naccumulated via alpha-blending in image space. 2DGS reconstructs surface\nattributes such as colors, depths, and normals through gradient descent.\nA 2D Gaussian is therefore defined in a local tangent plane in\nworld space, which is parameterized:\n𝑃(𝑢, 𝑣) = p𝑘+ 𝑠𝑢t𝑢𝑢+ 𝑠𝑣t𝑣𝑣= H(𝑢, 𝑣, 1, 1)⊤\n(4)\nwhere H =\n\u0014𝑠𝑢t𝑢\n𝑠𝑣t𝑣\n0\np𝑘\n0\n0\n0\n1\n\u0015\n=\n\u0014RS\np𝑘\n0\n1\n\u0015\n(5)\nwhere H ∈4 × 4 is a homogeneous transformation matrix repre-\nsenting the geometry of the 2D Gaussian. For the point u = (𝑢, 𝑣) in\n𝑢𝑣space, its 2D Gaussian value can then be evaluated by standard\nGaussian\nG(u) = exp\n\u0012\n−𝑢2 + 𝑣2\n2\n\u0013\n(6)\nThe center p𝑘, scaling (𝑠𝑢,𝑠𝑣), and the rotation (t𝑢, t𝑣) are learnable\nparameters. Following 3DGS [Kerbl et al. 2023], each 2D Gaussian\nprimitive has opacity 𝛼and view-dependent appearance 𝑐parame-\nterized with spherical harmonics.\n4.2\nSplatting\nOne common strategy for rendering 2D Gaussians is to project the\n2D Gaussian primitives onto the image space using the affine ap-\nproximation for the perspective projection [Zwicker et al. 2001a].\nHowever, as noted in [Zwicker et al. 2004], this projection is only\naccurate at the center of the Gaussian and has increasing approxi-\nmation error with increased distance to the center. To address this\nissue, Zwicker et al. proposed a formulation based on homogeneous\ncoordinates. Specifically, projecting the 2D splat onto an image plane\ncan be described by a general 2D-to-2D mapping in homogeneous\ncoordinates. Let W ∈4 × 4 be the transformation matrix from world\nspace to screen space. The screen space points are hence obtained\nby\nx = (𝑥𝑧,𝑦𝑧,𝑧,𝑧)⊤= W𝑃(𝑢, 𝑣) = WH(𝑢, 𝑣, 1, 1)⊤\n(7)\nwhere x represents a homogeneous ray emitted from the camera and\npassing through pixel (𝑥,𝑦) and intersecting the splat at depth 𝑧. To\nrasterize a 2D Gaussian, Zwicker et al. proposed to project its conic\ninto the screen space with an implicit method using M = (WH)−1.\nHowever, the inverse transformation introduces numerical insta-\nbility especially when the splat degenerates into a line segment\n(i.e., if it is viewed from the side). To address this issue, previous\nsurface splatting rendering methods discard such ill-conditioned\ntransformations using a pre-defined threshold. However, such a\nscheme poses challenges within a differentiable rendering frame-\nwork, as thresholding can lead to unstable optimization. To address\nthis problem, we utilize an explicit ray-splat intersection inspired\nby [Sigg et al. 2006].\nRay-splat Intersection: We efficiently locate the ray-splat inter-\nsections by finding the intersection of three non-parallel planes, a\nmethod initially designed for specialized hardware [Weyrich et al.\n2007]. Given an image coordinate x = (𝑥,𝑦), we parameterize the\nray of a pixel as the intersection of two orthogonal planes: the\nx-plane and the y-plane. Specifically, the x-plane is defined by a\nnormal vector (−1, 0, 0) and an offset 𝑥. Therefore, the x-plane can\nbe represented as a 4D homogeneous plane h𝑥= (−1, 0, 0,𝑥). Sim-\nilarly, the y-plane is h𝑦= (0, −1, 0,𝑦). Thus, the ray x = (𝑥,𝑦) is\ndetermined by the intersection of the 𝑥-plane and the 𝑦-planes.\nNext, we transform both planes into the local coordinates of\nthe 2D Gaussian primitives, the 𝑢𝑣-coordinate system. Note that\ntransforming points on a plane using a transformation matrix M\nis equivalent to transforming the homogeneous plane parameters\nusing the inverse transpose M−⊤[Vince 2008]. Therefore, applying\nM = (WH)−1 is equivalent to (WH)⊤, eliminating explicit matrix\ninversion and yielding:\nh𝑢= (WH)⊤h𝑥\nh𝑣= (WH)⊤h𝑦\n(8)\nAs introduced in Section 4.1, points on the 2D Gaussian plane are\nrepresented as (𝑢, 𝑣, 1, 1). At the same time, the intersection point\nshould fall on the transformed 𝑥-plane and 𝑦-plane. Thus,\nh𝑢· (𝑢, 𝑣, 1, 1)⊤= h𝑣· (𝑢, 𝑣, 1, 1)⊤= 0\n(9)\nThis leads to an efficient solution for the intersection point u(x):\n𝑢(x) = h2\n𝑢h4\n𝑣−h4\n𝑢h2\n𝑣\nh1\n𝑢h2\n𝑣−h2\n𝑢h1\n𝑣\n𝑣(x) = h4\n𝑢h1\n𝑣−h1\n𝑢h4\n𝑣\nh1\n𝑢h2\n𝑣−h2\n𝑢h1\n𝑣\n(10)\nwhere h𝑖\n𝑢, h𝑖\n𝑣are the 𝑖-th parameters of the 4D homogeneous plane\nparameters. We obtain the depth 𝑧of the intersected points via Eq. 7\nand evaluate the Gaussian value with Eq 6.\nDegenerate Solutions: When a 2D Gaussian is observed from a\nslanted viewpoint, it degenerates to a line in screen space. Therefore,\nit might be missed during rasterization. To deal with these cases and\nstabilize optimization, we employ the object-space low-pass filter\nintroduced in [Botsch et al. 2005]:\nˆ\nG(x) = max\nn\nG(u(x)), G( x −c\n𝜎\n)\no\n(11)\nwhere u(x) is given by (10) and c is the projection of center p𝑘.\nIntuitively, ˆ\nG(x) is lower-bounded by a fixed screen-space Gaussian\nlow-pass filter with center c𝑘and radius 𝜎. In our experiments, we\nset 𝜎=\n√\n2/2 to ensure sufficient pixels are used during rendering.\nRasterization: We follow a similar rasterization process as in\n3DGS [Kerbl et al. 2023]. First, a screen space bounding box is com-\nputed for each Gaussian primitive. Then, 2D Gaussians are sorted\nbased on the depth of their center and organized into tiles based on\ntheir bounding boxes. Finally, volumetric alpha blending is used to\n\n\n2D Gaussian Splatting for Geometrically Accurate Radiance Fields\n•\n5\nintegrate alpha-weighted appearance from front to back:\nc(x) =\n∑︁\n𝑖=1\nc𝑖𝛼𝑖ˆ\nG𝑖(u(x))\n𝑖−1\nÖ\n𝑗=1\n(1 −𝛼𝑗ˆ\nG𝑗(u(x)))\n(12)\nThe iterative process is terminated when the accumulated opacity\nreaches saturation.\n5\nTRAINING\nOur 2D Gaussian method, while effective in geometric modeling, can\nresult in noisy reconstructions when optimized only with photomet-\nric losses, a challenge inherent to 3D reconstruction tasks [Barron\net al. 2022b; Yu et al. 2022b; Zhang et al. 2020]. To mitigate this\nissue and improve the geometry reconstruction, we introduce two\nregularization terms: depth distortion and normal consistency.\nDepth Distortion: Different from NeRF, 3DGS’s volume rendering\ndoesn’t consider the distance between intersected Gaussian primi-\ntives. Therefore, spread out Gaussians might result in a similar color\nand depth rendering. This is different from surface rendering, where\nrays intersect the first visible surface exactly once. To mitigate this\nissue, we take inspiration from Mip-NeRF360 [Barron et al. 2022a]\nand propose a depth distortion loss to concentrate the weight dis-\ntribution along the rays by minimizing the distance between the\nray-splat intersections:\nL𝑑=\n∑︁\n𝑖,𝑗\n𝜔𝑖𝜔𝑗|𝑧𝑖−𝑧𝑗|\n(13)\nwhere 𝜔𝑖= 𝛼𝑖ˆ\nG𝑖(u(x)) Î𝑖−1\n𝑗=1(1 −𝛼𝑗ˆ\nG𝑗(u(x))) is the blending\nweight of the 𝑖−th intersection and 𝑧𝑖is the depth of the intersection\npoints. Unlike the distortion loss in Mip-NeRF360, where 𝑧𝑖is the\ndistance between sampled points and is not optimized, our approach\ndirectly encourages the concentration of the splats by adjusting the\nintersection depth 𝑧𝑖. Note that we implement this regularization\nterm efficiently with CUDA in a manner similar to [Sun et al. 2022b].\nNormal Consistency: As our representation is based on 2D Gauss-\nian surface elements, we must ensure that all 2D splats are locally\naligned with the actual surfaces. In the context of volume rendering\nwhere multiple semi-transparent surfels may exist along the ray,\nwe consider the actual surface at the median point of intersection,\nwhere the accumulated opacity reaches 0.5. We then align the splats’\nnormal with the gradients of the depth maps as follows:\nL𝑛=\n∑︁\n𝑖\n𝜔𝑖(1 −n⊤\n𝑖N)\n(14)\nwhere 𝑖indexes over intersected splats along the ray, 𝜔denotes the\nblending weight of the intersection point, n𝑖represents the normal\nof the splat that is oriented towards the camera, and N is the normal\nestimated by the nearby depth point p. Specifically, N is computed\nwith finite differences as follows:\nN(𝑥,𝑦) = ∇𝑥p × ∇𝑦p\n|∇𝑥p × ∇𝑦p|\n(15)\nBy aligning the splat normal with the estimated surface normal, we\nensure that 2D splats locally approximate the actual object surface.\nFinal Loss: Finally, we optimize our model from an initial sparse\npoint cloud using a set of posed images. We minimize the following\nloss function:\nL = L𝑐+ 𝛼L𝑑+ 𝛽L𝑛\n(16)\nwhere L𝑐is an RGB reconstruction loss combining L1 with the\nD-SSIM term from [Kerbl et al. 2023], while L𝑑and L𝑛are regu-\nlarization terms. We set 𝛼= 1000 for bounded scenes, 𝛼= 100 for\nunbounded scenes and 𝛽= 0.05 for all scenes.\n6\nEXPERIMENTS\nWe now present an extensive evaluation of our 2D Gaussian Splat-\nting reconstruction method, including appearance and geometry\ncomparison with previous state-of-the-art implicit and explicit ap-\nproaches. We then analyze the contribution of the proposed compo-\nnents.\n6.1\nImplementation\nWe implement our 2D Gaussian Splatting with custom CUDA ker-\nnels, building upon the framework of 3DGS [Kerbl et al. 2023]. We\nextend the renderer to output depth distortion maps, depth maps\nand normal maps for regularizations. During training, we increase\nthe number of 2D Gaussian primitives following the adaptive con-\ntrol strategy in 3DGS. Since our method does not directly rely on\nthe gradient of the projected 2D center, we hence project the gra-\ndient of 3D center p𝑘onto the screen space as an approximation.\nSimilarly, we employ a gradient threshold of 0.0002 and remove\nsplats with opacity lower than 0.05 every 3000 step. We conduct all\nthe experiments on a single GTX RTX3090 GPU.\nMesh Extraction: To extract meshes from reconstructed 2D splats,\nwe render depth maps of the training views using the median depth\nvalue of the splats projected to the pixels and utilize truncated signed\ndistance fusion (TSDF) to fuse the reconstruction depth maps, using\nOpen3D [Zhou et al. 2018]. We set the voxel size to 0.004 and the\ntruncated threshold to 0.02 during TSDF fusion. We also extend the\noriginal 3DGS to render depth and employ the same technique for\nsurface reconstruction for a fair comparison.\n6.2\nComparison\nDataset: We evaluate the performance of our method on vari-\nous datasets, including DTU [Jensen et al. 2014], Tanks and Tem-\nples [Knapitsch et al. 2017], and Mip-NeRF360 [Barron et al. 2022a].\nThe DTU dataset comprises 15 scenes, each with 49 or 69 images of\nresolution 1600 × 1200. We use Colmap [Schönberger and Frahm\n2016] to generate a sparse point cloud for each scene and down-\nsample the images into a resolution of 800 × 600 for efficiency. We\nuse the same training process for 3DGS [Kerbl et al. 2023] and\nSuGar [Guédon and Lepetit 2023] for a fair comparison.\nGeometry Reconstruction: In Table 1 and Table 3, we compare\nour geometry reconstruction to SOTA implicit (i.e., NeRF [Milden-\nhall et al. 2020], VolSDF [Yariv et al. 2021], and NeuS [Wang et al.\n2021]), explicit (i.e., 3DGS [Kerbl et al. 2023] and concurrent work\nSuGaR [Guédon and Lepetit 2023]) on Chamfer distance and train-\ning time using the DTU dataset. Our method outperforms all com-\npared methods in terms of Chamfer distance. Moreover, as shown\nin Table 2, 2DGS achieves competitive results with SDF models\n(i.e., NeuS [Wang et al. 2021] and Geo-Neus [Fu et al. 2022]) on the\n\n\n6\n•\nBinbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao\nGround truth\nOurs (color)\nOurs (normal)\n3DGS\nSuGaR\nOurs\nFig. 4. Visual comparisons (test-set view) between our method, 3DGS, and SuGaR using scenes from an unbounded real-world dataset - kitchen, bicycle,\nstump, and treehill. The first column displays the ground truth novel view, while the second and the third columns showcase our rendering results,\ndemonstrating high-fidelity novel view and accurate geometry. Columns four and five present surface reconstruction results from our method and 3DGS,\nutilizing TSDF fusion with a depth range of 6 meters. Background regions (outside this range) are colored light gray. The sixth column displays results from\nSuGaR. In contrast to both baselines, our method excels in capturing sharp edges and intricate geometry.\nTable 1. Quantitative comparison on the DTU Dataset [Jensen et al. 2014]. Our 2DGS achieves the highest reconstruction accuracy among other methods and\nprovides 100× speed up compared to the SDF based baselines.\n24\n37\n40\n55\n63\n65\n69\n83\n97\n105\n106\n110\n114\n118\n122\nMean\nTime\nimplicit\nNeRF [Mildenhall et al. 2021]\n1.90\n1.60\n1.85\n0.58\n2.28\n1.27\n1.47\n1.67\n2.05\n1.07\n0.88\n2.53\n1.06\n1.15\n0.96\n1.49\n> 12h\nVolSDF [Yariv et al. 2021]\n1.14\n1.26\n0.81\n0.49\n1.25\n0.70\n0.72\n1.29\n1.18\n0.70\n0.66\n1.08\n0.42\n0.61\n0.55\n0.86\n>12h\nNeuS [Wang et al. 2021]\n1.00\n1.37\n0.93\n0.43\n1.10\n0.65\n0.57\n1.48\n1.09\n0.83\n0.52\n1.20\n0.35\n0.49\n0.54\n0.84\n>12h\nexplicit\n3DGS [Kerbl et al. 2023]\n2.14\n1.53\n2.08\n1.68\n3.49\n2.21\n1.43\n2.07\n2.22\n1.75\n1.79\n2.55\n1.53\n1.52\n1.50\n1.96\n11.2 m\nSuGaR [Guédon and Lepetit 2023]\n1.47\n1.33\n1.13\n0.61\n2.25\n1.71\n1.15\n1.63\n1.62\n1.07\n0.79\n2.45\n0.98\n0.88\n0.79\n1.33\n∼1h\n2DGS-15k (Ours)\n0.48\n0.92\n0.42\n0.40\n1.04\n0.83\n0.83\n1.36\n1.27\n0.76\n0.72\n1.63\n0.40\n0.76\n0.60\n0.83\n5.5 m\n2DGS-30k (Ours)\n0.48\n0.91\n0.39\n0.39\n1.01\n0.83\n0.81\n1.36\n1.27\n0.76\n0.70\n1.40\n0.40\n0.76\n0.52\n0.80\n18.8 m\nTable 2. Quantitative results on the Tanks and Temples Dataset [Knapitsch\net al. 2017]. We report the F1 score and training time.\nNeuS\nGeo-Neus\nNeurlangelo\nSuGaR\n3DGS\nOurs\nBarn\n0.29\n0.33\n0.70\n0.14\n0.13\n0.36\nCaterpillar\n0.29\n0.26\n0.36\n0.16\n0.08\n0.23\nCourthouse\n0.17\n0.12\n0.28\n0.08\n0.09\n0.13\nIgnatius\n0.83\n0.72\n0.89\n0.33\n0.04\n0.44\nMeetingroom\n0.24\n0.20\n0.32\n0.15\n0.01\n0.16\nTruck\n0.45\n0.45\n0.48\n0.26\n0.19\n0.26\nMean\n0.38\n0.35\n0.50\n0.19\n0.09\n0.30\nTime\n>24h\n>24h\n>24h\n>1h\n14.3 m\n34.2 m\nTnT dataset, and significantly better reconstruction than explicit\nreconstruction methods (i.e., 3DGS and SuGaR). Notably, our model\ndemonstrates exceptional efficiency, offering a reconstruction speed\nthat is approximately 100 times faster compared to implicit recon-\nstruction methods and more than 3 times faster than the concurrent\nTable 3. Performance comparison between 2DGS (ours), 3DGS and SuGaR\non the DTU dataset [Jensen et al. 2014]. We report the averaged chamfer\ndistance, PSNR, reconstruction time, and model size.\nCD ↓\nPSNR ↑\nTime ↓\nMB (Storage) ↓\n3DGS [Kerbl et al. 2023]\n1.96\n35.76\n11.2 m\n113\nSuGaR [Guédon and Lepetit 2023]\n1.33\n34.57\n∼1 h\n1247\n2DGS-15k (Ours)\n0.83\n33.42\n5.5 m\n52\n2DGS-30k (Ours)\n0.80\n34.52\n18.8 m\n52\nwork SuGaR. Our approach can also achieve qualitatively better\nreconstructions with more appearance and geometry details and\nfewer outliers, as shown in Figure 5. We include the geometry recon-\nstruction results in supplementary Figure 10. Moreover, SDF-based\nreconstruction methods require predefining the spherical size for\ninitialization, which plays a critical role in the success of SDF recon-\nstruction. Bu contrast, our method leverages radiance field based\ngeometry modeling and is less sensitive to initialization.\n\n\n2D Gaussian Splatting for Geometrically Accurate Radiance Fields\n•\n7\nFig. 5. Qualitative comparison on the DTU benchmark [Jensen et al. 2014].\nOur 2DGS produces detailed and noise-free surfaces.\nTable 4. Quantitative results on Mip-NeRF 360 [Barron et al. 2022a] dataset.\nAll scores of the baseline methods are directly taken from their papers\nwhenever available. We report the performance of 3DGS, SuGaR and ours\nusing 30𝑘iterations.\nOutdoor Scene\nIndoor scene\nPSNR ↑\nSSIM ↑\nLIPPS ↓\nPSNR ↑\nSSIM ↑\nLIPPS ↓\nNeRF\n21.46\n0.458\n0.515\n26.84\n0.790\n0.370\nDeep Blending\n21.54\n0.524\n0.364\n26.40\n0.844\n0.261\nInstant NGP\n22.90\n0.566\n0.371\n29.15\n0.880\n0.216\nMERF\n23.19\n0.616\n0.343\n27.80\n0.855\n0.271\nMipNeRF360\n24.47\n0.691\n0.283\n31.72\n0.917\n0.180\nBakedSDF\n22.47\n0.585\n0.349\n27.06\n0.836\n0.258\nMobile-NeRF\n21.95\n0.470\n0.470\n-\n-\n-\n3DGS\n24.24\n0.705\n0.283\n30.99\n0.926\n0.199\nSuGaR\n22.76\n0.631\n0.349\n29.44\n0.911\n0.216\n2DGS (Ours)\n24.33\n0.709\n0.284\n30.39\n0.924\n0.182\nAppearance Reconstruction: Our method represents 3D scenes\nas radiance fields, providing high-quality novel view synthesis. In\nthis section, we compare our novel view renderings using the Mip-\nNeRF360 dataset against baseline approaches, as shown in Table 4\nand Figure 4. Note that, since the ground truth geometry is not\navailable in the Mip-NeRF360 dataset and we hence focus on quan-\ntitative comparison. Remarkably, our method consistently achieves\ncompetitive NVS results across state-of-the-art techniques while\nproviding geometrically accurate surface reconstruction.\n6.3\nAblations\nIn this section, we isolate the design choices and measure their effect\non reconstruction quality, including regularization terms and mesh\nextraction. We conduct experiments on the DTU dataset [Jensen et al.\n2014] with 15𝑘iterations and report the reconstruction accuracy,\ncompleteness and average reconstruction quality. The quantitative\neffect of the choices is reported in Table 5.\nRegularization: We first examine the effects of the proposed nor-\nmal consistency and depth distortion regularization terms. Our\nmodel (Table 5 E) provides the best performance when applying\nboth regularization terms. We observe that disabling the normal\nconsistency (Table 5 A) can lead to incorrect orientation, as shown\nin Figure 6 A. Additionally, the absence of depth distortion (Table 5\nB) results in a noisy surface, as shown in Figure 6 B.\nTable 5. Quantitative studies for the regularization terms and mesh extrac-\ntion methods on the DTU dataset.\nAccuracy ↓\nCompletion ↓\nAveragy ↓\nA. w/o normal consistency\n1.35\n1.13\n1.24\nB. w/o depth distortion\n0.89\n0.87\n0.88\nC. w / expected depth\n0.88\n1.01\n0.94\nD. w / Poisson\n1.25\n0.89\n1.07\nE. Full Model\n0.79\n0.86\n0.83\nInput\n(A) w/o. NC\n(B) w/o. DD\nFull Model\nFig. 6. Qualitative studies for the regularization effects. From left to right –\ninput image, surface normals without normal consistency, without depth\ndistortion, and our full model. Disabling the normal consistency loss leads\nto noisy surface orientations; conversely, omitting depth distortion regular-\nization results in blurred surface normals. The complete model, employing\nboth regularizations, successfully captures sharp and flat features.\nMesh Extraction: We now analyze our choice for mesh extraction.\nOur full model (Table 5 E) utilizes TSDF fusion for mesh extraction\nwith median depth. One alternative option is to use the expected\ndepth instead of the median depth. However, it yields worse recon-\nstructions as it is more sensitive to outliers, as shown in Table 5\nC. Further, our approach surpasses screened Poisson surface recon-\nstruction (SPSR)(Table 5 D) [Kazhdan and Hoppe 2013] using 2D\nGaussians’ center and normal as inputs, due to SPSR’s inability to\nincorporate the opacity and the size of 2D Gaussian primitives.\n7\nCONCLUSION\nWe presented 2D Gaussian splatting, a novel approach for geomet-\nrically accurate radiance field reconstruction. We utilized 2D Gauss-\nian primitives for 3D scene representation, facilitating accurate and\nview consistent geometry modeling and rendering. We proposed\ntwo regularization techniques to further enhance the reconstructed\ngeometry. Extensive experiments on several challenging datasets\nverify the effectiveness and efficiency of our method.\nLimitations: While our method successfully delivers accurate ap-\npearance and geometry reconstruction for a wide range of objects\nand scenes, we also discuss its limitations: First, we assume surfaces\nwith full opacity and extract meshes from multi-view depth maps.\nThis can pose challenges in accurately handling semi-transparent\nsurfaces, such as glass, due to their complex light transmission prop-\nerties. Secondly, our current densification strategy favors texture-\nrich over geometry-rich areas, occasionally leading to less accurate\nrepresentations of fine geometric structures. A more effective densi-\nfication strategy could mitigate this issue. Finally, our regularization\noften involves a trade-off between image quality and geometry, and\ncan potentially lead to oversmoothing in certain regions.\n\n\n8\n•\nBinbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao\nAcknowledgement: BH and SG are supported by NSFC #62172279,\n#61932020, Program of Shanghai Academic Research Leader. ZY, AC\nand AG are supported by the ERC Starting Grant LEGO-3D (850533)\nand DFG EXC number 2064/1 - project number 390727645.\nREFERENCES\nKara-Ali Aliev, Artem Sevastopolsky, Maria Kolos, Dmitry Ulyanov, and Victor Lem-\npitsky. 2020. Neural point-based graphics. In Computer Vision–ECCV 2020: 16th\nEuropean Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXII 16.\nSpringer, 696–712.\nJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-\nBrualla, and Pratul P. Srinivasan. 2021. Mip-NeRF: A Multiscale Representation for\nAnti-Aliasing Neural Radiance Fields. ICCV (2021).\nJonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman.\n2022a. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings\nof the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5470–5479.\nJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman.\n2022b. Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields. CVPR (2022).\nJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman.\n2023. Zip-NeRF: Anti-Aliased Grid-Based Neural Radiance Fields. ICCV (2023).\nMario Botsch, Alexander Hornung, Matthias Zwicker, and Leif Kobbelt. 2005. High-\nquality surface splatting on today’s GPUs. In Proceedings Eurographics/IEEE VGTC\nSymposium Point-Based Graphics, 2005. IEEE, 17–141.\nAnpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. 2022. TensoRF:\nTensorial Radiance Fields. In European Conference on Computer Vision (ECCV).\nHanlin Chen, Chen Li, and Gim Hee Lee. 2023b. NeuSG: Neural Implicit Surface Re-\nconstruction with 3D Gaussian Splatting Guidance. arXiv preprint arXiv:2312.00846\n(2023).\nZhiqin Chen, Thomas Funkhouser, Peter Hedman, and Andrea Tagliasacchi. 2023a.\nMobilenerf: Exploiting the polygon rasterization pipeline for efficient neural field\nrendering on mobile architectures. In Proceedings of the IEEE/CVF Conference on\nComputer Vision and Pattern Recognition. 16569–16578.\nZhang Chen, Zhong Li, Liangchen Song, Lele Chen, Jingyi Yu, Junsong Yuan, and Yi\nXu. 2023c. NeuRBF: A Neural Fields Representation with Adaptive Radial Basis\nFunctions. In Proceedings of the IEEE/CVF International Conference on Computer\nVision. 4182–4194.\nSara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and\nAngjoo Kanazawa. 2022. Plenoxels: Radiance Fields without Neural Networks. In\nCVPR.\nQiancheng Fu, Qingshan Xu, Yew-Soon Ong, and Wenbing Tao. 2022. Geo-Neus:\nGeometry-Consistent Neural Implicit Surfaces Learning for Multi-view Reconstruc-\ntion. Advances in Neural Information Processing Systems (NeurIPS) (2022).\nJian Gao, Chun Gu, Youtian Lin, Hao Zhu, Xun Cao, Li Zhang, and Yao Yao. 2023. Re-\nlightable 3D Gaussian: Real-time Point Cloud Relighting with BRDF Decomposition\nand Ray Tracing. arXiv:2311.16043 (2023).\nAntoine Guédon and Vincent Lepetit. 2023. SuGaR: Surface-Aligned Gaussian Splatting\nfor Efficient 3D Mesh Reconstruction and High-Quality Mesh Rendering. arXiv\npreprint arXiv:2311.12775 (2023).\nPeter Hedman, Pratul P Srinivasan, Ben Mildenhall, Jonathan T Barron, and Paul De-\nbevec. 2021. Baking neural radiance fields for real-time view synthesis. In Proceedings\nof the IEEE/CVF International Conference on Computer Vision. 5875–5884.\nWenbo Hu, Yuling Wang, Lin Ma, Bangbang Yang, Lin Gao, Xiao Liu, and Yuewen\nMa. 2023. Tri-MipRF: Tri-Mip Representation for Efficient Anti-Aliasing Neural\nRadiance Fields. In ICCV.\nEldar Insafutdinov and Alexey Dosovitskiy. 2018. Unsupervised learning of shape and\npose with differentiable point clouds. Advances in neural information processing\nsystems 31 (2018).\nRasmus Jensen, Anders Dahl, George Vogiatzis, Engin Tola, and Henrik Aanæs. 2014.\nLarge scale multi-view stereopsis evaluation. In Proceedings of the IEEE conference\non computer vision and pattern recognition. 406–413.\nYingwenqi Jiang, Jiadong Tu, Yuan Liu, Xifeng Gao, Xiaoxiao Long, Wenping Wang, and\nYuexin Ma. 2023. GaussianShader: 3D Gaussian Splatting with Shading Functions\nfor Reflective Surfaces. arXiv preprint arXiv:2311.17977 (2023).\nMichael Kazhdan and Hugues Hoppe. 2013. Screened poisson surface reconstruction.\nACM Transactions on Graphics (ToG) 32, 3 (2013), 1–13.\nBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 2023.\n3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Transactions on\nGraphics 42, 4 (July 2023). https://repo-sam.inria.fr/fungraph/3d-gaussian-splatting/\nLeonid Keselman and Martial Hebert. 2022. Approximate differentiable rendering with\nalgebraic surfaces. In European Conference on Computer Vision. Springer, 596–614.\nLeonid Keselman and Martial Hebert. 2023. Flexible techniques for differentiable\nrendering with 3d gaussians. arXiv preprint arXiv:2308.14737 (2023).\nArno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. 2017. Tanks and\nTemples: Benchmarking Large-Scale Scene Reconstruction. ACM Transactions on\nGraphics 36, 4 (2017).\nGeorgios Kopanas, Julien Philip, Thomas Leimkühler, and George Drettakis. 2021. Point-\nBased Neural Rendering with Per-View Optimization. In Computer Graphics Forum,\nVol. 40. Wiley Online Library, 29–43.\nChristoph Lassner and Michael Zollhofer. 2021. Pulsar: Efficient sphere-based neural\nrendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern\nRecognition. 1440–1449.\nZhaoshuo Li, Thomas Müller, Alex Evans, Russell H Taylor, Mathias Unberath, Ming-\nYu Liu, and Chen-Hsuan Lin. 2023. Neuralangelo: High-Fidelity Neural Surface\nReconstruction. In IEEE Conference on Computer Vision and Pattern Recognition\n(CVPR).\nZhihao Liang, Qi Zhang, Ying Feng, Ying Shan, and Kui Jia. 2023. GS-IR: 3D Gaussian\nSplatting for Inverse Rendering. arXiv preprint arXiv:2311.16473 (2023).\nLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Christian Theobalt. 2020.\nNeural Sparse Voxel Fields. NeurIPS (2020).\nJonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. 2024. Dynamic\n3D Gaussians: Tracking by Persistent Dynamic View Synthesis. In 3DV.\nLars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas\nGeiger. 2019. Occupancy Networks: Learning 3D Reconstruction in Function Space.\nIn Conference on Computer Vision and Pattern Recognition (CVPR).\nBen Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ra-\nmamoorthi, and Ren Ng. 2020. NeRF: Representing Scenes as Neural Radiance Fields\nfor View Synthesis. In ECCV.\nBen Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ra-\nmamoorthi, and Ren Ng. 2021. Nerf: Representing scenes as neural radiance fields\nfor view synthesis. Commun. ACM 65, 1 (2021), 99–106.\nThomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. 2022. Instant\nNeural Graphics Primitives with a Multiresolution Hash Encoding. ACM Trans.\nGraph. 41, 4, Article 102 (July 2022), 15 pages.\nMichael Niemeyer, Lars Mescheder, Michael Oechsle, and Andreas Geiger. 2020. Differ-\nentiable Volumetric Rendering: Learning Implicit 3D Representations without 3D\nSupervision. In Conference on Computer Vision and Pattern Recognition (CVPR).\nMichael Oechsle, Songyou Peng, and Andreas Geiger. 2021. UNISURF: Unifying Neural\nImplicit Surfaces and Radiance Fields for Multi-View Reconstruction. In International\nConference on Computer Vision (ICCV).\nJeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Love-\ngrove. 2019. DeepSDF: Learning Continuous Signed Distance Functions for Shape\nRepresentation. In The IEEE Conference on Computer Vision and Pattern Recognition\n(CVPR).\nHanspeter Pfister, Matthias Zwicker, Jeroen Van Baar, and Markus Gross. 2000. Surfels:\nSurface elements as rendering primitives. In Proceedings of the 27th annual conference\non Computer graphics and interactive techniques. 335–342.\nShenhan Qian, Tobias Kirschstein, Liam Schoneveld, Davide Davoli, Simon Giebenhain,\nand Matthias Nießner. 2023. GaussianAvatars: Photorealistic Head Avatars with\nRigged 3D Gaussians. arXiv preprint arXiv:2312.02069 (2023).\nChristian Reiser, Songyou Peng, Yiyi Liao, and Andreas Geiger. 2021. KiloNeRF: Speed-\ning up Neural Radiance Fields with Thousands of Tiny MLPs. In International\nConference on Computer Vision (ICCV).\nChristian Reiser, Rick Szeliski, Dor Verbin, Pratul Srinivasan, Ben Mildenhall, Andreas\nGeiger, Jon Barron, and Peter Hedman. 2023. Merf: Memory-efficient radiance fields\nfor real-time view synthesis in unbounded scenes. ACM Transactions on Graphics\n(TOG) 42, 4 (2023), 1–12.\nDarius Rückert, Linus Franke, and Marc Stamminger. 2022. Adop: Approximate dif-\nferentiable one-pixel point rendering. ACM Transactions on Graphics (ToG) 41, 4\n(2022), 1–14.\nJohannes Lutz Schönberger and Jan-Michael Frahm. 2016. Structure-from-Motion\nRevisited. In Conference on Computer Vision and Pattern Recognition (CVPR).\nJohannes Lutz Schönberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm.\n2016. Pixelwise View Selection for Unstructured Multi-View Stereo. In European\nConference on Computer Vision (ECCV).\nThomas Schöps, Torsten Sattler, and Marc Pollefeys. 2019. Surfelmeshing: Online\nsurfel-based mesh reconstruction. IEEE transactions on pattern analysis and machine\nintelligence 42, 10 (2019), 2494–2507.\nYahao Shi, Yanmin Wu, Chenming Wu, Xing Liu, Chen Zhao, Haocheng Feng, Jingtuo\nLiu, Liangjun Zhang, Jian Zhang, Bin Zhou, Errui Ding, and Jingdong Wang. 2023.\nGIR: 3D Gaussian Inverse Rendering for Relightable Scene Factorization. Arxiv\n(2023). arXiv:2312.05133\nChristian Sigg, Tim Weyrich, Mario Botsch, and Markus H Gross. 2006. GPU-based\nray-casting of quadratic surfaces.. In PBG@ SIGGRAPH. 59–65.\nCheng Sun, Min Sun, and Hwann-Tzong Chen. 2022a. Direct Voxel Grid Optimization:\nSuper-fast Convergence for Radiance Fields Reconstruction. In CVPR.\nCheng Sun, Min Sun, and Hwann-Tzong Chen. 2022b. Improved Direct Voxel Grid\nOptimization for Radiance Fields Reconstruction. arxiv cs.GR 2206.05085 (2022).\nJohn Vince. 2008. Geometric algebra for computer graphics. Springer Science & Business\nMedia.\n\n\n2D Gaussian Splatting for Geometrically Accurate Radiance Fields\n•\n9\nPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping\nWang. 2021. NeuS: Learning Neural Implicit Surfaces by Volume Rendering for\nMulti-view Reconstruction. Advances in Neural Information Processing Systems 34\n(2021), 27171–27183.\nYiming Wang, Qin Han, Marc Habermann, Kostas Daniilidis, Christian Theobalt, and\nLingjie Liu. 2023. NeuS2: Fast Learning of Neural Implicit Surfaces for Multi-view\nReconstruction. In Proceedings of the IEEE/CVF International Conference on Computer\nVision (ICCV).\nTim Weyrich, Simon Heinzle, Timo Aila, Daniel B Fasnacht, Stephan Oetiker, Mario\nBotsch, Cyril Flaig, Simon Mall, Kaspar Rohrer, Norbert Felber, et al. 2007. A\nhardware architecture for surface splatting. ACM Transactions on Graphics (TOG)\n26, 3 (2007), 90–es.\nThomas Whelan, Renato F Salas-Moreno, Ben Glocker, Andrew J Davison, and Stefan\nLeutenegger. 2016. ElasticFusion: Real-time dense SLAM and light source estimation.\nThe International Journal of Robotics Research 35, 14 (2016), 1697–1716.\nOlivia Wiles, Georgia Gkioxari, Richard Szeliski, and Justin Johnson. 2020. SynSin:\nEnd-to-end View Synthesis from a Single Image. In Proceedings of the IEEE/CVF\nConference on Computer Vision and Pattern Recognition (CVPR).\nTianyi Xie, Zeshun Zong, Yuxing Qiu, Xuan Li, Yutao Feng, Yin Yang, and Chenfanfu\nJiang. 2023. PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynam-\nics. arXiv preprint arXiv:2311.12198 (2023).\nYunzhi Yan, Haotong Lin, Chenxu Zhou, Weijie Wang, Haiyang Sun, Kun Zhan, Xi-\nanpeng Lang, Xiaowei Zhou, and Sida Peng. 2023. Street Gaussians for Modeling\nDynamic Urban Scenes. (2023).\nYao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan. 2018. MVSNet: Depth\nInference for Unstructured Multi-view Stereo. European Conference on Computer\nVision (ECCV) (2018).\nLior Yariv, Jiatao Gu, Yoni Kasten, and Yaron Lipman. 2021. Volume rendering of neural\nimplicit surfaces. Advances in Neural Information Processing Systems 34 (2021),\n4805–4815.\nLior Yariv, Peter Hedman, Christian Reiser, Dor Verbin, Pratul P. Srinivasan, Richard\nSzeliski, Jonathan T. Barron, and Ben Mildenhall. 2023. BakedSDF: Meshing Neural\nSDFs for Real-Time View Synthesis. arXiv (2023).\nLior Yariv, Yoni Kasten, Dror Moran, Meirav Galun, Matan Atzmon, Basri Ronen, and\nYaron Lipman. 2020. Multiview Neural Surface Reconstruction by Disentangling\nGeometry and Appearance. Advances in Neural Information Processing Systems 33\n(2020).\nWang Yifan, Felice Serena, Shihao Wu, Cengiz Öztireli, and Olga Sorkine-Hornung.\n2019. Differentiable surface splatting for point-based geometry processing. ACM\nTransactions on Graphics (TOG) 38, 6 (2019), 1–14.\nAlex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and Angjoo Kanazawa. 2021.\nPlenOctrees for Real-time Rendering of Neural Radiance Fields. In ICCV.\nZehao Yu, Anpei Chen, Bozidar Antic, Songyou Peng, Apratim Bhattacharyya, Michael\nNiemeyer, Siyu Tang, Torsten Sattler, and Andreas Geiger. 2022a. SDFStudio: A Uni-\nfied Framework for Surface Reconstruction. https://github.com/autonomousvision/\nsdfstudio\nZehao Yu, Anpei Chen, Binbin Huang, Torsten Sattler, and Andreas Geiger. 2024. Mip-\nSplatting: Alias-free 3D Gaussian Splatting. Conference on Computer Vision and\nPattern Recognition (CVPR) (2024).\nZehao Yu and Shenghua Gao. 2020. Fast-MVSNet: Sparse-to-Dense Multi-View Stereo\nWith Learned Propagation and Gauss-Newton Refinement. In Conference on Com-\nputer Vision and Pattern Recognition (CVPR).\nZehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sattler, and Andreas Geiger.\n2022b. MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface\nReconstruction. Advances in Neural Information Processing Systems (NeurIPS) (2022).\nKai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. 2020. NeRF++: Analyzing\nand Improving Neural Radiance Fields. arXiv:2010.07492 (2020).\nQian-Yi Zhou, Jaesik Park, and Vladlen Koltun. 2018. Open3D: A Modern Library for\n3D Data Processing. arXiv:1801.09847 (2018).\nWojciech Zielonka, Timur Bagautdinov, Shunsuke Saito, Michael Zollhöfer, Jus-\ntus Thies, and Javier Romero. 2023.\nDrivable 3D Gaussian Avatars.\n(2023).\narXiv:2311.08581 [cs.CV]\nMatthias Zwicker, Hanspeter Pfister, Jeroen Van Baar, and Markus Gross. 2001a. EWA\nvolume splatting. In Proceedings Visualization, 2001. VIS’01. IEEE, 29–538.\nMatthias Zwicker, Hanspeter Pfister, Jeroen Van Baar, and Markus Gross. 2001b. Surface\nsplatting. In Proceedings of the 28th annual conference on Computer graphics and\ninteractive techniques. 371–378.\nMatthias Zwicker, Jussi Rasanen, Mario Botsch, Carsten Dachsbacher, and Mark Pauly.\n2004. Perspective accurate splatting. In Proceedings-Graphics Interface. 247–254.\nA\nDETAILS OF DEPTH DISTORTION\nWhile Barron et al. [Barron et al. 2022b] calculates the distortion\nloss with samples on the ray, we operate Gaussian primitives, where\nthe intersected depth may not be ordered. To this end, we adopt\nColor\nDepth\nGround truth\nOurs \n3DGS\nFig. 7. Affine-approximation adopted in [Zwicker et al. 2001b][Kerbl et al.\n2023] causes perspective distortion and inaccurate depth.\nan L2 loss and transform the intersected depth 𝑧to NDC space to\ndown-weight distant Gaussian primitives, 𝑚= NDC(𝑧), with near\nand far plane empirically set to 0.2 and 1000. We implemented our\ndepth distortion loss based on [Sun et al. 2022b], also powered by\ntile-based rendering. Here we show that the nested algorithm can\nbe implemented in a single forward pass:\nL =\n𝑁−1\n∑︁\n𝑖=0\n𝑖−1\n∑︁\n𝑗=0\n𝜔𝑖𝜔𝑗(𝑚𝑖−𝑚𝑗)2\n=\n𝑁−1\n∑︁\n𝑖=0\n𝜔𝑖\n \n𝑚2\n𝑖\n𝑖−1\n∑︁\n𝑗=0\n𝜔𝑗+\n𝑖−1\n∑︁\n𝑗=0\n𝜔𝑗𝑚2\n𝑗−2𝑚𝑖\n𝑖−1\n∑︁\n𝑗=0\n𝜔𝑗𝑚𝑗\n!\n=\n𝑁−1\n∑︁\n𝑖=0\n𝜔𝑖\n\u0010\n𝑚2\n𝑖𝐴𝑖−1 + 𝐷2\n𝑖−1 −2𝑚𝑖𝐷𝑖−1\n\u0011\n,\n(17)\nwhere 𝐴𝑖= Í𝑖\n𝑗=0 𝜔𝑗, 𝐷𝑖= Í𝑖\n𝑗=0 𝜔𝑗𝑚𝑗and 𝐷2\n𝑖= Í𝑖\n𝑗=0 𝜔𝑗𝑚2\n𝑗.\nSpecifically, we let 𝑒𝑖=\n\u0010\n𝑚2\n𝑖𝐴𝑖−1 + 𝐷2\n𝑖−1 −2𝑚𝑖𝐷𝑖−1\n\u0011\nso that the\ndistortion loss can be “rendered” as L𝑖= Í𝑖\n𝑗=0 𝜔𝑗𝑒𝑗. Here, L𝑖mea-\nsures the depth distortion up to the 𝑖-th Gaussian. During marching\nGaussian front-to-back, we simultaneously accumulate 𝐴𝑖, 𝐷𝑖and\n𝐷2\n𝑖, preparing for the next distortion computation L𝑖+1. Similarly,\nthe gradient of the depth distortion can be back-propagated to the\nprimitives back-to-front. Different from implicit methods where\n𝑚are the pre-defined sampled depth and non-differentiable, we\nadditionally back-propagate the gradient through the intersection\n𝑚, encouraging the Gaussians to move tightly together directly.\nB\nDEPTH CALCULATIONS\nMean depth:There are two optional depth computations used for\nour meshing process. The mean (expected) depth is calculated by\nweighting the intersected depth:\n𝑧mean =\n∑︁\n𝑖\n𝜔𝑖𝑧𝑖/(\n∑︁\n𝑖\n𝜔𝑖+ 𝜖)\n(18)\nwhere 𝜔𝑖= 𝑇𝑖𝛼𝑖ˆ\nG𝑖(u(x) is the weight contribution of the 𝑖-th\nGaussian and 𝑇𝑖= Î𝑖−1\n𝑗=1(1 −𝛼𝑗ˆ\nG𝑗(u(x))) measures its visibility.\nIt is important to normalize the depth with the accumulated alpha\n𝐴= Í\n𝑖𝜔𝑖to ensure that a 2D Gaussian can be rendered as a planar\n2D disk in the depth visualization.\n\n\n10\n•\nBinbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao\n(a) Ground-truth\n(b) MipNeRF360 [Barron et al. 2022b], SSIM=0.813\n(c) 3DGS, normals from depth points\n(d) 3DGS [Kerbl et al. 2023], SSIM=0.834\n(e) Our model (2DGS), normals from depth points\n(f) Our model (2DGS), SSIM=0.845\nFig. 8. We visualize the depth maps generated by MipNeRF360 [Barron et al. 2022b], 3DGS [Kerbl et al. 2023], and our method. The depth maps for 3DGS (d)\nand 2DGS (f) are rendered using Eq. 18 and visualized following MipNeRF360. To highlight the surface smoothness, we further visualize the normal estimated\nfrom depth points using Eq. 15 for both 3DGS (c) and ours (e). While MipNeRF360 is capable of producing plausibly smooth depth maps, its sampling process\nmay result in the loss of detailed structures. Both 3DGS and 2DGS excel at modeling thin structures; however, as illustrated in (c) and (e), the depth map of\n3DGS exhibits significant noise. In contrast, our approach generates sampled depth points with normals consistent with the rendered normal map (refer to\nFigure 1b), thereby enhancing depth fusion during the meshing process.\nTable 6. PSNR↑, SSIM↑, LIPPS↓scores for MipNeRF360 dataset.\nbicycle\nflowers\ngarden\nstump\ntreehill\nroom\ncounter\nkitchen\nbonsai\nmean\n3DGS\n24.71\n21.09\n26.63\n26.45\n22.33\n31.50\n29.07\n31.13\n32.26\n27.24\nSuGaR\n23.12\n19.25\n25.43\n24.69\n21.33\n30.12\n27.57\n29.48\n30.59\n25.73\nOurs\n24.82\n20.99\n26.91\n26.41\n22.52\n30.86\n28.45\n30.62\n31.64\n27.03\n3DGS\n0.729\n0.571\n0.834\n0.762\n0.627\n0.922\n0.913\n0.926\n0.943\n0.803\nSuGaR\n0.639\n0.486\n0.776\n0.686\n0.566\n0.910\n0.892\n0.908\n0.932\n0.755\nOurs\n0.731\n0.573\n0.845\n0.764\n0.630\n0.918\n0.908\n0.927\n0.940\n0.804\n3DGS\n0.265\n0.377\n0.147\n0.266\n0.362\n0.231\n0.212\n0.138\n0.214\n0.246\nSuGaR\n0.344\n0.416\n0.220\n0.335\n0.429\n0.245\n0.232\n0.164\n0.221\n0.290\nOurs\n0.271\n0.378\n0.138\n0.263\n0.369\n0.214\n0.197\n0.125\n0.194\n0.239\nTable 7. PSNR scores for Synthetic NeRF dataset [Mildenhall et al. 2021].\nMic\nChair\nShip\nMaterials\nLego\nDrums\nFicus\nHotdog\nAvg.\nPlenoxels\n33.26\n33.98\n29.62\n29.14\n34.10\n25.35\n31.83\n36.81\n31.76\nINGP-Base\n36.22\n35.00\n31.10\n29.78\n36.39\n26.02\n33.51\n37.40\n33.18\nMip-NeRF\n36.51\n35.14\n30.41\n30.71\n35.70\n25.48\n33.29\n37.48\n33.09\nPoint-NeRF\n35.95\n35.40\n30.97\n29.61\n35.04\n26.06\n36.13\n37.30\n33.30\n3DGS\n35.36\n35.83\n30.80\n30.00\n35.78\n26.15\n34.87\n37.72\n33.32\n2DGS (Ours)\n35.09\n35.05\n30.60\n29.74\n35.10\n26.05\n35.57\n37.36\n33.07\nMedian depth:We compute the median depth as the largest “visible”\ndepth, considering 𝑇𝑖= 0.5 as the pivot for surface and free space:\n𝑧median = max{𝑧𝑖|𝑇𝑖> 0.5}.\n(19)\nWe find our median depth computation is more robust to [Luiten\net al. 2024]. When a ray’s accumulated alpha does not reach 0.5,\nwhile Luiten et al. sets a default value of 15, our computation selects\nthe last Gaussian, which is more accurate and suitable for training.\nC\nADDITIONAL RESULTS\nOur 2D Gaussian Splatting method achieves comparable perfor-\nmance even without the need for regularizations, as Table 7 shows.\nAdditionally, we have provided a breakdown of various per-scene\nmetrics for the MipNeRF360 dataset [Barron et al. 2022b] in Tables 6.\nFigure 8 presents a comparison of our rendered depth maps with\nthose from 3DGS and MipNeRF360.\n\n\n2D Gaussian Splatting for Geometrically Accurate Radiance Fields\n•\n11\nscan24\nscan37\nscan40\nscan55\nscan63\nscan65\nscan69\nscan83\nscan97\nscan105\nscan106\nscan110\nscan114\nscan118\nscan122\nscan24\nscan37\nscan40\nscan55\nscan63\nscan65\nscan69\nscan83\nscan97\nscan105\nscan106\nscan110\nscan114\nscan118\nscan122\n2DGS\n3DGS\nFig. 9. Comparison of surface reconstruction using our 2DGS and 3DGS [Kerbl et al. 2023]. Meshes are extracted by applying TSDF to the depth maps.\n\n\n12\n•\nBinbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao\nBarn\nTruck\nCaterpillar\nMeetingRoom\nIgnatius\nFig. 10. Qualitative studies for the Tanks and Temples dataset [Knapitsch et al. 2017].\nFig. 11. Appearance rendering results from reconstructed 2D Gaussian disks, including DTU, TnT, and Mip-NeRF360 datasets.\n(A) Semi-transparent\n(B) High light\nFig. 12. Illustration of limitations: Our 2DGS struggles with the accurate reconstruction of semi-transparent surfaces, for example, the glass shown in example\n(A). Moreover, our method tends to create holes in areas with high light intensity, as shown in (B).","difficulty":"hard","domain":"Multi-Document QA","length":"short","question":"Both 3D Gaussian Splatting (3DGS) and 2D Gaussian Splatting (2DGS) seek to project Gaussians onto screen space for novel view synthesis, but their methods of handling surface projection and depth consistency differ. The 3D Gaussian splat is governed by:\n\n\nG(p) = \\exp\\left(-\\frac{1}{2}(p - p_k)^\\top \\Sigma^{-1} (p - p_k)\\right)\n\n\nwhere  \\Sigma  represents the covariance matrix that controls the orientation and scaling of the Gaussian in 3D space. In contrast, 2D Gaussian Splatting collapses the volume into 2D elliptical disks embedded in 3D space to improve multi-view consistency.\n\nGiven the key differences between 3DGS and 2DGS, which of the following best explains how 2D Gaussian Splatting achieves superior surface projection accuracy, and what is the main mathematical reasoning behind this improvement?","sub_domain":"Academic"}

Source: https://huggingface.co/datasets/zai-org/LongBench-v2

initial import

Posting: /agents

GET /api/v1/write?intent=publish&task_id=b0a2722f-c7f1-5506-be8c-365a7b70e03a&body={url_encoded_text}&agent_name={optional_name}&nonce={optional_random_id}
