# LongBench v2 / 66f14c70821e116aacb271ee

task_id: 6acb208a-1e41-5ea1-ac7f-be286674edf5
task_key: train--66f14c70821e116aacb271ee
task_revision_id: 3

{"choice_A":"The data set in DAIR-V2X includes actual measured data of V2V and V2I, while the data set in V2X-Sim also includes V2V and V2I, but it is simulated data.","choice_B":"The dataset in DAIR-V2X is measured data and takes into account the time asynchrony caused by communication, while the dataset in V2X-Sim does not take this into account.","choice_C":"Neither the DAIR-V2X nor the V2X-Sim datasets consider the problem of posture errors.","choice_D":"DAIR-V2X is the first measured dataset that includes both V2V and V2I。","context":"DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Cooperative\n3D Object Detection\nHaibao Yu1, Yizhen Luo1,3, Mao Shu2, Yiyi Huo1,4, Zebang Yang1,3, Yifeng Shi2, Zhenglong Guo2,\nHanyu Li2, Xing Hu2, Jirui Yuan1, Zaiqing Nie1*\n1Institute for AI Industry Research(AIR), Tsinghua University\n2 Baidu Inc. 3 Department of Computer Science and Technology, Tsinghua University\n4 University of Chinese Academy of Science\n{yuhaibao@air.,luoyz18@mails.,yzb19@mails.,yuanjirui@air.,zaiqing@air.}tsinghua.edu.cn,\n{shumao,shiyifeng,guozhenglong,lihanyu02,huxing}@baidu.com, huoyiyi18@mails.ucas.ac.cn\nAbstract\nAutonomous driving faces great safety challenges for a\nlack of global perspective and the limitation of long-range\nperception capabilities.\nIt has been widely agreed that\nvehicle-infrastructure cooperation is required to achieve\nLevel 5 autonomy. However, there is still NO dataset from\nreal scenarios available for computer vision researchers\nto work on vehicle-infrastructure cooperation-related prob-\nlems.\nTo accelerate computer vision research and inno-\nvation for Vehicle-Infrastructure Cooperative Autonomous\nDriving (VICAD), we release DAIR-V2X Dataset, which is\nthe first large-scale, multi-modality, multi-view dataset from\nreal scenarios for VICAD. DAIR-V2X comprises 71254 Li-\nDAR frames and 71254 Camera frames, and all frames\nare captured from real scenes with 3D annotations. The\nVehicle-Infrastructure Cooperative 3D Object Detection\nproblem (VIC3D) is introduced, formulating the problem of\ncollaboratively locating and identifying 3D objects using\nsensory inputs from both vehicle and infrastructure. In ad-\ndition to solving traditional 3D object detection problems,\nthe solution of VIC3D needs to consider the temporal asyn-\nchrony problem between vehicle and infrastructure sensors\nand the data transmission cost between them.\nFurther-\nmore, we propose Time Compensation Late Fusion (TCLF),\na late fusion framework for the VIC3D task as a bench-\nmark based on DAIR-V2X. Find data, code, and more up-\nto-date information at https://thudair.baai.ac.cn/index and\nhttps://github.com/AIR-THU/DAIR-V2X.\n1. Introduction\nAutonomous driving (AD) is arguably one of the hottest\ntopics currently occupying public attention and imagina-\n*Corresponding author. 3,4 Work done while at AIR.\ntion.\nThe success of deep neural networks brings the\npromise of solving AD’s core requirement to perceive the\nsurrounding environment from point cloud [15,20,28], im-\nages [7, 18] or multi-modality data [21, 24].\nDespite its\nFigure 1.\nDatasets available for 3D Object Detection in au-\ntonomous driving. DAIR-V2X is the first real-world V2X dataset\nfor VICAD.\ngreat progress recently, autonomous driving still faces great\nsafety challenges for a lack of global perspective and the\nlimitation of long-range perception capability. It has been\nwidely agreed that vehicle-infrastructure cooperation is re-\nquired to achieve Level 5 autonomy. Utilizing both vehi-\ncle and infrastructure sensors brings a number of signifi-\ncant advantages, including providing a global perspective\nfar beyond the current horizon and covering blind spots.\nAdvances in communications like V2X (vehicle to every-\nthing) have made it possible to utilize data from infrastruc-\nture sensors [3,22]. However, there is still NO dataset from\nreal scenarios available for researchers to work on vehicle-\ninfrastructure cooperation-related problems.\nTo accelerate computer vision research and innovation\n21361\n\n\nTable 1. A detailed comparison between autonomous driving-related datasets. - indicates that specific information is not provided. In\nparticular, DAIR-V2X is composed of DAIR-V2X-C, DAIR-V2X-V and DAIR-V2X-I, where DAIR-V2X-C is captured by both vehicle\nand infrastructure sensors, DAIR-V2X-V is captured by vehicle sensors, and DAIR-V2X-I is captured by infrastructure sensors.\nDataset\nYear\nReal/Simulated\nView\nImage\nPointcloud\n3D boxes\nClasses\nKITTI [10]\n2012\nreal\nsingle vehicle\n15k\n15k\n200k\n8\nnuScenes [2]\n2019\nreal\nsingle vehicle\n1.4M\n400k\n1.4M\n23\nWaymo Open [23]\n2019\nreal\nsingle vehicle\n1M\n200k\n12M\n4\nApolloScape [12]\n2018\nreal\nsingle vehicle\n144k\n0\n70k\n8-35\nBBD100K [30]\n2020\nreal\nsingle vehicle\n100M\n0\n0\n10\nONCE [17]\n2021\nreal\nsingle vehicle\n7M\n1M\n417k\n5\nSYNTHIA [19]\n2016\nsimulated\nsingle vehicle\n213k\n0\n-\n13\nV2X-Sim [16]\n2021\nsimulated\nmulti-vehicle\n0\n10k\n26.6k\n2\nhighD [13]\n2018\nreal\ninfrastructure (UAV)\n1.53M\n0\n0\n1\nDAIR-V2X (Our)\n2021\nreal\nvehicle-infrastructure cooperative\n71k\n71k\n1.2M\n10\n- DAIR-V2X-C\n2021\nreal\nvehicle-infrastructure cooperative\n39k\n39k\n464k\n10\n- DAIR-V2X-V\n2021\nreal\nsingle vehicle\n22k\n22k\n239k\n10\n- DAIR-V2X-I\n2021\nreal\ninfrastructure\n10k\n10k\n493k\n10\nfor Vehicle-Infrastructure Cooperative Autonomous Driv-\ning (VICAD), we release DAIR-V2X Dataset, which is the\nfirst large-scale, multi-modality, multi-view dataset for VI-\nCAD. It contains 71254 LiDAR frames and 71254 Cam-\nera frames captured in intersection scenes where a well-\nequipped vehicle passes through intersections with infras-\ntructure sensors deployed. 40% of the frames are captured\nfrom infrastructure sensors and 60% of the frames are cap-\ntured from vehicle sensors. All of them are precisely la-\nbeled by expert annotators. The dataset covers 10 km of\ncity roads, 10 km of highway, 28 intersections, and 38 km2\nof driving regions with diverse weather and lighting varia-\ntions. More details could be found in Tab. 1.\nIn this paper, the Vehicle-Infrastructure Cooperative 3D\nObject Detection (VIC3D) task is introduced, formulating\nthe problem of cooperatively locating and identifying 3D\nobjects using sensory inputs from both vehicle and infras-\ntructure. In addition to solving traditional 3D object detec-\ntion problems, the solution of VIC3D needs to consider the\ntemporal asynchrony problem and data transmission cost\nbetween vehicle and infrastructure sensors.\nTo resolve the VIC3D object detection task and facilitate\nfuture research, we also introduce our VIC3D object detec-\ntion benchmark in this paper. For data with less temporal\nasynchrony problems, we implement both early fusion and\nlate fusion approaches. Results show that the average preci-\nsion of fusion methods is 10 to 20 points higher than detec-\ntors that only use information from a single view. Results\nalso show that early fusion can achieve better performance\nthan late fusion but requires more data transmission. With\nthe DAIR-V2X dataset, we expect more future research to\nachieve a performance-bandwidth trade-off. For data with\nsevere temporal asynchrony, we propose a Time Compensa-\ntion Late Fusion framework, which can effectively alleviate\nthe temporal asynchrony problem.\nThe key contributions of our work are as follows:\n• We release the DAIR-V2X dataset, which is the first\nlarge-scale dataset for vehicle-infrastructure coopera-\ntive autonomous driving. All frames are captured from\nreal scenarios with 3D annotations.\n• We formulate the problem of cooperatively locating\nand identifying 3D objects using sensory inputs from\nboth vehicle and infrastructure as VIC3D.\n• We introduce benchmarks for VIC3D object detection\nand single-view 3D object detection tasks. The results\nshow the effectiveness of vehicle-infrastructure coop-\neration in VIC3D object detection. Especially, we pro-\npose the Time Compensation Late Fusion framework\nto alleviate the temporal asynchrony problem.\n2. Relative Work\n2.1. Autonomous Driving Datasets\nIn recent years, an increasing number of autonomous\ndriving datasets have been released and greatly promoted\nthe development of autonomous driving research. Datasets\nlike SYNTHIA [19] and Cityscapes [5] mainly focus on\n2D annotations for images. KITTI [10] and nuScenes [2]\nare multi-modality datasets providing camera images as\nwell as LiDAR point clouds. Nevertheless, all datasets men-\ntioned above only provide data from a single-vehicle view.\nV2X-SIM [16] is an attempt to generate a multi-vehicle\nview dataset, but the dataset was generated by a simula-\ntor rather than captured from real scenarios.\nCompared\nwith those datasets, our DAIR-V2X dataset is the first large-\nscale, multi-modality, multi-view dataset captured from real\nscenarios for VICAD, and contains data captured from the\nVehicle-Infrastructure Cooperative view. Tab. 1 shows the\ncomparison of our dataset with the others. In our DAIR-\n21362\n\n\nFigure 2. a) Acquisition system with infrastructure sensors. b) Acquisition system with vehicle sensors. c) Infrastructure-view image and\npoint cloud with 3d annotation. Paired vehicle-view and infrastructure-view information complement each other in the perspective of view.\nd) Vehicle-view image and point cloud with 3d annotation.\nV2X, we also provide a Repo3D [29] dataset composed of\nmulti-source infrastructure images and 3D annotations, for\nthose who are interested in Mono3D object detection and\ndomain adaptation.\n2.2. 3D Detection\n3D object detection serves as the prerequisite for the suc-\ncess of autonomous driving. Many techniques have been\nintroduced and can be roughly classified into three cate-\ngories.\na) Image-based 3D Detection refers to methods\nthat detect 3D objects directly from 2D images. ImVox-\nelNet [7] is a good example to make predictions from im-\nages. b) Pointcloud-based 3D Detection stands for manners\nthat make 3D object detection merely from point clouds.\nPointPillars [15], SECOND [27], and 3DSSD [28] are such\napproaches that achieve convincing detection results from\npoint clouds. c) Multimodality-based 3D Detection uses\nboth images and point clouds to make predictions. Point-\npainting [24] and MVXNet [21] are practices of fusing im-\nage and LiDAR features to predict 3D bounding boxes.\nWhile 3D object detection has made great progress recently,\nthere are still some tough problems that remain to solve\nsuch as blind spots and weak long-distance perception. To\nexplore how to utilize the infrastructure information to solve\nthe problems mentioned above, we conduct VIC3D object\ndetection based on our dataset proposed in this paper.\n2.3. Multi-Sensor Fusion\nMulti-sensor fusion [26] is the integration of heteroge-\nneous information collected by different sensors to alleviate\nthe uncertainty and vulnerability of systems that rely on a\nsingle sensor. Based on the fusion stage, multi-sensor fu-\nsion can be categorized into early fusion, intermediate fu-\nsion, and late fusion. a) In early fusion, raw data from dif-\nferent sensors are directly transferred and fused [9]. b) In in-\ntermediate fusion, intermediate representations like features\nextracted from the models are fused [4,21]. c) In late fusion,\nthe prediction outputs like 3D information of the objects are\nfused [11]. VIC3D can be considered as a variant of the\nmulti-sensor problem, so previous fusion methods can be\ntaken into consideration to integrate the infrastructure in-\nformation. However, in addition to the multi-sensor fusion\nchallenges, VIC3D faces difficulties caused by the temporal\nasynchrony problem and the data transmission constraint.\n2.4. V2X Cooperative Perception\nV2X aims to build a communication system between ve-\nhicles and other devices in a complex traffic environment.\nCurrent V2X research mainly focuses on V2V (Vehicle-to-\nVehicle) and V2I (Vehicle-to-Infrastructure) area. V2VNet\n[25] is a pioneering work in V2V that broadcasts com-\npressed intermediate features and propagates message re-\nceived from nearby vehicles to generate motion forecasts.\nWorks of V2I [6,31] leverage infrastructure LiDAR data to\ngenerate and broadcast detection results. However, none of\nthese approaches have been verified on a dataset captured\nfrom real scenarios. This may cause a huge gap between\ntheory and practice. Therefore, we release the DAIR-V2X\ndataset to boost further study in this field.\n3. The DAIR-V2X Dataset\nIn order to facilitate research on VICAD, we re-\nlease DAIR-V2X, a large-scale, multi-modality, multi-view\ndataset from real scenarios with 3D annotations for vehicle\ninfrastructure cooperation. Here we describe how we set up\ninfrastructure and vehicle sensors, select interesting scenes,\nannotate the dataset and protect the privacy of third parties.\n3.1. Setup\nEquipment. Equipment for data collection are composed\nof infrastructure sensors and vehicle sensors. a) Infrastruc-\nture sensors. Each of the 28 intersections selected from\nBeijing High-level Autonomous Driving Demonstration\n21363\n\n\nTable 2. Key Sensor Specifications in DAIR-V2X. Veh. stands for\nvehicle view, and Inf. stands for infrastructure view.\nSensor\nDetails\nInf. LiDAR\n300 beams, 10Hz capture frequency, 100o\nhorizontal FOV, −30o to 10o vertical FOV,\n≤280m range, ±3cm accuracy\nInf. Camera\nRGB, 25Hz capture frequency, 1920x1080\nresolution, JPEG compressed\nVeh. LiDAR\n40 beams, 10Hz capture frequency, 360o\nhorizontal FOV, −30o to 10o vertical FOV,\n≤200m range, ±0.33o vertical resolution\nVeh. Camera\nRGB, 20Hz capture frequency, 1920x1080\nresolution, JPEG compressed\nVeh. GPS & IMU\n1000HZ update rate\nArea are deployed with four pairs of 300-beam LiDAR and\nhigh-resolution camera. The DAIR-V2X dataset picks only\none pair of them. b) Vehicle sensors. One 40-beam LiDAR\nand one high-quality camera looking forward are mounted\non top of the autonomous vehicles.\nSpecific layout is\nposted in Figure 2, and precise details are displayed in\nTable 2.\nCoordinate. There are 5 types of coordinate systems on\nDAIR-V2X, i.e., the LiDAR coordinate, the camera coor-\ndinate, the image coordinate, the world coordinate, and the\npositioning coordinate. The origin of the LiDAR coordinate\nsystem is located at the center of the LiDAR sensor, the x-\naxis is positive forwards, the y-axis is positive to the left,\nand the z-axis is positive upwards. The infrastructure Li-\nDAR coordinate system is converted from its original sys-\ntem which has an inclination angle with the ground. The\nreal-time relative pose of the equipped vehicle is obtained\nfrom GPS/IMU combined with SLAM and a local map.\nThere is also manual secondary labeling confirmation to en-\nsure calibration accuracy. The Lidar-to-Camera transforma-\ntion is obtained by multiplying Lidar-to-World and World-\nto-Camera transformations.\n3.2. Data Acquisition\nCollection. We drive a well-equipped vehicle in the col-\nlection area and save the corresponding vehicle frames and\ninfrastructure frames respectively. After the collection of\nraw data, we manually select 100 representative scenes of\n20s duration. Such scenes include vehicle data and infras-\ntructure data, where vehicles drive through intersections de-\nployed with equipment. We sample key frames at 10Hz\nfrom both sides to form DAIR-V2X-C. In DAIR-V2X-C, it\nis important to note that the timestamp difference between\na vehicle frame and its closest infrastructure frame could\nbe slightly varied, due to the asynchronous triggering be-\ntween vehicle sensors and infrastructure sensors. We sam-\nple 22K frames from additionally about 350 vehicle-only\nsegments of 60s duration to form DAIR-V2X-V, and sam-\nple 10K frames from additionally about 150 infrastructure-\nonly segments duration to form DAIR-V2X-I, to enlarge the\ndataset. Compared to the single-view data in DAIR-V2X-\nC, DAIR-V2X-V and DAIR-V2X-I contains more diverse\nscenes and will be more challenging to only improve the\nsingle-view performance.\nAnnotation.\nWith multiple validation steps and refine-\nment processes, expert annotators make high-quality an-\nnotations for infrastructure frames and vehicle frames\nrespectively.\nSpecifically, annotators exhaustively label\neach of the 10 object classes in every image and point\ncloud frame with its category attribute, occlusion state,\ntruncated state, and a 7-dimensional cuboid modeled as\nx, y, z, width, length, height, and yaw angle.\n10 cate-\ngories include different vehicles, pedestrian, different cy-\nclists. Moreover, experts also meticulously annotate objects\nin camera images with a rectangle bounding box modeled\nas x, y, width, and length.\nTo be mentioned, we also conduct semi-automatic label-\ning for the cooperative annotations with vehicle and infras-\ntructure frame pairs. We first select vehicle and infrastruc-\nture frame pairs from DAIR-V2X-C. The timestamp differ-\nences between the two frames of the selected pairs are less\nthan 10ms (We call it the Synchronous Case which is de-\nfined in Section 4.1. To obtain more cooperative annota-\ntions, we extend the threshold from 10ms to 30ms). Next,\nwe convert infrastructure 3D boxes into vehicle LiDAR co-\nordinate system and fuse the vehicle annotations and infras-\ntructure annotations. For each 3D box in the infrastructure\nannotation, if we can not find any 3D box in the vehicle an-\nnotation that has the same location and category, we add the\ninfrastructure 3D box into the vehicle annotations; in this\nway, we get the vehicle-infrastructure cooperative annota-\ntions. We manually supervise and adjust the cooperative\nannotations to generate more accurate annotations. Here we\ntake 9331 infrastructure frames and vehicle frames as well\nas the cooperative annotations to form the VIC-Sync dataset\nfor our VIC3D object detection benchmark.\nProtection. The whole dataset is desensitized before public\nrelease. Complied with local laws and regulations, we erase\nall localization information, including road name, map data,\nand positioning information, to make sure our dataset meets\nrequirements. In addition, we utilize professional labeling\ntools to blur all the information suspected of privacy viola-\ntion, including road signs, license plates and faces, to pro-\ntect privacy and avoid violating personal rights.\n4. Task & Metrics\nAutonomous driving faces great safety challenges for a\nlack of global perspective and the limitation of long-range\nperception capabilities. Since 3D object detection is one\nof the key perception tasks in autonomous driving, in this\npaper, we focus on the vehicle-infrastructure cooperative\n(VIC) 3D object detection task, the vehicle receives and\n21364\n\n\nintegrates information from infrastructure to localize and\nrecognize objects surrounding itself. Compared with tradi-\ntional multi-sensor 3D object detection tasks, VIC3D object\ndetection has the following alternative characteristics:\n• Transmission Cost. Limited by physical communica-\ntion conditions, fewer data should be transmitted from\ninfrastructure to reduce bandwidth consumption, alle-\nviate time delay, and satisfy real-time requirements.\nThus, the solution to VIC3D object detection needs to\nbalance the trade-off between the performance and the\ntransmission cost.\n• Temporal Asynchrony. Timestamps of data from the\nvehicle sensors and the infrastructure sensors are dif-\nferent due to the asynchronous triggering and time\ndelay caused by transmission cost, to generate the\ntemporal-spatial error. Therefore, temporal synchro-\nnization should be considered in solving VIC3D.\nTo better formulate the VIC3D object detection task, we\nwill give a detailed definition to the VIC3D object detection\nand then provide two metrics to measure detection perfor-\nmance and transmission cost in this section.\n4.1. VIC3D Object Detection\nVIC3D object detection can be formulated as the opti-\nmization problem of effectively integrating infrastructure\nand vehicle information to localize and recognize 3D\nobjects considering transmission cost.\nHere we discuss\nwhat the input and output of VIC3D should be.\nInput. The input of VIC3D is composed of data from the\nvehicle and the infrastructure.\n• Vehicle Frame Iv(tv): captured at time tv as well as its\nrelative pose Mv(tv), where Iv(·) denotes the captur-\ning function of vehicle sensors.\n• Infrastructure Frame Ii(ti): captured at time ti as well\nas its relative pose Mv(ti), where Ii(·) denotes the\ncapturing function of infrastructure sensors.\nNote that ti should be earlier than tv because there is a time\ndelay caused by data transmission from the infrastructure\nto the vehicle. Considering that the objects would move\nso slightly in the tiny time interval that the spatial offset\ncan be ignored, we take the case that |tv −ti| ≤10ms\nas Synchronous Case (i.e. tv ≈ti). Similarly, we take\nthe case that |tv −ti| > 10ms as Asynchronous Case.\nIn addition, we allow using more infrastructure frames\nprevious to Ii(ti) in solving VIC3D to make full use of the\ninfrastructure computing resources.\nGround Truth. The outputs of VIC3D object detection\ncontain 3D information like the location, category, and ori-\nentation of objects surrounding the vehicle.\nThe corre-\nsponding ground truth of VIC3D is the fusion result of in-\nfrastructure and vehicle ground truth, which could be for-\nmulated as:\n  G T =  GT_{v} \\cup GT_{i}, \n(1)\nwhere GTv is the ground truth for vehicle sensor percep-\ntion and GTi is the ground truth for infrastructure sensor\nperception.\nVIC3D is mainly used to improve the perception perfor-\nmance of the self-driving vehicle. We are more concerned\nabout a certain range of egocentric surroundings and the 3D\ninformation of objects at time tv than at ti. Therefore, GTv\nand GTi should both be based on time tv. However, the\ntimestamp of the input frames captured from the infrastruc-\nture and captured from the vehicle could be different that\ntv ̸= ti. This not only brings challenges to fusing the in-\nfrastructure information in model prediction but also creates\nhuge problems to generate the ground truth. That’s because\nobjects annotated with infrastructure frame at time ti may\nmove to different locations at time tv, and we cannot di-\nrectly get the infrastructure frame at time tv to annotate.\nIn response to these difficulties, we discuss how we gen-\nerate the ground truth for VIC3D based on DAIR-V2X.\n• Synchronous Case (i.e. tv ≈ti). Under this condition,\nan object that appears in vehicle frame Iv(tv) should\nhave the same spatial location as it appears in infras-\ntructure frame Ii(ti). Therefore, we can directly take\nthe vehicle-infrastructure cooperative 3D annotations\nobtained by semi-automatic labeling illustrated in Sec-\ntion 3.2 as ground truth.\n• Asynchronous Case (i.e. tv ̸= ti). If we can find such\ninfrastructure frame Ii(t′\ni) satisfying |tv −t\n′\ni| ≤10ms,\nwe can generate ground truth with Ii(t\n′\ni). If not, we\nhave to estimate the 3D states of objects at tv to gener-\nate ground truth. This work can be carried out based on\nthe tracking ID and kinematic equation after we pro-\nvide the tracking ID in future work.\n4.2. Evaluation Metrics.\nVIC3D object detection has two major goals: better\ndetection performance and less transmission cost.\nWe\ndescribe the metrics for such two goals below.\nAverage Precision. AP (Average precision) is a popular\nmetric for measuring the object detectors performance [8].\nWe also use AP to evaluate the 3d detection performance\nwith cooperative annotations as ground truth. Since we are\nmore concerned about egocentric surroundings, we remove\nobjects outside the designed area. Here we set the designed\narea as a rectangular area as [0, -39.12, 100, 39.12].\nTransmission Cost. We use AB (Average Byte) to measure\nthe transmission cost. Here Byte is a unit of digital informa-\ntion that consists of eight bits. To simplify the problem, we\nignore the time consumption of data encoders and decoders\nduring transmission. That means the less transmission cost,\nthe less time delay. Data to be transmitted from the infras-\ntructure can be one or a combination of the following forms.\n21365\n\n\nTable 3. VIC3D object detection Benchmark on DAIR-V2X-C.\nModality\nFusion\nModel\nDataset\nAP3D(IoU=0.5)\nAPBEV (IoU=0.5)\nAB\nOverall\n0-30m\n30-50m\n50-100m\nOverall\n0-30m\n30-50m\n50-100m\n(Byte)\nImage\nVeh.-Only\nImvoxelNet [7]\nVIC-Sync\n12.03\n16.25\n7.25\n2.28\n13.62\n17.66\n8.58\n2.82\n0\nInf.-Only\nImvoxelNet [7]\nVIC-Sync\n19.93\n27.34\n17.61\n14.43\n25.31\n32.02\n23.28\n20.38\n102.32\nLate Fusion\nImvoxelNet [7]\nVIC-Sync\n26.56\n34.20\n17.20\n9.81\n31.40\n37.75\n21.21\n12.99\n102.32\nPointcloud\nVeh.-Only\nPointPillars [15]\nVIC-Sync\n31.33\n27.48\n25.58\n12.63\n35.06\n30.55\n28.65\n14.16\n0\nInf.-Only\nPointPillars [15]\nVIC-Sync\n17.62\n16.54\n10.98\n9.17\n24.40\n21.47\n16.00\n13.07\n336.16\nLate Fusion\nPointPillars [15]\nVIC-Sync\n41.90\n37.65\n32.72\n18.84\n47.96\n42.40\n37.65\n22.08\n336.16\nEarly Fusion\nPointPillars [15]\nVIC-Sync\n50.03\n53.07\n60.38\n33.05\n53.73\n55.80\n64.08\n36.17\n1382275.75\nPointcloud\nLate Fusion\nPointPillars [15]\nVIC-Async-1\n40.21\n34.17\n29.40\n15.50\n46.41\n38.05\n34.10\n19.20\n341.08\nLate Fusion\nPointPillars [15]\nVIC-Async-2\n35.29\n32.16\n28.07\n13.44\n40.65\n35.62\n32.35\n15.88\n306.79\nEarly Fusion\nPointPillars [15]\nVIC-Async-1\n47.47\n48.88\n58.86\n30.89\n51.67\n52.70\n63.09\n34.72\n1362216.0\nPointcloud\nTCLF\nPointPillars [15]\nVIC-Async-1\n40.79\n34.67\n29.69\n15.76\n46.80\n38.24\n34.27\n19.40\n539.60\nTCLF\nPointPillars [15]\nVIC-Async-2\n36.72\n33.91\n29.41\n14.52\n41.67\n36.78\n33.36\n17.18\n506.70\n• Raw data such as images or point clouds contains com-\nplete information but requires much transmission cost.\n• Intermediate representation requires less transmission\ncost while retaining valuable information, which may\nachieve a better performance-transmission trade-off.\nSurely, this requires a more sophisticated design to ex-\ntract suitable intermediate representation.\n• Object-level outputs directly provide 3D object infor-\nmation. Although it is transmission-efficient, it may\nlose valuable information.\n• Other auxiliary information like scene flows help to al-\nleviate temporal asynchrony problems.\n5. Benchmark\nIn this section, we provide a VIC3D object detection\nbenchmark and a Single-View (SV) 3D object detection\nbenchmark on our DAIR-V2X dataset, analyze their char-\nacteristics and suggest avenues for future research.\n5.1. Benchmark for VIC3D object detection\nWe provide a benchmark for VIC3D object detection on\nthe VIC-Sync dataset extracted from DAIR-V2X-C, which\nis illustrated in Section 3.2. The dataset is composed of\n9311 pairs of infrastructure and vehicle frames as well as\ntheir cooperative annotations as ground truth. Besides, we\ntake the temporal asynchrony between the infrastructure\nframe and the vehicle frame into consideration in the bench-\nmark, which is mainly caused by the difference in the sam-\npling rate and transmission delay. To simulate the tempo-\nral asynchrony phenomenon, we replace each infrastructure\nframe in the VIC-Sync dataset with the infrastructure frame\nwhich is k-th frame previous to the original infrastructure\nframe to construct the VIC-Async-k dataset for the bench-\nmark. In our experiments, we set k = 1, 2. We split VIC-\nSync and VIC-Async-k datasets to train/valid/test part as\n5:2:3 respectively. We use cooperative annotations to eval-\nuate the detection results under the vehicle-egocentric view.\nThe experiment results are presented in Table 3.\n5.1.1\nBaselines\nHere we present several baselines with different modalities\nand fusion methods for VIC3D object detection.\nLiDAR detection baseline with Late Fusion.\nTo demon-\nstrate the performance improvement by utilizing both in-\nfrastructure and vehicle data, we implement a late fusion\nframework with an infrastructure detector and a vehicle de-\ntector. Firstly, we choose PointPillars [15] as the 3D detec-\ntor and train the two detectors with infrastructure-view and\nvehicle-view data in VIC-Sync separately. Then, we con-\nvert the infrastructure predictions into the vehicle LiDAR\ncoordinate system and merge the prediction results with a\nmatcher based on the Euclidian distance measurement and\nthe Hungarian method [14] to generate fusion results.\nTo illustrate the temporal asynchrony problem, we also\nimplement the LiDAR detection late fusion baseline on the\nVIC-Async-k dataset. In addition, based on tracking and\nstate estimation we propose the Time Compensation Late\nFusion (TCLF) framework. The TCLF is mainly composed\nof the following three parts: 1) Estimating the velocity of\nthe objects with two adjacent infrastructure frames. 2) Esti-\nmating the state of the infrastructure objects at tv. 3) Fusing\nthe estimated infrastructure predictions and vehicle predic-\ntions following the way of LiDAR late fusion baseline. The\ndetails of the TCLF framework could be seen in Fig. 3.\nNote that we also report the evaluation results only with\nthe infrastructure data and only with the vehicle data, which\nare named as Veh.-Only and Inf.-Only respectively. The\nevaluation results are presented in Tab. 3.\nImage detection baseline with Late Fusion.\nTo exam-\nine image-only VIC3D object detection, we also implement\nthe late fusion framework only with infrastructure images\nand vehicle images. We choose ImvoxelNet [7] as the 3D\ndetector and train infrastructure detector and vehicle detec-\ntor with the corresponding part of VIC-Sync training data\nseparately. We implement the image detection late fusion\nfollowing the LiDAR detection late fusion.\n21366\n\n\nFigure 3. Time Compensation Late Fusion (TCLF) Framework. ∆t denotes the sampling interval of infrastructure sensors. We predict and\nmatch the boxes between two infrastructure frames. For matched vehicles, we compute their velocities directly. For unmatched vehicles,\nwe feed the position and motion information of the current scene into an MLP to predict their velocities. Finally, we can approximate the\npositions of vehicles at tv by linear interpolation, and fuse the results of the vehicle frame.\nLiDAR detection baseline with Early Fusion.\nTo ex-\nplore the fusion effect at the raw data level, we implement\nthe early fusion with PointPillars [15] as the 3D detector\non the VIC-Sync dataset. We first convert the infrastructure\npoint cloud in the VIC-Sync dataset into the vehicle LiDAR\ncoordinate system, then fuse the infrastructure point cloud\nand vehicle point cloud. We directly train and evaluate the\ndetector with the fused point cloud. Further to illustrate the\ntemporal asynchrony problem, we also implement the early\nfusion with PointPillars [15] on the VIC-Async-k dataset.\nFigure 4.\nPrediction results of the vehicle frame (orange) and\ninfrastructure frame (blue). We observe that infrastructure data\n(thick blue boxes) supplements the blind spot and extends the per-\nception field for the vehicle.\nFigure 5. Prediction results with and without time compensation.\nThe results of TCLF (blue) have a larger overlap with ground truth\n(black) than the results without time compensation (orange).\n5.1.2\nAnalysis\nHere we analyze the properties of the methods for the\nVIC3D object detection benchmark in Section 5.1.1.\nCooperative-view vs. Single-view.\nWe compare the per-\nformance of the methods whether using both infrastructure\ndata and vehicle data. In Tab. 3, the detection performance\nof late fusion is much better than the performance of Veh.-\nOnly or Inf.-Only, whether it is Image-based or LiDAR-\nbased or it is based on VIC-Sync dataset or VIC-Async-\nk dataset.\nFor example, the LiDAR detection with Late\nFusion achieves overall 41.90 AP points for 3D detection\nand overall 47.96 AP points for BEV detection on the VIC-\nSync dataset. However, the LiDAR detection only with ve-\nhicle data just achieves overall 31.33% AP for 3D detec-\ntion and overall 35.06% AP for BEV detection, and the Li-\nDAR detection only with infrastructure data just achieves\noverall 17.62% AP for 3D detection and overall 24.40%\nAP for BEV detection. The experiment results demonstrate\nthat fusing the infrastructure information can effectively im-\nprove the perception performance of the vehicle. This is\nmainly because infrastructure data provides supplementary\ninformation that makes up for the vehicle’s perception field.\nA visualization example is shown in Fig. 4.\nTemporal Asynchrony vs Time Compensation.\nTempo-\nral asynchrony brings challenges to fusing the infrastructure\ndata. Compared with the results on the VIC-Sync dataset,\nthe performance of LiDAR detection with fusion drops sig-\nnificantly on VIC-Async-k (2 points on VIC-Async-1 and\n6 points on VIC-Async-2). The decline is mainly due to\nthe state changes of moving objects, resulting in matching\ndifficulties and fusion errors. However, our TCLF can ef-\nfectively improve the performance of late fusion up to 0.5%\nAP and 1.5% AP on VIC-Async-1 and VIC-Async-2 re-\nspectively, which demonstrates that time compensation can\neffectively alleviate the temporal asynchrony problems es-\n21367\n\n\nTable 4. SV3D Detection Benchmark on DAIR-V2X-V\nModality\nModel\nVehicle3D(IoU=0.5)\nPedestrian3D(IoU=0.25)\nCyclist3D(IoU=0.25)\nEasy\nMiddle\nHard\nEasy\nMiddle\nHard\nEasy\nMiddle\nHard\nImage\nImvoxelNet [7]\n38.37\n24.28\n21.54\n4.54\n4.54\n4.54\n10.38\n9.09\n9.09\nPointCloud\nPointPillars [15]\n61.76\n49.02\n43.45\n33.40\n24.68\n22.39\n38.24\n33.80\n32.35\nPointCloud\nSECOND [27]\n69.44\n59.63\n57.63\n43.45\n39.06\n38.78\n44.21\n39.49\n37.74\nImage+PointCloud\nMVXNet [21]\n69.86\n60.74\n59.31\n47.73\n43.37\n42.49\n45.68\n41.84\n40.55\nTable 5. SV3D Detection Benchmark on DAIR-V2X-I\nModality\nModel\nVehicle3D(IoU=0.5)\nPedestrian3D(IoU=0.25)\nCyclist3D(IoU=0.25)\nEasy\nMiddle\nHard\nEasy\nMiddle\nHard\nEasy\nMiddle\nHard\nImage\nImvoxelNet [7]\n44.78\n37.58\n37.55\n6.81\n6.746\n6.73\n21.06\n13.57\n13.17\nPointCloud\nPointPillars [15]\n63.07\n54.00\n54.01\n38.53\n37.20\n37.28\n38.46\n22.60\n22.49\nPointCloud\nSECOND [27]\n71.47\n53.99\n54.00\n55.16\n52.49\n52.52\n54.68\n31.05\n31.19\nImage+PointCloud\nMVXNet [21]\n71.04\n53.71\n53.76\n55.83\n54.45\n54.40\n54.05\n30.79\n31.06\npecially when the time delay is larger. A visualization ex-\nample is provided in Fig. 5.\nEarly Fusion vs. Late Fusion.\nCompared with late fu-\nsion, early fusion achieves up to 8% AP higher under both\nBEV and 3D benchmarks, whether it is based on the VIC-\nSync dataset or the VIC-Async-1 dataset. However, early\nfusion should transmit the whole point cloud and suffers an\nextremely high transmission cost, which is about 4000 times\nmore than late fusion. For more practical applications, we\nencourage future research on achieving better performance\nwhile consuming less transmission bandwidth. We will also\nrelease the feature fusion for the benchmark in the future.\n5.2. Benchmark for SV3D Detection\nWe present an extensive 3D detection benchmark for\nthose who are interested in Single-View (SV) 3D detection\ntasks based on DAIR-V2X-V and DAIR-V2X-I datasets.\nCompared with the single-side data in DAIR-V2X-C, the\ntwo datasets are more diverse and could be more challeng-\ning to implement 3D object detection. Hence, we encourage\nresearchers who just aim at improving the performance of\nvehicle 3D object detection or infrastructure 3D object on\nDAIR-V2X-V and DAIR-V2X-I.\nWe split DAIR-V2X-V and DAIR-V2X-I datasets to\ntrain/valid/test part as 5:2:3 respectively. We present a num-\nber of baselines with methods based on different modalities\non the two datasets respectively: ImvoxelNet [7], PointPil-\nlars [15], SECOND [27] and MVXNet [21]. We evaluate\n3D object detection performance using the PASCAL criteria\nas KITTI [10], that distant objects are filtered out based on\ntheir bounding box height in the image plane. Three types\nof modes are used for evaluation, including Easy, Moder-\nate, and Hard modes. We implement these baselines with\nMMDetection3D Framework [1].\nEvaluation results are\nshown in Tab. 4 and Tab. 5.\n6. Conclusion\nIn this paper, we introduce DAIR-V2X, the first large-\nscale,\nmulti-modality,\nmulti-view dataset for vehicle-\ninfrastructure cooperative autonomous driving, and all\nframes are captured from real scenes with 3D annotations.\nWe also define VIC3D object detection to formulate the\nproblem of collaboratively locating and identifying 3D ob-\njects using sensory input from both vehicle and infras-\ntructure. In addition to solving traditional 3D object de-\ntection problems, the solution of VIC3D needs to con-\nsider the temporal asynchrony problem between vehicle and\ninfrastructure sensors and the data transmission cost be-\ntween them.\nTo facilitate future research, we provide a\nVIC3D benchmark for detection models with our proposed\nTime Compensation Late Fusion framework, as well as ex-\ntensive benchmarks for 3D detection on vehicle-view and\ninfrastructure-view datasets. Results show that integrating\ndata from infrastructure sensors achieves an average of 15%\nAP higher than single-vehicle 3D detection, and TCLF can\nalleviate temporal asynchrony problems.\nAcknowledgements\nWe thank Fan Yang, Ruiwen Zhang, Wenyue Wu, and\nXiao Wang from Baidu Inc. for the support in data process-\ning. We thank Jilei Mao, Taohua Zhou, Yingjuan Tang, Zan\nMao, and Zhiwen Yang for their support in the benchmark\nconstruction. Thanks to Beijing High-level Autonomous\nDriving Demonstration Area, Beijing Connected and Au-\ntonomous Vehicles Technology Co., Ltd, Baidu Apollo, and\nBeijing Academy of Artificial Intelligence for their support\nthroughout the dataset construction and release process.\n21368\n\n\nReferences\n[1] MMDetection3D: OpenMMLab next-generation platform\nfor general 3D object detection, 2020. 8\n[2] Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora,\nVenice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi-\nancarlo Baldan, and Oscar Beijbom.\nnuscenes: A multi-\nmodal dataset for autonomous driving. In 2020 IEEE Confer-\nence on Computer Vision and Pattern Recognition (CVPR),\npages 11618–11628, 2020. 2\n[3] Shanzhi Chen, Jinling Hu, Yan Shi, Ying Peng, Jiayi Fang,\nRui Zhao, and Li Zhao. Vehicle-to-everything (v2x) services\nsupported by lte-based systems and 5g. IEEE Communica-\ntions Standards Magazine, 1(2):70–76, 2017. 1\n[4] Xiaozhi Chen, Huimin Ma, Ji Wan, Bo Li, and Tian Xia.\nMulti-view 3d object detection network for autonomous\ndriving. In 2017 IEEE Conference on Computer Vision and\nPattern Recognition (CVPR), pages 6526–6534, 2017. 3\n[5] Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo\nRehfeld,\nMarkus Enzweiler,\nRodrigo Benenson,\nUwe\nFranke, Stefan Roth, and Bernt Schiele.\nThe cityscapes\ndataset for semantic urban scene understanding. In Proceed-\nings of the IEEE conference on computer vision and pattern\nrecognition, pages 3213–3223, 2016. 2\n[6] Yuepeng Cui, Hao Xu, Jianqing Wu, Yuan Sun, and Junxuan\nZhao. Automatic vehicle tracking with roadside lidar data\nfor the connected-vehicles system. IEEE Intelligent Systems,\n34(3):44–51, 2019. 3\n[7] Anton Konushin Danila Rukhovich,\nAnna Vorontsova.\nImvoxelnet: Image to voxels projection for monocular and\nmulti-view general-purpose 3d object detection.\narXiv\npreprint arXiv:2106.01178, 2021. 1, 3, 6, 8\n[8] Mark Everingham, Luc Van Gool, Christopher KI Williams,\nJohn Winn, and Andrew Zisserman. The pascal visual object\nclasses (voc) challenge. International journal of computer\nvision, 88(2):303–338, 2010. 5\n[9] Hongbo Gao, Bo Cheng, Jianqiang Wang, Keqiang Li, Jian-\nhui Zhao, and Deyi Li.\nObject classification using cnn-\nbased fusion of vision and lidar in autonomous vehicle en-\nvironment.\nIEEE Transactions on Industrial Informatics,\n14(9):4224–4231, 2018. 3\n[10] Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we\nready for autonomous driving? the kitti vision benchmark\nsuite. In 2012 IEEE Conference on Computer Vision and\nPattern Recognition (CVPR), pages 3354–3361, 2012. 2, 8\n[11] Mohammad-Hashem Haghbayan,\nFahimeh Farahnakian,\nJonne Poikonen, Markus Laurinen, Paavo Nevalainen, Juha\nPlosila, and Jukka Heikkonen. An efficient multi-sensor fu-\nsion approach for object detection in maritime environments.\nIn 2018 21st International Conference on Intelligent Trans-\nportation Systems (ITSC), pages 2163–2170, 2018. 3\n[12] Xinyu Huang, Peng Wang, Xinjing Cheng, Dingfu Zhou,\nQichuan Geng, and Ruigang Yang. The apolloscape open\ndataset for autonomous driving and its application.\nIEEE\ntransactions on pattern analysis and machine intelligence,\n42(10):2702–2719, 2019. 2\n[13] Robert Krajewski, Julian Bock, Laurent Kloeker, and Lutz\nEckstein.\nThe highd dataset: A drone dataset of natural-\nistic vehicle trajectories on german highways for validation\nof highly automated driving systems.\nIn 2018 21st Inter-\nnational Conference on Intelligent Transportation Systems\n(ITSC), pages 2118–2125. IEEE, 2018. 2\n[14] Harold W. Kuhn. The hungarian method for the assignment\nproblem. In 50 Years of Integer Programming, 2010. 6\n[15] Alex H. Lang, Sourabh Vora, Holger Caesar, Lubing Zhou,\nJiong Yang, and Oscar Beijbom. Pointpillars: Fast encoders\nfor object detection from point clouds. Proceedings of the\nIEEE Conference on Computer Vision and Pattern Recogni-\ntion, 2019. 1, 3, 6, 7, 8\n[16] Yiming Li, Shunli Ren, Pengxiang Wu, Siheng Chen, Chen\nFeng, and Wenjun Zhang.\nLearning distilled collabo-\nration graph for multi-agent perception.\narXiv preprint\narXiv:2111.00643, 2021. 2\n[17] Jiageng Mao, Minzhe Niu, Chenhan Jiang, Hanxue Liang,\nJingheng Chen, Xiaodan Liang, Yamin Li, Chaoqiang Ye,\nWei Zhang, Zhenguo Li, et al.\nOne million scenes\nfor autonomous driving:\nOnce dataset.\narXiv preprint\narXiv:2106.11037, 2021. 2\n[18] Cody Reading, Ali Harakeh, Julia Chae, and Steven L\nWaslander.\nCategorical depth distribution network for\nmonocular 3d object detection.\nIn Proceedings of the\nIEEE/CVF Conference on Computer Vision and Pattern\nRecognition, pages 8555–8564, 2021. 1\n[19] German Ros, Laura Sellart, Joanna Materzynska, David\nVazquez, and Antonio M Lopez. The synthia dataset: A large\ncollection of synthetic images for semantic segmentation of\nurban scenes.\nIn Proceedings of the IEEE conference on\ncomputer vision and pattern recognition, pages 3234–3243,\n2016. 2\n[20] Shaoshuai Shi, Xiaogang Wang, and Hongsheng Li. Pointr-\ncnn: 3d object proposal generation and detection from point\ncloud. In Proceedings of the IEEE/CVF conference on com-\nputer vision and pattern recognition, pages 770–779, 2019.\n1\n[21] Vishwanath A. Sindagi, Yin Zhou, and Oncel Tuzel. Mvx-\nnet: Multimodal voxelnet for 3d object detection. In 2019 In-\nternational Conference on Robotics and Automation (ICRA),\npages 7276–7282, 2019. 1, 3, 8\n[22] Carlos Renato Storck and F´\natima Duarte-Figueiredo.\nA\n5g v2x ecosystem providing internet of vehicles. Sensors,\n19(3):550, 2019. 1\n[23] Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien\nChouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou,\nYuning Chai, Benjamin Caine, et al. Scalability in perception\nfor autonomous driving: Waymo open dataset. In Proceed-\nings of the IEEE/CVF Conference on Computer Vision and\nPattern Recognition, pages 2446–2454, 2020. 2\n[24] Sourabh Vora, Alex H. Lang, Bassam Helou, and Oscar Bei-\njbom. Pointpainting: Sequential fusion for 3d object detec-\ntion. In 2020 IEEE Conference on Computer Vision and Pat-\ntern Recognition (CVPR), pages 4603–4611, 2020. 1, 3\n[25] Tsun-Hsuan Wang, Sivabalan Manivasagam, Ming Liang,\nBinh Yang, Wenyuan Zeng, James Tu, and Raquel Urtasun.\nV2vnet: Vehicle-to-vehicle communication for joint percep-\ntion and prediction. In ECCV, 2020. 3\n[26] Zhangjing Wang, Yu Wu, and Qingqing Niu. Multi-sensor\nfusion in automated driving:\nA survey.\nIEEE Access,\n8:2847–2868, 2020. 3\n21369\n\n\n[27] Yan Yan, Yuxing Mao, and Bo Li. Second: Sparsely embed-\nded convolutional detection. Sensors, 2018. 3, 8\n[28] Zetong Yang, Yanan Sun, Shu Liu, and Jiaya Jia.\n3dssd:\nPoint-based 3d single stage object detector.\nProceedings\nof the IEEE Conference on Computer Vision and Pattern\nRecognition, 2020. 1, 3\n[29] Xiaoqing Ye, Mao Shu, Hanyu Li, Yifeng Shi, Yingying\nLi, Guangjie Wang, Xiao Tan, and Errui Ding.\nRope3d:\nThe roadside perception dataset for autonomous driving and\nmonocular 3d object detection.\nIn IEEE/CVF Conference\non Computer Vision and Pattern Recognition (CVPR), June\n2022. 3\n[30] Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying\nChen, Fangchen Liu, Vashisht Madhavan, and Trevor Dar-\nrell. Bdd100k: A diverse driving dataset for heterogeneous\nmultitask learning. In Proceedings of the IEEE/CVF con-\nference on computer vision and pattern recognition, pages\n2636–2645, 2020. 2\n[31] Junxuan Zhao, Hao Xu, Hongchao Liu, Jianqing Wu, Yichen\nZheng, and Dayong Wu. Detection and tracking of pedestri-\nans and vehicles using roadside lidar sensors. Transportation\nResearch Part C: Emerging Technologies, 100:68–87, 2019.\n3\n21370\n\n\n10914\nIEEE ROBOTICS AND AUTOMATION LETTERS, VOL. 7, NO. 4, OCTOBER 2022\nV2X-Sim: Multi-Agent Collaborative Perception\nDataset and Benchmark for Autonomous Driving\nYiming Li\n, Student Member, IEEE, Dekun Ma, Ziyan An, Zixun Wang, Student Member, IEEE, Yiqi Zhong,\nSiheng Chen\n, and Chen Feng\n, Member, IEEE\nAbstract—Vehicle-to-everything (V2X) communication tech-\nniques enable the collaboration between vehicles and many other\nentities in the neighboring environment, which could fundamen-\ntally improve the perception system for autonomous driving. How-\never, the lack of a public dataset signiﬁcantly restricts the re-\nsearch progress of collaborative perception. To ﬁll this gap, we\npresent V2X-Sim, a comprehensive simulated multi-agent percep-\ntion dataset for V2X-aided autonomous driving. V2X-Sim pro-\nvides: (1) multi-agent sensor recordings from the road-side unit\n(RSU) and multiple vehicles that enable collaborative perception,\n(2) multi-modality sensor streams that facilitate multi-modality\nperception, and (3) diverse ground truths that support various\nperception tasks. Meanwhile, we build an open-source testbed and\nprovide a benchmark for the state-of-the-art collaborative percep-\ntion algorithms on three tasks, including detection, tracking and\nsegmentation. V2X-Sim seeks to stimulate collaborative perception\nresearch for autonomous driving before realistic datasets become\nwidely available.\nIndex Terms—Deep learning for visual perception, multi-robot\nsystems, data sets for robotic vision.\nI. INTRODUCTION\nP\nERCEPTION is a fundamental capability for autonomous\nvehicles, which allows them to represent, identify, and\ninterpret sensory input for understanding the complex surround-\nings. In literature, single-vehicle perception has been intensively\nstudied thanks to the well-established driving datasets [1]–[3],\nand researchers have proposed various algorithms to deal with\ndifferent downstream tasks [4]–[6].\nDespite recent advances in single-vehicle perception, the\nindividual viewpoint often results in degraded perception in\nManuscript received 24 February 2022; accepted 30 June 2022. Date of\npublication 21 July 2022; date of current version 23 August 2022. This letter\nwas recommended for publication by Associate Editor I. Gilitschenski and\nEditor C. C. Lerma upon evaluation of the reviewers’ comments. This work\nwas supported in part by the NSF CPS Program under Grant CMMI-1932187\nand CNS-2121391, in part by the National Natural Science Foundation of China\nunder Grant 62171276, in part by the Science and Technology Commission of\nShanghai Municipal under Grant 21511100900, and in part by CALT under\nGrant 2021-01. (Corresponding authors: Siheng Chen; Chen Feng.)\nYiming Li, Dekun Ma, Ziyan An, Zixun Wang, and Chen Feng are with\nthe New York University, Brooklyn, NY 11201 USA (e-mail: yimingli9702\n@gmail.com; dm4524@nyu.edu; annieziyan1222@gmail.com; craddywang@\ngmail.com; cfeng@nyu.edu).\nYiqi Zhong is with the University of Southern California, Los Angeles 91103\nUSA (e-mail: yiqizhon@usc.edu).\nSiheng Chen is with the Cooperative Medianet Innovation Center, Shanghai\nJiao Tong University and Shanghai AI Laboratory, Shanghai 200240, China\n(e-mail: sihengc@sjtu.edu.cn).\nDigital Object Identiﬁer 10.1109/LRA.2022.3192802\nOur dataset and code are available at https://ai4ce.github.io/V2X-Sim/\nFig. 1.\n(a) Intersection for vehicle-to-everything (V2X) communication. (b)\nWorkﬂow of multi-agent collaborative perception with intermediate-/feature-\nbased strategy. We benchmark collaborative object detection, multi-object track-\ning, and semantic segmentation in the bird’s eye view (BEV).\nlong-range or occluded areas. A promising solution to this prob-\nlem is through vehicle-to-everything (V2X) [7], a cutting-edge\ncommunication technology that enables dialogue between a\nvehicle and other entities, including vehicle-to-vehicle (V2V)\nand vehicle-to-infrastructure (V2I). With the aid of V2X com-\nmunication, we are able to upgrade single-vehicle perception\nto collaborative perception, which introduces more viewpoints\nto help autonomous vehicles see further, better and even see\nthrough occlusion, thereby fundamentally enhancing the capa-\nbility of perception.\nCollaborative perception naturally draws on communica-\ntion and perception. Its development requires expertise from\nboth communities. Recently, the communication community\nhas made enormous efforts to promote such a study [8]–[10];\nhowever, only a few works have been proposed from the per-\nspective of perception [11]–[15]. One major reason for this is\nthe lack of well-designed and organized collaborative perception\ndatasets. Due to the immaturity of V2X and the cost of simul-\ntaneously operating multiple autonomous vehicles, it is very\nexpensive and laborious to build such a real dataset for research\ncommunities. Therefore, we synthesize a comprehensive and\npublicly available dataset, named as V2X-Sim, to advance the\nstudy of collaborative perception for V2X-communication-aided\nautonomous driving.\nTo generate V2X-Sim, we employ SUMO [16], a micro-trafﬁc\nsimulation, to produce numerically-realistic trafﬁc ﬂow, and\nCARLA [17], a widely-used open-source simulator for au-\ntonomous driving research, to retrieve well-synchronized sensor\nstreams from multiple vehicles as well as the road-side unit\n(RSU). Meanwhile, multi-modality sensor streams of different\nentities are recorded to enable cross-modality perception. In\naddition, diverse annotations including bounding boxes, vehicle\n2377-3766 © 2022 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission.\nSee https://www.ieee.org/publications/rights/index.html for more information.\nAuthorized licensed use limited to: BEIJING UNIVERSITY OF POST AND TELECOM. Downloaded on November 07,2023 at 07:46:33 UTC from IEEE Xplore.  Restrictions apply. \n\n\nLI et al.: V2X-SIM: MULTI-AGENT COLLABORATIVE PERCEPTION DATASET AND BENCHMARK FOR AUTONOMOUS DRIVING\n10915\nTABLE I\nCOMPARISON OF COLLABORATIVE PERCEPTION DATASETS FOR AUTONOMOUS\nDRIVING. THERE ARE NO PUBLIC DATASETS WHICH SUPPORT BOTH V2V AND\nV2I RESEARCH: MULTI-AGENT DATA ARE EITHER GENERATED BY\nSIMULATORS [12], [20] OR CREATED BY SELECTING CONSECUTIVE FRAMES\nFROM SINGLE-AGENT REAL DATASETS [11], [21], [22]. SEVERAL WORKS\nCOLLECT DATA FROM MULTIPLE INFRASTRUCTURE SENSORS: [23] IN\nSIMULATION, [24], [25] IN REAL WORLD. OUR DATASET IS THE FIRST PUBLIC\nMULTI-AGENT MULTI-MODALITY DATASET WHICH SUPPORTS DIFFERENT\nCOLLABORATIVE PERCEPTION TASKS.\ntrajectories, and semantic labels are provided to facilitate various\ndownstream tasks. To better serve multi-agent, multi-modality,\nand multi-task perception research for autonomous driving,\nwe further provide a benchmark for three crucial perception\ntasks (collaborative detection, tracking, and segmentation) on\nthe proposed dataset using the state-of-the-art collaboration\nstrategies [12], [13], [18], [19]. In summary, our contributions\nare two-fold:\nr We propose V2X-Sim, a comprehensive V2X perception\ndataset for autonomous driving, to support multi-agent\nmulti-modality multi-task perception research.\nr We create an open-source testbed for collaborative percep-\ntion methods, and provide a benchmark on three tasks to\nencourage more research in this ﬁeld.\nII. RELATED WORK\nAutonomous driving dataset: Since the pioneering dataset\nKITTI [2] was released, the autonomous driving community has\nbeentryingtoincreasethedatasetcomprehensivenessintermsof\ndriving scenarios, sensor modalities, and data annotations. Re-\ngarding driving scenarios, current datasets cover crowded urban\nscenes [29], adverse weather conditions [30], night scenes [31],\nand multiple cities [1] to enrich the data distribution. As for\nsensor modalities, nuScenes [1] collects data with Radar, RGB\ncameras, and LiDAR in a 360◦viewpoint; WoodScape [32]\ncaptures data with ﬁsheye cameras. Regarding data annotations,\nsemantic labels in both images [33]–[36] and point clouds [37],\n[38] are provided to enable semantic segmentation; 2D/3D\nbox trajectories are offered [39], [40] to facilitate tracking and\nprediction. In summary, most real datasets emphasize the data\ncomprehensiveness in single-vehicle situations, but ignore the\nmulti-vehicle scenarios.\nV2X system and dataset: By sharing information with other\nvehicles or the RSU, V2X mitigates the shortcomings of single-\nvehicle perception and planning such as the limited sensing\nrange and frequent occlusion [7]. Previous research [41] has\ndeveloped an enhanced cooperative microscopic trafﬁc model\nin V2X scenarios, and studied the effect of V2X in trafﬁc\ndisturbance scenarios. [42] proposes a multi-modal cooperative\nperception system that provides see-through, lifted-seat, satellite\nand all-around views to drivers. More recently, COOPER [11]\ninvestigates raw-data level collaborative perception to improve\nthe detection capability for autonomous driving. V2VNet [12]\nproposes intermediate-feature level collaboration to promote the\nvehicle’s perception and prediction capability. Several works\nuse multiple infrastructure sensors to jointly perceive the en-\nvironment and employ output-level collaboration with vehicle-\nto-infrastructure communication [23], [24]. As for the datasets,\n[11], [21], [22] simulate the V2V scenarios with different frames\nfrom KITTI [2]. Yet, these datasets are unrealistic multi-agent\ndatasets for the measurements are not captured by different\nviewpoints. Some other works use a platoon strategy for data\ncapture [43], [44], but they are biased because the observations\nwere highly correlated with each other. The most relevant work\nis V2V-Sim [12], which is based on a high-quality LiDAR sim-\nulator [26]. Unfortunately, V2V-Sim does not include the V2I\nscenario and is not publicly available. Moreover, OPV2V [20]\nand CODD [28] only support the detection task in the V2V\nscenario. Existing collaborative perception datasets are summa-\nrized in Table I: V2X-Sim1 is currently the most comprehensive\none with multi-agent multi-modality sensory streams in both\nV2V and V2I scenarios, and can support various downstream\ntasks such as multi-agent collaborative detection, tracking, and\nsemantic segmentation.\nIII. V2X-SIM DATASET\nV2X-Sim could enable more research on the collaboration\nstrategy among vehicles to achieve a more robust perception.\nThis could fundamentally beneﬁt autonomous driving, intelli-\ngent transportation systems, smart cities, etc..\nA. Sensor Suite on Vehicles and RSU\nMulti-modality data is essential for robust perception. To\nensure the comprehensiveness of our dataset, we equip each\nvehicle with a sensor suite composed of RGB cameras, LiDAR,\nGPS, IMU, and RSU with RGB cameras and LiDAR.\nSensor conﬁguration: On both the vehicle and RSU, the\ncamera and LiDAR cover 360◦horizontally to enable full-view\nperception. Speciﬁcally, each vehicle carries six RGB cameras\nfollowing the nuScenes conﬁguration [1]; the RSU is equipped\nwith four RGB cameras toward four directions at the cross-\nroad. We employ depth camera, semantic segmentation camera,\nsemantic LiDAR, and BEV semantic segmentation camera in\nCARLA to obtain the corresponding ground-truth for each RGB\ncamera. Note that the BEV semantic segmentation camera is\nbased on orthogonal projection while the ego-vehicle seman-\ntic segmentation camera uses perspective projection. Table II\nsummarizes the detailed sensor speciﬁcation.\nSensor layout and coordinate system: The overall sensor\nlayout and coordinate system is shown in Fig. 2. The BEV\nsemantic camera shares the same x, y position with LiDAR yet\nis placed higher to ensure a certain size of ﬁeld of view. Note that\nwe invert the y-axis in CARLA and use a right-hand coordinate\nsystem following nuScenes [1].\nDiverse annotations: To assist downstream tasks including\ndetection, tracking and semantic segmentation, we provide var-\nious annotations such as 3D bounding boxes, pixel-wise and\npoint-wisesemanticlabels.Eachboxisdeﬁnedbythelocationof\nits center in x, y, z coordinates, and its width, length, and height.\nIn total, there are 23 categories such as the pedestrian, building,\nground, etc. In addition, precise depth values are provided for\ndepth estimation.\n1This work extends the LiDAR-based V2V data in our previous work [13]\nwith more modalities, scenarios and downstream tasks.\nAuthorized licensed use limited to: BEIJING UNIVERSITY OF POST AND TELECOM. Downloaded on November 07,2023 at 07:46:33 UTC from IEEE Xplore.  Restrictions apply. \n\n\n10916\nIEEE ROBOTICS AND AUTOMATION LETTERS, VOL. 7, NO. 4, OCTOBER 2022\nTABLE II\nSENSOR SUITE OF VEHICLE (V) AND INTERSECTION (I).\nB. CARLA-SUMO Co-Simulation\nWeconsideritarealisticV2Xscenariowhenmultiplevehicles\nwith their own routes are simultaneously located in the same\nintersection. Each intersection is also equipped with one RSU\nwith sensing capability. Regarding the trafﬁc simulation, there\nare several non-public simulators which can explicitly generate\ndata tailored for collaborative perception such as scenes with\nocclusion, and sensor range limitations [45], [46]. Yet in this\nwork, we use open-source CARLA-SUMO co-simulation for\ntrafﬁc ﬂow simulation and data recording. Vehicles are spawned\nin CARLA via SUMO to roam around in the town with random\nroutes. Hundreds of vehicles are spawned in different towns\n(Town03, Town04 and Town05 that have crossroads as well as\nT-junctionsinboththecrowdeddowntownandsuburbhighway).\nWe record several log ﬁles in different towns. Afterwards, at\ndifferent junctions, we read out 100 scenes from the log ﬁles.\nEach scene has a 20-second duration, and we choose M(M =\n2, 3, 4, 5) vehicles as well as one RSU in each scene as intelligent\nagents to share information. Example scenarios are shown in\nFig. 3(a).\nC. Downstream Tasks\nOur dataset not only supports single-agent perception tasks\nsuch as 3D object detection, tracking, image-/point-cloud-based\nsemantic segmentation, depth estimation, but also enables col-\nlaborative perception such as collaborative 3D object detection,\ntracking, and collaborative BEV semantic segmentation in ur-\nban driving scenes. We provide a benchmark for collaborative\nperception algorithms.\nD. Dataset Statistics\nAnnotation statistics: We provide statistics on the annotations\nand objects to highlight the inclusiveness and diversity of our\ndataset. In Fig. 3(b) we analyze the size of the cars’ bounding\nboxes within a 70 m range from ego vehicles in each frame. The\ngreat variation of car sizes indicates that our scenes contain a\ndiverse set of car makes and models that well includes most of\nthe common real-world vehicles. Fig. 3(c) shows the annotation\ncount in each frame for vehicles within 0-30 m, 30-50 m, and\n50-70 m ranges from each ego vehicle. It suggests that our\ndataset features both crowded scenes (up to 100 annotations\nwithin 50-70 m from the ego vehicle) and less crowded driving\nscenarios (as low as 10 annotations within 30 m from the ego\nvehicle). Fig. 3(d) contains statistics on the number of LiDAR\npoints per annotation for single-agent and multi-agent scenarios\nFig. 2.\nSensor layout and coordinate systems.\nrespectively. The number of total LiDAR points of each object\nannotation increases when there are more than one agents ob-\nserving the same object. Speciﬁcally, for a single agent, there\nare 183.83 points in each annotation on average, but the number\ngoes up to 875.59 points per annotation for multiple agents.\nScene features: We analyze the distance between each two ego\nvehicles for every frame, as shown in Fig. 3(e). An overwhelm-\ning percentage of the ego vehicles are presented within 20-30\nmeters from each other, suggesting they are geographically\nclosely connected. The speed of cars within 70 m from ego\nvehicles are shown in Fig. 3(f). Given the fact that our scenes\nare selected near intersections, we notice that a major fraction\nof vehicles are slower than 10 km/h. However, the maximum\nspeed is as high as 90+ km/h. Fig. 3(g) shows the percentage of\nannotations observed by a certain number of ego vehicles, up to\n5. Over 60% of the annotations are observed by at least two ego\nvehicles.\nIV. COLLABORATIVE PERCEPTION BENCHMARK\nWe benchmark three crucial perception tasks in autonomous\ndriving within the collaboration setting: detection, tracking, and\nsemantic segmentation. The three tasks have been extensively\nstudied since they generate essential perception knowledge for\nautonomous vehicles to make safer decisions. For performance\nevaluation, we follow the same evaluation protocol of the three\ntasksinthesingle-agentscenario,exceptthatweutilizetheinfor-\nmation shared by other agents while the single-agent perception\ndo not have access to such information. We report the results in\ntwo scenarios: (1) V2V only, and (2) V2X (V2V + V2I).\nDataset format and split: Our V2X-Sim dataset follows the\nsame storage format of nuScenes [1], an authoritative multi-\nmodality single-agent autonomous driving dataset. nuScenes\ncollected real-world single-agent driving data to promote the\nsingle-agent autonomous driving research; while we simulate\nmulti-agent scenarios to facilitate the next-generation V2X-\naided collaborative perception technology. Each scene in our\ndataset contains a 20 s trafﬁc ﬂow at a certain intersection\nof three CARLA towns, and the multi-modality multi-agent\nsensory streams are recorded at 5 Hz, meaning each scene is\ncomposed of 100 frames. We generate 100 scenes with a total\nof 10,000 frames, and in each scene, 2-5 vehicles are selected\nas the collaboration agents. We use 8,000/1,000/1,000 frames\nfor training/validation/testing respectively, and we ensure that\nthere is no overlap in terms of the intersections across train-\ning/validation/testing set. Each frame has data sampled from\nmultiple agents (vehicles and RSU). There are 37,200 samples\nin the training set, 5,000 samples in the validation set, and 5,000\nsamples in the test set. The split is shared across tasks.\nAuthorized licensed use limited to: BEIJING UNIVERSITY OF POST AND TELECOM. Downloaded on November 07,2023 at 07:46:33 UTC from IEEE Xplore.  Restrictions apply. \n\n\nLI et al.: V2X-SIM: MULTI-AGENT COLLABORATIVE PERCEPTION DATASET AND BENCHMARK FOR AUTONOMOUS DRIVING\n10917\nFig. 3.\n(a) Visualizations of the bird’s eye view point cloud from different scenes. Gray denotes the point cloud captured by the RSU LiDAR. Each color (except\nfor gray) represents a vehicle, and the orange boxes denote the vehicles in the scene. (b) Statistics for car bounding box sizes. (c) Counts for annotations per\nkeyframe where the annotated vehicles are presented within 0-30 m, 30-50 m, and 50-70 m of the ego vehicles. (d) Counts for LiDAR points per annotation.\n(e) Statistics of the distance between every two ego vehicles for all frames. (f) Speed of cars located within 70 m from ego vehicles. (g) Percentage of annotated\nvehicles observed by 1-5 ego vehicles.\nImplementation details: Bird’s-eye-view (BEV) is a widely\nused and powerful representation in autonomous driving be-\ncause it can describe the surrounding objects and overall context\nvia a compact top-down 2D map [47]. Therefore, BEV-based\nrepresentation is adopted in all three tasks: we use a 3D voxel\ngrid to represent the 3D world, employ binary representation,\nand assign each voxel a positive label if the voxel contains\npoint cloud data. Since the generated 3D voxel grid can be\nconsidered as a pseudo-image whose height dimension is the\nchannel dimension, we can perform the efﬁcient 2D convolution\ninstead of the heavy 3D convolution. Speciﬁcally, we crop the\npoints located in the region of [−32, 32] × [−32, 32] × [−3, 2]\nmeters for vehicles deﬁned in the ego-vehicle Euclidean coor-\ndinate and [−32, 32] × [−32, 32] × [−8, −3] meters for RSU.\nThe width and length of each voxel are 0.25 m, and the height\nis 0.4 m, meaning the generated BEV-based pseudo-image has\na dimension of 256 × 256 × 13 (W × L × H). Note that the\nmodels in all the tasks consume the 3D BEV map and generate\nperception results in 2D BEV.\nBenchmark models: We aim to benchmark collaborative\nperception strategies rather than the well-studied individual\nperception methods. We consider early/intermediate/late/no\ncollaboration models for the benchmark. The intermediate mod-\nels, including DiscoNet [13], V2VNet [12], When2com [18],\nand Who2com [19], are based on the communication of the\nintermediate features of the neural network. The methods in\nour benchmark are as follows:\nr Lower-bound: The single-agent perception model without\ncollaboration which processes a single-view point cloud is\nconsidered as the lower-bound.\nr Co-lower-bound: Collaborative lower-bound fuses the out-\nput from different single-agent perception models.\nr Upper-bound: The early collaboration model which trans-\nmits raw point cloud data is the upper-bound.\nr DiscoNet [13]: DiscoNet uses a directed collaboration\ngraph with matrix-valued edge weight to adaptively high-\nlight the informative spatial regions and reject the noisy\nregions of the messages sent by the partners. After adaptive\nmessage fusion, the updated features will be transmitted to\nthe output head for perception.\nr V2VNet [12]: V2VNet uses a pose-aware graph neural\nnetwork to propagate agents’ information, and employs\na convolutional gated recurrent unit based module to ag-\ngregate other agents’ information. After several rounds of\nneural message passing, the updated features are fed into\nthe output head to generate perception results.\nr When2com [18]: When2com employs attention-based\nmechanism for communication group construction: the\npartners with satisfactory correlation scores will be se-\nlected as the collaborators. After the attention-score-based\nweighted fusion, the updated features will be fed into\nthe output head for perceptions.\nThe model with pose\ninformation is marked by ∗.\nr Who2com [19]: Who2com shares a similar idea with\nWhen2com, yet it uses handshake mechanism to select the\ncollaborator: the partner with the highest score will be se-\nlected as the collaborator. The model with pose information\nis marked by ∗.\nWe implement a 3D perception pipeline that can be integrated\nwith all of the communication methods mentioned above. Since\nthe source codes of V2VNet is not publicly available, we re-\nimplement the V2VNet in PyTorch according to its pseudo-\ncode. For when2com/who2com, we borrow their communica-\ntion modules from its ofﬁcial code. All of the intermediate\ncollaboration modules use the same architecture and conduct\ncollaboration at the same intermediate feature layer. Moreover,\nall of the methods are trained with the same setting to ensure\nthat the performance gain comes from the collaboration instead\nof irrelevant techniques.\nA. Collaborative Object Detection in BEV\nProblem deﬁnition: As the most crucial perception task in\nautonomous driving, 3D object detection aims to recognize and\nlocalize the objects in 3D scenes given a single frame, with the\nfollowing tracking, prediction and planning modules all heavily\nrelying on the detections. The models consume voxelized point\ncloud and output BEV bounding boxes.\nBackbone and evaluation: We use a classic anchor-based\ndetector composed of a convolutional encoder, a convolutional\nAuthorized licensed use limited to: BEIJING UNIVERSITY OF POST AND TELECOM. Downloaded on November 07,2023 at 07:46:33 UTC from IEEE Xplore.  Restrictions apply. \n\n\n10918\nIEEE ROBOTICS AND AUTOMATION LETTERS, VOL. 7, NO. 4, OCTOBER 2022\nFig. 4.\nVisualizations of BEV detection on V2X-Sim. Red and green boxes are the predictions and ground-truths respectively.\nTABLE III\nQUANTITATIVE RESULTS OF COLLABORATIVE BEV DETECTION ON V2X-SIM\nRSU Denotes the Road-Side Unit. AP Denotes the Average Precision.\nΔ is the Absolute Gain in AP Introduced by RSU\ndecoder, and an output header for classiﬁcation and regres-\nsion [48]. Regarding the loss function, we use the binary cross-\nentropy loss to supervise the box classiﬁcation and the smooth\nL1 loss to supervise the box regression, following [48]. We\nemploy the generic BEV detection evaluation metric: Average\nPrecision (AP) at Intersection-over-Union (IoU) thresholds of\n0.5 and 0.7. We target the vehicle detection and report the results\non the test set.\nQuantitative results: Table III demonstrates the quantitative\ncomparisons on AP (@IoU = 0.5/0.7). We ﬁnd that: (1) the\nupper-boundperformsbestamongstallmethods,anditimproves\nlower-bound signiﬁcantly by 41.1% and 51.6% in terms of\nAP@0.5 and AP@0.7 in the scenario of V2V only, validating\nthe necessity of collaboration; (2) V2V and V2I jointly can\ngenerally improve the perception over V2V only with more\nviewpoints, e.g., adding V2I can bring an improvement of 9.4%\nfor upper-bound, and 5.5% for V2VNet in terms of AP@0.5;\n(3) among the intermediate models, DiscoNet achieves the\nbest performance via the well-designed distilled collaboration\ngraph, V2VNet achieves the second best performance via the\npowerful neural message passing, and When2com/Who2com\nonly achieve comparable performance with lower-bound since\nthe attention-mechanism is not suitable in point-cloud-based\ncollaborative perception: the agents usually need complemen-\ntary information rather than a highly similar one; (4) the\nlate collaboration model (co-lower-bound) hurts the detection\nperformance because of introducing extra false positives from\nother vehicles.\nQualitative results: The qualitative results are shown in Fig. 4.\nWe see that the collaboration can fundamentally mitigate the\nproblems of long-range perception and occlusion.\nB. Collaborative Multi-Object Tracking in BEV\nProblem deﬁnition: Different from detection, multi-object\ntracking requires the generation of temporally consistent per-\nception results. Multi-object tracking in BEV is to use bounding\nboxes, object categories, and object identities to track different\nobjects within a temporal sequence.\nEvaluation metrics: We mainly utilize HOTA (Higher Order\nTracking Accuracy) [49] to evaluate our BEV tracking perfor-\nmance. HOTA can evaluate detection, association, and local-\nization performance via a single uniﬁed metric. In addition, the\nclassic multi-object tracking accuracy (MOTA) and multi-object\ntracking precision (MOTP) [50] are also employed. MOTA can\nmeasure detection errors and association errors. MOTP solely\nmeasures localization accuracy.\nBaseline tracker: We implement SORT [51] as our baseline\ntracker. Given the detection results, SORT will combine the\nKalman Filter and Hungarian algorithm to achieve an accurate\nand efﬁcient tracking performance.\nQuantitative results: Quantitative comparisons of BEV track-\ning are shown in Table IV. Similar to BEV detection, upper-\nbound achieves the best performance in terms of MOTA and\nHOTA. Meanwhile, adding V2I can improve MOTA largely yet\ncannot help too much in MOTP. Co-lower-bound shows good\nperformanceinlocalizationaccuracy(MOTP).Amoreadvanced\ntracker is required to exploit the collaboration for ﬁlling the\nperformance gap.\nC. Collaborative Semantic Segmentation in BEV\nProblem deﬁnition: We aim to conduct semantic segmentation\nin BEV using only geometry point cloud. In the collabora-\ntive perception scenarios, there are measurements collected by\nmultiple agents with distinct viewpoints. Therefore, there are\nmore information in the scene, facilitating the semantic scene\nunderstanding.\nAuthorized licensed use limited to: BEIJING UNIVERSITY OF POST AND TELECOM. Downloaded on November 07,2023 at 07:46:33 UTC from IEEE Xplore.  Restrictions apply. \n\n\nLI et al.: V2X-SIM: MULTI-AGENT COLLABORATIVE PERCEPTION DATASET AND BENCHMARK FOR AUTONOMOUS DRIVING\n10919\nFig. 5.\nVisualizations of collaborative BEV semantic segmentation.\nTABLE IV\nQUANTITATIVE RESULTS OF BEV TRACKING ON V2X-SIM. MOTA: MULTIPLE OBJECT TRACKING ACCURACY. MOTP: MULTIPLE OBJECT TRACKING PRECISION.\nHOTA: HIGHER ORDER TRACKING ACCURACY. DETA: DETECTION ACCURACY. ASSA: ASSOCIATION ACCURACY. DETRE: DETECTION RECALL. DETPR:\nDETECTION PRECISION. ASSRE: ASSOCIATION RECALL. ASSPR: ASSOCIATION PRECISION. LOCA: LOCALIZATION ACCURACY. THE NUMBER TO THE LEFT OF ()\nDENOTES THE PERFORMANCE IN V2V SOLELY. THE NUMBER IN () REPRESENTS THE PERFORMANCE GAIN BY ADDING V2I\nTABLE V\nQUANTITATIVE RESULTS OF BEV SEGMENTATION ON V2X-SIM. THE NUMBER TO THE LEFT OF () DENOTES THE PERFORMANCE IN V2V SOLELY\nThe Number in () Represents the Performance Gain by Adding V2I\nBaseline segmentation method and evaluation metrics: We\nfollow the backbone architecture as well as the loss function\nof U-Net [52] in our baseline method. The input is a BEV-\nbased voxelized point cloud, and the output is BEV semantic\nsegmentation. We label and predict seven categories as listed\nin Table V, while the remaining is unlabeled. In our bench-\nmark, we evaluate the segmentation performance using mean\nIntersection-over-Union (mIoU).\nQuantitative results: As illustrated in Table V, we ﬁnd that:\n(1) V2VNet, DiscoNet, and upper-bound achieve comparable\nperformance in terms of terrain and road categories; (2) the\nattention-based methods (i.e., when2com, who2com) performs\nworse because the attention-based mechanisms try to ﬁnd the\ncollaboration partners with high correlation scores. Whereas,\nin 3D perception, the collaborators with complementary in-\nformation should be prioritized during the collaboration. Such\nnonalignment can make it quite hard for the attention model\nto learn; (3) there is a large gap between lower-bound and\nupper-boundregardingthetwosafety-critical categories: vehicle\n(45.93% v.s. 64.09%) and pedestrian (20.59% v.s. 31.54%),\nproving the values of collaboration; (4) employing V2V and\nV2I jointly can generally enhance the vehicle segmentation over\nusing V2V solely; (5) co-lower-bound performs better than the\nlower-bound.\nAuthorized licensed use limited to: BEIJING UNIVERSITY OF POST AND TELECOM. Downloaded on November 07,2023 at 07:46:33 UTC from IEEE Xplore.  Restrictions apply. \n\n\n10920\nIEEE ROBOTICS AND AUTOMATION LETTERS, VOL. 7, NO. 4, OCTOBER 2022\nFig. 6.\nExperimental results of a robustness test conducted under various magnitudes of pose noise. AUG. means augmentation which adds pose noise during\ntraining.\nFig. 7.\nExperimental results of a bandwidth test conducted under various compression ratios.\nQualitative results: Fig. 5 shows the semantics segmentation\nresults. The results of upper-bound restore the semantic informa-\ntion with rich point cloud data. Intermediate-based collaboration\nstrategies V2VNet and DiscoNet can achieve satisfactory per-\nformance yet When2com and Who2com hurt the performance\ncompared to the lower-bound.\nD. Discussions on Pose Noise and Compression Ratio\nWe further examine the robustness of different intermediate\nmodels against realistic pose noise (Gaussian noise with a mean\nof 0.05 m-0.25 m and a standard deviation of 0.02 cm), as\nshown in Fig. 6. We can see that the models perform comparably\nwith or without pose noise in the training phase, and all the\nintermediate methods have shown stable performance against\nthe pose noise. The reason is that the intermediate feature map\nhas a relatively low spatial resolution (each grid in the feature\nmap denotes a coverage of 2m × 2m), thus is not vulnerable\nto noisy pose. Meanwhile, we employ an 1 × 1 autoencoder to\nfurther compress the feature channel of the transmitted feature\nmap. We test the models with different compression ratios, and\nwe found that the jointly learned 1 × 1 autoencoder can even\nimprove the performance a little bit, and most intermediate\nmodels achieve comparable performance at different levels of\ncompression, as shown in Fig. 7.\nV. CONCLUSION\nWe propose V2X-Sim dataset based on CARLA-SUMO\nco-simulation, in order to enable multi-agent collaborative per-\nception research in autonomous driving. Our dataset provides\nmulti-agent multi-modality sensor streams captured by the ve-\nhicles and road-side unit (RSU) in realistic trafﬁc ﬂows. Diverse\nannotations are provided to support a variety of 3D perception\ntasks. In addition, we benchmark several state-of-the-art col-\nlaborative perception methods in collaborative BEV detection,\ntracking, and semantic segmentation tasks. Future works include\nthe simulation of latency issues as well as the development of\nnovel evaluation metrics in collaborative perception tasks. We\nbelieve our work can inspire many relevant research areas in-\ncluding but not limited to autonomous driving, computer vision,\nmulti-robot system, communication engineering, and machine\nlearning.\nACKNOWLEDGMENT\nTheauthorswouldliketothankanonymousreviewersfortheir\nhelpful suggestions, and NYU high performance computing\n(HPC) for the support.\nREFERENCES\n[1] H. Caesar et al., “NuScenes: A multimodal dataset for autonomous driv-\ning,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2020, pp. 11618–\n11628.\n[2] A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving?\nthe KITTI vision benchmark suite,” in Proc. IEEE Conf. Comput. Vis.\nPattern Recognit., 2012, pp. 3354–3361.\n[3] P. Sun et al., “Scalability in perception for autonomous driving: Waymo\nopen dataset,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2020,\npp. 2443–2451.\n[4] E. Arnold, O. Y. Al-Jarrah, M. Dianati, S. Fallah, D. Oxtoby, and A.\nMouzakitis, “A survey on 3D object detection methods for autonomous\ndriving applications,” IEEE Trans. Intell. Transp. Syst., vol. 20, no. 10,\npp. 3782–3795, Oct. 2019.\nAuthorized licensed use limited to: BEIJING UNIVERSITY OF POST AND TELECOM. Downloaded on November 07,2023 at 07:46:33 UTC from IEEE Xplore.  Restrictions apply. \n\n\nLI et al.: V2X-SIM: MULTI-AGENT COLLABORATIVE PERCEPTION DATASET AND BENCHMARK FOR AUTONOMOUS DRIVING\n10921\n[5] S. M. Marvasti-Zadeh, L. Cheng, H. Ghanei-Yakhdan, and S. Kasaei,\n“Deep learning for visual tracking: A comprehensive survey,” IEEE Trans.\nIntell. Transp. Syst., vol. 23, no. 5, pp. 3943–3968, May 2022.\n[6] S. Minaee, Y. Boykov, F. Porikli, A. Plaza, N. Kehtarnavaz, and D.\nTerzopoulos, “Image segmentation using deep learning: A survey,” IEEE\nTrans.PatternAnal.Mach.Intelli.,vol.44,no.7,pp. 3523–3542,Jul.2022.\n[7] Z. MacHardy, A. Khan, K. Obana, and S. Iwashina, “V2X access technolo-\ngies: Regulation, research, and remaining challenges,” IEEE Commun.\nSurv. Tut., vol. 20, no. 3, pp. 1858–1877, Jul.–Sep. 2018.\n[8] M. Muhammad and G. A. Safdar, “Survey on existing authentication\nissues for cellular-assisted V2X communication,” Veh. Commun., vol. 12,\npp. 50–65, 2018.\n[9] M. Hasan, S. Mohan, T. Shimizu, and H. Lu, “Securing vehicle-to-\neverything (V2X) communication platforms,” IEEE Trans. Intell. Veh.,\nvol. 5, no. 4, pp. 693–713, Dec. 2020.\n[10] V. Mannoni, V. Berg, S. Sesia, and E. Perraud, “A comparison of the V2X\ncommunication systems: Its-g5 and C-V2X,” in Proc. IEEE 89th Veh.\nTechnol. Conf., 2019, pp. 1–5.\n[11] Q. Chen, S. Tang, Q. Yang, and S. Fu, “Cooper: Cooperative perception\nfor connected autonomous vehicles based on 3D point clouds,” in Proc.\nIEEE 39th Int. Conf. Distrib. Comput. Syst., 2019, pp. 514–524.\n[12] T.-H. Wang et al., “V2VNet: Vehicle-to-vehicle communication for joint\nperception and prediction,” in Proc. Eur. Conf. Comput. Vis., 2020,\npp. 605–621.\n[13] Y. Li, S. Ren, P. Wu, S. Chen, C. Feng, and W. Zhang, “Learning dis-\ntilled collaboration graph for multi-agent perception,” in Proc. Neural Inf.\nProcess. Syst., 2021, pp. 29541–29552.\n[14] Y. Yuan and M. Sester, “Comap: A synthetic dataset for collective multi-\nagent perception of autonomous driving,” Int. Arch. Photogrammetry,\nRemote Sens. Spatial Inf. Sci., vol. 43, pp. 255–263, 2021.\n[15] Y. Yuan, H. Cheng, and M. Sester, “Keypoints-based deep feature fusion\nfor cooperative vehicle detection of autonomous driving,” IEEE Robot.\nAutomat. Lett., vol. 7, no. 2, pp. 3054–3061, Apr. 2022.\n[16] D. Krajzewicz, J. Erdmann, M. Behrisch, and L. Bieker, “Recent devel-\nopment and applications of SUMO-simulation of urban mobility,” Int. J.\nAdv. Syst. Meas., vol. 5, no. 3/4, pp. 128–138, 2012.\n[17] A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun, “CARLA:\nAn open urban driving simulator,” in Proc. 1st Annu. Conf. Robot Learn.,\n2017, pp. 1–16.\n[18] Y.-C. Liu, J. Tian, N. Glaser, and Z. Kira, “When2com: Multi-agent\nperception via communication graph grouping,” in Proc. IEEE Conf.\nComput. Vis. Pattern Recognit., 2020, pp. 4105–4114.\n[19] Y.-C. Liu, J. Tian, C.-Y. Ma, N. Glaser, C.-W. Kuo, and Z. Kira,\n“Who2com: Collaborative perception via learnable handshake commu-\nnication,” in Proc. IEEE Int. Conf. Robot. Automat., 2020, pp. 6876–6883.\n[20] R. Xu, H. Xiang, X. Xia, X. Han, J. Liu, and J. Ma, “OPV2V: An open\nbenchmark dataset and fusion pipeline for perception with vehicle-to-\nvehicle communication,” in Proc. IEEE Int. Conf. Robot. Automat., 2022,\npp. 2583–2589.\n[21] Z. Xiao, Z. Mo, K. Jiang, and D. Yang, “Multimedia fusion at semantic\nlevel in vehicle cooperactive perception,” in Proc. IEEE Int. Conf. Multi-\nmedia Expo Workshops, 2018, pp. 1–6.\n[22] Y. Maalej, S. Sorour, A. Abdel-Rahim, and M. Guizani, “Vanets meet\nautonomous vehicles: A multimodal 3D environment learning approach,”\nin Proc. IEEE Glob. Commun. Conf., 2017, pp. 1–6.\n[23] E. Arnold, M. Dianati, R. de Temple, and S. Fallah, “Cooperative per-\nception for 3D object detection in driving scenarios using infrastructure\nsensors,” IEEE Trans. Intell. Transp. Syst., vol. 23, no. 3, pp. 1852–1864,\nMar. 2022.\n[24] M. Howe, I. Reid, and J. Mackenzie, “Weakly supervised training of\nmonocular 3D object detectors using wide baseline multi-view trafﬁc\ncamera data,” in Proc. Brit. Mach.Vis., 2021, p. 394.\n[25] H. Yu et al., “DAIR-V2X: A large-scale dataset for vehicle-infrastructure\ncooperative 3D object detection,” in Proc. IEEE/CVF Conf. Comput. Vis.\nPattern Recognit., 2022, pp. 21361–21370.\n[26] S. Manivasagam et al., “LiDArsim: Realistic LiDAR simulation by lever-\naging the real world,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit.,\n2020, pp. 11164–11173.\n[27] R. Xu, Y. Guo, X. Han, X. Xia, H. Xiang, and J. Ma, “OpenCDA: An open\ncooperative driving automation framework integrated with co-simulation,”\nin Proc. IEEE Int. Intell. Transp. Syst. Conf., 2021, pp. 1155–1162.\n[28] E. Arnold, S. Mozaffari, and M. Dianati, “Fast and robust registration of\npartially overlapping point clouds,” IEEE Robot. Automat. Lett., vol. 7,\nno. 2, pp. 1502–1509, Apr. 2022.\n[29] A. Patil, S. Malla, H. Gang, and Y.-T. Chen, “The H3D dataset for full-\nsurround 3D multi-object detection and tracking in crowded urban scenes,”\nin Proc. Int. Conf. Robot. Automat., 2019, pp. 9552–9557.\n[30] M. Pitropov et al., “Canadian adverse driving conditions dataset,” Int. J.\nRobot. Res., vol. 40, pp. 681–690, 2021.\n[31] Q.-H. Pham et al., “A*3D dataset: Towards autonomous driving in chal-\nlenging environments,” in Proc. IEEE Int. Conf. Robot. Automat., 2020,\npp. 2267–2273.\n[32] S. Yogamani et al., “Woodscape: A multi-task, multi-camera ﬁsheye\ndataset for autonomous driving,” in Proc. IEEE/CVF Int. Conf. Comput.\nVis., 2019, pp. 9307–9317.\n[33] X. Huang et al., “The apolloscape dataset for autonomous driving,” in\nProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, 2018,\npp. 1067–10676.\n[34] M. Cordts et al., “The cityscapes dataset for semantic urban scene un-\nderstanding,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2016,\npp. 3213–3223.\n[35] G. Ros, L. Sellart, J. Materzynska, D. Vázquez, and A. M. López, “The\nsynthia dataset: A large collection of synthetic images for semantic seg-\nmentation of urban scenes,” in Proc. IEEE Conf. Comput. Vis. Pattern\nRecognit., 2016, pp. 3234–3243.\n[36] G. Neuhold, T. Ollmann, S. R. Buló, and P. Kontschieder, “The mapillary\nvistas dataset for semantic understanding of street scenes,” in Proc. IEEE\nInt. Conf. Comput. Vis., 2017, pp. 5000–5009.\n[37] J. Behley et al., “SemanticKITTI: A dataset for semantic scene under-\nstanding of LiDAR sequences,” in Proc. IEEE/CVF Int. Conf. Comput.\nVis., 2019, pp. 9296–9306.\n[38] Q. Hu, B. Yang, S. Khalid, W. Xiao, A. Trigoni, and A. Markham, “To-\nwards semantic segmentation of urban-scale 3D point clouds: A dataset,\nbenchmarks and challenges,” in Proc. IEEE Conf. Comput. Vis. Pattern\nRecognit., 2021, pp. 4977–4987.\n[39] M.-F. Chang et al., “Argoverse: 3D tracking and forecasting with rich\nmaps,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2019,\npp. 8740–8749.\n[40] S. Ettinger et al., “Large scale interactive motion forecasting for au-\ntonomous driving: The waymo open motion dataset,” in Proc. IEEE/CVF\nInt. Conf. Comput. Vis., 2021, pp. 9710–9719.\n[41] D. Jia and D. Ngoduy, “Enhanced cooperative car-following trafﬁc model\nwith the combination of V2V and V2I communication,” Transp. Res. Part\nB-methodological, vol. 90, pp. 172–191, 2016.\n[42] S.-W. Kim et al., “Multivehicle cooperative driving using cooperative per-\nception: Design and experimental validation,” IEEE Trans. Intell. Transp.\nSyst., vol. 16, no. 2, pp. 663–680, Apr. 2015.\n[43] Z. Y. Rawashdeh and Z. Wang, “Collaborative automated driv-\ning: A machine learning-based method to enhance the accuracy of\nshared information,” in Proc. Int. Conf. Intell. Transp. Syst., 2018,\npp. 3961–3966.\n[44] Q. Chen et al., “DSRC and radar object matching for cooperative\ndriver assistance systems,” in Proc. IEEE Intell. Veh. Symp., 2015,\npp. 1348–1354.\n[45] S. Suo, S. Regalado, S. Casas, and R. Urtasun, “Trafﬁcsim: Learning\nto simulate realistic multi-agent behaviors,” in Proc. IEEE/CVF Conf.\nComput. Vis. Pattern Recognit., 2021, pp. 10400–10409.\n[46] S. Tan et al., “Scenegen: Learning to generate realistic trafﬁc\nscenes,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.,\n2021, pp. 892–901.\n[47] P. Wu, S. Chen, and D. Metaxas, “Motionnet: Joint perception and motion\nprediction for autonomous driving based on bird’s eye view maps,” in Proc.\nIEEE Conf. Comput. Vis. Pattern Recognit., 2020, pp. 11382–11392.\n[48] W. Luo, B. Yang, and R. Urtasun, “Fast and furious: Real time end-to-end\n3D detection, tracking and motion forecasting with a single convolu-\ntional net,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2018,\npp. 3569–3577.\n[49] J. Luiten et al., “HOTA: A higher order metric for evaluating multi-object\ntracking,” Int. J. Comput. Vis., vol. 129, no. 2, pp. 548–578, Oct. 2020.\n[50] K. Bernardin and R. Stiefelhagen, “Evaluating multiple object tracking\nperformance: The clear mot metrics,” EURASIP J. Image Video Process.,\npp. 1–10, 2008, Art. no. 246309.\n[51] A. Bewley, Z. Ge, L. Ott, F. Ramos, and B. Upcroft, “Simple online\nand realtime tracking,” in Proc. IEEE Int. Conf. Image Process., 2016,\npp. 3464–3468.\n[52] O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks\nfor biomedical image segmentation,” in Proc. Int. Med. Image Comput.\nComput.-Assist. Interv., 2015, pp. 234–241.\nAuthorized licensed use limited to: BEIJING UNIVERSITY OF POST AND TELECOM. Downloaded on November 07,2023 at 07:46:33 UTC from IEEE Xplore.  Restrictions apply.","difficulty":"hard","domain":"Multi-Document QA","length":"short","question":"What is the difference between the datasets of the two papers?","sub_domain":"Academic"}

Source: https://huggingface.co/datasets/zai-org/LongBench-v2

initial import

Posting: /agents

GET /api/v1/write?intent=publish&task_id=6acb208a-1e41-5ea1-ac7f-be286674edf5&body={url_encoded_text}&agent_name={optional_name}&nonce={optional_random_id}
