SentX Blog Meet Victoria

123D Gives Autonomous-Driving Datasets a Common Language — Here Is What It Establishes and Where It Stops

October 5, 2026 · 3 min read

123D is a data-infrastructure contribution rather than a new perception or planning algorithm. The problem it targets is familiar to anyone who has worked across autonomous-driving benchmarks: the field runs on a patchwork of datasets — nuScenes, Waymo Open Perception, Argoverse 2 Sensor, PandaSet, KITTI-360, Waymo Open Motion, nuPlan and PAI-AV — and each ships with its own file layouts, timestamping conventions, annotation schemas and calibration formats. Transferring a detector from one dataset to another can cost accuracy. 123D folds all of them into a single representation behind a common API, so downstream code reads one interface regardless of where the data came from. The paper reports integrating the eight real-world datasets plus a synthetic set, L3AD, recorded in CARLA simulation.

The core mechanism is straightforward once described. Every recording is converted into timestamped event streams: one stream per modality — cameras, with support for multiple projection models, lidar, ego-states, 3D bounding boxes, traffic-light states, and any user-defined type. Streams carry no prescribed sampling rate; each event keeps its own timestamp, and static metadata such as vehicle dimensions and sensor calibration lives in the file's schema rather than being repeated row by row. A separate synchronization table precomputes which rows across modalities refer to the same instant, with configurable strategies: preserve the original keyframes, choose a reference modality, or resample everything to a target rate. The access layer can bypass the table for arbitrary-timestamp reads, which is what lets the framework serve both synchronous and asynchronous consumers.

The paper demonstrates the approach with two worked applications: detection transfer and planning. In the detection study, standard detectors (PETR and BEVFormer-S, both with ResNet50 backbones) are trained individually on five selected sources — nuScenes, Waymo Open Perception, Argoverse 2 Sensor, nuPlan and CARLA — and on a uniform mixture of those same five, under a fixed 30,000-frame budget, so that diversity is isolated from raw volume; PandaSet, KITTI-360 and PAI-AV are excluded from training and used for held-out transfer evaluation, and storage constraints led the authors to use the nuPlan-mini and PAI-AV-NCore subsets. The evaluation is deliberately narrow — vehicle class only, within a 50-meter radius, scored with a modified NDS that drops velocity and attribute errors. The authors' own conclusion is measured: cross-domain and vehicle transfer remains challenging for both perception and behavior tasks, simple data mixing reduces the gap to some extent, and sim-to-real transfer proves no easier than cross-rig transfer between real datasets, even when the simulator replicates the target rig. The planning example, built on reinforcement learning, lands in the same place: behavior transfer stays difficult, with partial benefit from mixing. All reported numbers are the authors' measurements on their own harness, and this explainer is a source-based interpretation of the paper's inspected sections, not an independent reproduction of the results.

Three limitations are documented in the paper itself. First, coverage is limited to common sensor modalities; additional sensors such as radar and richer semantic annotations are named development targets, not delivered features. Second, there is no explicit support for web-scale datasets or continuous data streaming; the authors note that Arrow's cloud-access capabilities were only preliminarily tested. Third, the framework handles standard vehicles; trucks, mobile robots and other platforms are targeted for later expansion. The work is an author-submitted arXiv paper, and the description above reflects the sections available in that version.

For practitioners, the decision is therefore fairly clean. If you need to run existing detection or planning pipelines across several messy, heterogeneous driving datasets, 123D removes most of the glue code, and it preserves original annotations untouched rather than re-labeling them. If your task depends on radar, web-scale ingestion, live streaming, non-standard platforms, or certified accuracy, it sits outside the described framework and the question is simply unanswered by this paper. The honest bottom line is the paper's own: consolidation makes mixing data easy, but teaching models to transfer across domains remains the hard part. 123D shrinks the plumbing problem, not the learning problem.

Sources

  1. 123D: Unifying Multi-Modal Autonomous Driving Data at Scale — arXiv (author-submitted research)
Meet Victoria