SJTU_GVI: a GNSS-Visual-IMU rosbag dataset for ground-vehicle SLAM
A GNSS-Visual-IMU Dataset for SLAM
At a glance
- What is it?
- SJTU_GVI is a five-sequence rosbag benchmark recorded on a car with a Ublox ZED-F9P GNSS receiver, cameras and an IMU. It is small enough to download, but the sequences are short and the ground truth is a set of files rather than a finished evaluation pipeline.
- Who is it for?
- Adopt SJTU_GVI if you are working on GNSS-visual-inertial odometry for ground vehicles and want a small rosbag set with raw ZED-F9P measurements to sanity-check a pipeline before moving to a larger benchmark. Skip it if you need long trajectories, many scenarios, or a maintained evaluation harness: five sequences of 109 to 349 seconds will not separate methods that only diverge over kilometres.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Probably not. The repository last received commits 22 months ago, on November 22, 2024.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What SJTU_GVI is, and the gap it fills
Most public SLAM datasets force a choice. The large multi-sensor sets give you variety but take hours to download and days to preprocess. The small ones are usually visual-inertial only, which means a GNSS-fused estimator has nothing to fuse. SJTU_GVI sits in the middle: it is a GNSS-Visual-IMU benchmark dataset for SLAM, published alongside the M2C-GVIO paper in Satellite Navigation, and the README positions it explicitly against M2DGR, the authors' earlier ground-robot dataset. The stated differences are that SJTU_GVI is more light-weight and easier to download, was recorded on a real car at high speed, and captures GNSS raw measurements from a Ublox ZED-F9P receiver.
The intended audience is narrow and specific. If you are testing GNSS-visual-inertial odometry, or a GNSS-SLAM method that consumes raw pseudorange and carrier-phase style measurements rather than a fused position fix, this dataset gives you the sensor combination without the download burden. If your work is purely visual or purely LiDAR, the GNSS stream adds nothing you need. The repository itself is thin: three top-level entries (README.md, calibrations/, data/), no code, no evaluation scripts. It is a data drop, not a toolkit.
The five sequences and what is actually in them
All five sequences were collected on the same day, 2021-12-23, which is the most important constraint in the dataset. Seq1 is 2.37G and 349 seconds; Seq2 is 921M and 109 seconds; Seq3 is 1.19G and 146 seconds; Seq4 is 850M and 112 seconds; Seq5 is 1.26G and 159 seconds. Total download is roughly 6.6 GB, which is the whole point of the project: a full set fits on a laptop drive and can be pulled over a normal connection.
The trade-off is coverage. Every sequence is under six minutes, and because they share a collection date they share weather, lighting and roughly the same urban environment. That is enough to check that a GNSS-visual-inertial estimator initialises, tracks and does not diverge, and enough to reproduce the numbers in the M2C-GVIO paper. It is not enough to make claims about robustness across seasons, cities or sensor configurations. The README also notes the platform was a real car at high speed, which matters for motion manifold work: the dynamics are vehicle-like, not pedestrian-like, and an estimator tuned for handheld motion may behave differently here.
Getting the rosbags: Baidu Pan and the extraction code
There is no package manager, no release tarball and no direct HTTP download. Each sequence is a link to Baidu Pan, and the README gives a single extraction code that applies to the set: "yj66". The repository homepage field is empty, so the GitHub page is the only entry point.
In practice the workflow is: open the sequence row you want, follow the Rosbag link, and enter the code when Baidu Pan asks for it.
# no install command is documented; the data is fetched by hand
# extraction code for the Baidu Pan links, per the README:
yj66Expect a rosbag file once the download completes. The README does not list the topics inside the bags, so before writing any launch file, inspect the bag and confirm which topics carry the camera images, the IMU measurements and the GNSS raw measurements. Two other pieces are in the repository rather than the bags: camera intrinsic and camera-IMU extrinsic calibration results under calibrations/, and the ground truth under data/. Both are plain files in the tree, so clone the repository to get them.
git clone https://github.com/sjtuyinjie/SJTU_GVI.git
cd SJTU_GVI
ls calibrations dataThe listing should show the calibration files and the ground truth files described in the README. If it does not, you are on a different branch or the clone failed.
Ground truth and calibration are files, not a pipeline
The README states that camera intrinsics and camera-IMU extrinsics are provided under calibrations/, and that the ground truth is provided under data/. That is the entire description. There is no documented format for the ground truth, no stated reference source, no accuracy figure, and no script that aligns an estimated trajectory to it. Anyone who has used a benchmark with an official evaluation tool will notice the absence immediately: you will be writing your own ATE or RPE script, and you will be deciding for yourself how to interpolate the ground truth onto your estimate's timestamps.
This is a real cost, not a cosmetic one. Trajectory alignment choices (Sim(3) versus SE(3), timestamp association windows) can move ATE numbers by amounts comparable to the differences between methods. On a dataset with published baselines, the harness pins those choices down. Here nothing does, so cross-paper comparisons on SJTU_GVI should be read with that in mind. The upside is that the ground truth files are in the repository where you can inspect them directly rather than behind a download form.
Where SJTU_GVI is the wrong choice
Three cases stand out. First, long-horizon drift evaluation. A 349-second maximum trajectory cannot expose the slow scale drift or GNSS outage recovery behaviour that motivates GNSS fusion in the first place; the sequences are too short for the failure to develop. Second, multi-scenario generalisation studies. All five sequences come from one day and one platform, so a method that wins here has demonstrated very little about robustness. Third, anything requiring LiDAR or a richer sensor suite. The name says GNSS, visual and IMU, and that is what the README describes.
There is also a practical failure mode worth naming: the distribution channel. Baidu Pan links are awkward outside mainland China, and a link that works today may require an account or a client tomorrow. The README offers no mirror. If your team cannot reliably pull from Baidu Pan, the dataset's main selling point, being light-weight and easy to download, does not apply to you.
M2DGR and other ground-vehicle SLAM benchmarks
The README itself frames the comparison: SJTU_GVI is presented as different from M2DGR, the authors' earlier multi-sensor, multi-scenario SLAM dataset for ground robots, in being more light-weight and easier to download, recorded at high speed on a car, and carrying Ublox ZED-F9P GNSS raw measurements. The difference in approach is one of scope versus convenience. M2DGR targets breadth: more sensors and more scenarios, at the cost of a much heavier download. SJTU_GVI targets a single well-instrumented platform and a small footprint.
That makes the two complementary rather than competing. A common pattern would be to develop and debug a GNSS-visual-inertial pipeline on SJTU_GVI because iteration is cheap, then validate on the heavier set before claiming generalisation. What SJTU_GVI does not offer, and M2DGR is described as offering, is scenario diversity. Note also that the README asks users who employ M2DGR in academic work to cite the M2DGR RA-L paper and the M2C-GVIO paper, so the citation obligations are stated for both.
Licence, maintenance and what to verify before you build on it
The repository does not state a licence, and no releases are listed. That is a gap you should resolve before using the data in anything you intend to distribute: without an explicit licence, the terms under which you may redistribute the rosbags, calibration files or ground truth are simply not stated, and the README's citation request is not a substitute for a licence grant. This is a factual observation about the repository, not legal advice; if redistribution matters to your project, ask the authors directly.
Maintenance is also undocumented. No release history is given, and the README describes a fixed set of five sequences collected on 2021-12-23 with no stated plan to extend it. Treat SJTU_GVI as a static artifact tied to the M2C-GVIO paper rather than something that will grow. The upgrade cost is therefore near zero in one sense (nothing to upgrade) and non-trivial in another: if you need more sequences or a different platform, you are looking at a different dataset, not a newer version of this one. The concrete first checks are whether the Baidu Pan links still resolve, whether the calibrations/ and data/ directories contain what the README says, and whether the bag topics line up with your estimator's expected inputs.
Editorial conclusion
Adopt SJTU_GVI if you are working on GNSS-visual-inertial odometry for ground vehicles and want a small rosbag set with raw ZED-F9P measurements to sanity-check a pipeline before moving to a larger benchmark. Skip it if you need long trajectories, many scenarios, or a maintained evaluation harness: five sequences of 109 to 349 seconds will not separate methods that only diverge over kilometres. Before committing, download one rosbag, confirm the extraction code works, and check that the topics inside match the sensors your estimator expects, because the README lists sensors but not topic names.
Frequently asked questions
What sensors does the SJTU_GVI dataset include?
The name and README describe GNSS, visual and IMU data, with GNSS raw measurements captured by a Ublox ZED-F9P receiver. Camera intrinsics and camera-IMU extrinsics are provided separately under calibrations/.
How do I download the SJTU_GVI rosbags?
Each of the five sequences links to Baidu Pan, and the README gives the extraction code "yj66" for the links. There is no package manager or direct HTTP download documented.
Where is the ground truth for SJTU_GVI?
The README states the ground truth is provided in the data/ directory of the repository, alongside calibration results in calibrations/. No format or evaluation script is documented.
How long are the SJTU_GVI sequences?
Durations range from 109 seconds for Seq2 to 349 seconds for Seq1, with sizes from 850M to 2.37G. All five were collected on 2021-12-23.