Datasets:
image imagewidth (px) 518 518 | label class label 2
classes |
|---|---|
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
1input | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
0gt | |
1input |
Mira-Scene Dataset
This repository contains the public data used by the Mira-Scene single-image 3D scene reconstruction project: the BlendSwap evaluation benchmark and the training releases derived from Objaverse Outpaint and 3D-FRONT. The benchmark and training data are published in the same Hugging Face dataset repository.
The training data is distributed as verified tar.gz shards. The archives
retain the relative paths expected by the Mira-Scene training loaders. The
dataset does not grant rights beyond the licenses and terms of the source
datasets; please acknowledge Objaverse, 3D-FRONT, and 3D-FUTURE when using the
training data.
Repository layout
blendswap_eval/ # evaluation scenes
objaverse_outpaint/ # Objaverse Outpaint training shards and indexes
3dfront/ # 3D-FRONT training shards and indexes
manifests/ # public checksums and release indexes
Evaluation benchmark
The BlendSwap benchmark is organized as one directory per scene:
blendswap_eval/
βββ <scene-name>/
βββ input/
β βββ scene.png
β βββ scene_fg.png
β βββ mask_*.png
β βββ floor_mask.png
βββ depth/
β βββ ...
βββ gt/
βββ camera.json
βββ scene_camera.glb
βββ canonical_coord_map_*.npy
βββ canonical_coord_map_*.png
βββ canonical_pcd_*.ply
βββ voxel_*.npy
βββ mesh_*.glb
βββ ...
input/ contains the input image and instance masks. depth/ contains depth
artifacts used by the evaluation workflow. gt/ contains per-object
canonical-coordinate maps, point clouds, voxel annotations, meshes, and camera
metadata.
The current BlendSwap snapshot contains 15 scenes and is approximately 1.46 GiB. From a Mira-Scene checkout, run:
python eval_scripts/run_eval.py \
--data-dir /path/to/mira-scene-data/blendswap_eval \
--config infer_scripts/config/local.yaml \
--gpu-ids 0
For multi-GPU evaluation, pass a comma-separated list to --gpu-ids.
Training data
| Subset | Views | Objects / scenes | Shards | Compressed | Extracted |
|---|---|---|---|---|---|
| Objaverse Outpaint | 109,944 | 42,972 meshes | 41 | 129.4 GB | 217.1 GB |
| 3D-FRONT | 40,043 | 9,481 scenes | 75 | 268.0 GB | 388.8 GB |
| Total | 149,987 | β | 116 | 397.4 GB | 605.9 GB |
The release contains 550,962 files. Reserve about 1 TB if compressed shards and extracted data must be kept at the same time.
Training data layout
The archives are split into the following public directories:
objaverse_outpaint/
βββ summary.json
βββ metadata.jsonl.gz
βββ shards/part-*.tar.gz
3dfront/
βββ preprocess_train.json
βββ shards/part-*.tar.gz
After extraction, objaverse_outpaint/summary.json is the index consumed by
ObjaverseSceneDepthDataset. 3dfront/preprocess_train.json is the scene
index consumed by the 3D-FRONT loader. The shard archives contain the images,
masks, depth, poses, HDF5 renderings, and mesh files required by training.
The manifests/ directory is not required for normal training. It contains
public integrity metadata: per-member SHA256 records in
manifests/files/, per-shard records in manifests/shards/, and the aggregate
manifests/shards.json. It does not contain private filesystem paths,
credentials, or internal verification reports.
Download
Download only the subset needed for a run, or download both subsets into the same directory:
hf download Yang-Tian/Mira-Scene-Dataset \
--repo-type dataset \
--local-dir ./mira-scene-data \
--include "objaverse_outpaint/*"
hf download Yang-Tian/Mira-Scene-Dataset \
--repo-type dataset \
--local-dir ./mira-scene-data \
--include "3dfront/*"
To download the optional integrity metadata:
hf download Yang-Tian/Mira-Scene-Dataset \
--repo-type dataset \
--local-dir ./mira-scene-data \
--include "manifests/*"
The archive upload is split into resumable commits. If you publish a local
copy, use the repository helper's training target; it reports progress every
60 seconds and resumes through Hugging Face's .cache/.huggingface/ metadata:
python scripts_bak2/upload_hf.py training \
--training-dir /path/to/mira_scene_hf_release_v1/release \
--workers 4
The download contains compressed archives. Extract each subset into the same dataset root, preserving the archive paths:
cd ./mira-scene-data
for archive in objaverse_outpaint/shards/*.tar.gz; do
tar -xzf "$archive"
done
for archive in 3dfront/shards/*.tar.gz; do
tar -xzf "$archive"
done
Do not place the archives under an additional nested directory. The resulting
dataset root should contain objaverse_outpaint/ and 3dfront/ directly.
Training configuration
Install Mira-Scene and UniDataset, then follow the training instructions in
example_train/README.md.
Set the public training configuration's dataset root to the directory
containing objaverse_outpaint/ and 3dfront/. The published indexes can be
used directly; no index conversion or preprocessing step is required.
The release includes an Objaverse loader index and public shard checksums for download verification. The manifests are optional and are not read by the training loader.
Source datasets and citation
The training data derives from Objaverse, 3D-FRONT, and 3D-FUTURE. Please follow each source dataset's applicable license, terms of use, and citation requirements.
License and source terms
This dataset combines Mira-Scene-generated annotations and rendered data with assets derived from Objaverse, 3D-FRONT, 3D-FUTURE, and BlendSwap. Each source dataset and asset retains its own license and terms of use. The MIT License for the Mira-Scene code does not apply to the dataset contents. Users are responsible for checking and complying with the terms of the relevant source datasets.
- Downloads last month
- 1,341