Dataset Viewer
Auto-converted to Parquet Duplicate
Search is not available for this dataset
image
imagewidth (px)
518
518
label
class label
2 classes
0gt
0gt
0gt
0gt
0gt
0gt
1input
1input
1input
1input
1input
1input
1input
1input
1input
0gt
0gt
0gt
0gt
0gt
0gt
0gt
1input
1input
1input
1input
1input
1input
1input
1input
1input
1input
0gt
0gt
0gt
0gt
0gt
1input
1input
1input
1input
1input
1input
1input
1input
0gt
0gt
0gt
0gt
0gt
0gt
0gt
1input
1input
1input
1input
1input
1input
1input
1input
1input
1input
0gt
0gt
0gt
0gt
0gt
0gt
1input
1input
1input
1input
1input
1input
1input
1input
1input
0gt
0gt
0gt
0gt
0gt
0gt
1input
1input
1input
1input
1input
1input
1input
1input
1input
0gt
0gt
0gt
0gt
0gt
0gt
0gt
1input
End of preview. Expand in Data Studio

Mira-Scene Dataset

This repository contains the public data used by the Mira-Scene single-image 3D scene reconstruction project: the BlendSwap evaluation benchmark and the training releases derived from Objaverse Outpaint and 3D-FRONT. The benchmark and training data are published in the same Hugging Face dataset repository.

The training data is distributed as verified tar.gz shards. The archives retain the relative paths expected by the Mira-Scene training loaders. The dataset does not grant rights beyond the licenses and terms of the source datasets; please acknowledge Objaverse, 3D-FRONT, and 3D-FUTURE when using the training data.

Repository layout

blendswap_eval/       # evaluation scenes
objaverse_outpaint/   # Objaverse Outpaint training shards and indexes
3dfront/              # 3D-FRONT training shards and indexes
manifests/            # public checksums and release indexes

Evaluation benchmark

The BlendSwap benchmark is organized as one directory per scene:

blendswap_eval/
└── <scene-name>/
    β”œβ”€β”€ input/
    β”‚   β”œβ”€β”€ scene.png
    β”‚   β”œβ”€β”€ scene_fg.png
    β”‚   β”œβ”€β”€ mask_*.png
    β”‚   └── floor_mask.png
    β”œβ”€β”€ depth/
    β”‚   └── ...
    └── gt/
        β”œβ”€β”€ camera.json
        β”œβ”€β”€ scene_camera.glb
        β”œβ”€β”€ canonical_coord_map_*.npy
        β”œβ”€β”€ canonical_coord_map_*.png
        β”œβ”€β”€ canonical_pcd_*.ply
        β”œβ”€β”€ voxel_*.npy
        β”œβ”€β”€ mesh_*.glb
        └── ...

input/ contains the input image and instance masks. depth/ contains depth artifacts used by the evaluation workflow. gt/ contains per-object canonical-coordinate maps, point clouds, voxel annotations, meshes, and camera metadata.

The current BlendSwap snapshot contains 15 scenes and is approximately 1.46 GiB. From a Mira-Scene checkout, run:

python eval_scripts/run_eval.py \
  --data-dir /path/to/mira-scene-data/blendswap_eval \
  --config infer_scripts/config/local.yaml \
  --gpu-ids 0

For multi-GPU evaluation, pass a comma-separated list to --gpu-ids.

Training data

Subset Views Objects / scenes Shards Compressed Extracted
Objaverse Outpaint 109,944 42,972 meshes 41 129.4 GB 217.1 GB
3D-FRONT 40,043 9,481 scenes 75 268.0 GB 388.8 GB
Total 149,987 β€” 116 397.4 GB 605.9 GB

The release contains 550,962 files. Reserve about 1 TB if compressed shards and extracted data must be kept at the same time.

Training data layout

The archives are split into the following public directories:

objaverse_outpaint/
β”œβ”€β”€ summary.json
β”œβ”€β”€ metadata.jsonl.gz
└── shards/part-*.tar.gz

3dfront/
β”œβ”€β”€ preprocess_train.json
└── shards/part-*.tar.gz

After extraction, objaverse_outpaint/summary.json is the index consumed by ObjaverseSceneDepthDataset. 3dfront/preprocess_train.json is the scene index consumed by the 3D-FRONT loader. The shard archives contain the images, masks, depth, poses, HDF5 renderings, and mesh files required by training.

The manifests/ directory is not required for normal training. It contains public integrity metadata: per-member SHA256 records in manifests/files/, per-shard records in manifests/shards/, and the aggregate manifests/shards.json. It does not contain private filesystem paths, credentials, or internal verification reports.

Download

Download only the subset needed for a run, or download both subsets into the same directory:

hf download Yang-Tian/Mira-Scene-Dataset \
  --repo-type dataset \
  --local-dir ./mira-scene-data \
  --include "objaverse_outpaint/*"
hf download Yang-Tian/Mira-Scene-Dataset \
  --repo-type dataset \
  --local-dir ./mira-scene-data \
  --include "3dfront/*"

To download the optional integrity metadata:

hf download Yang-Tian/Mira-Scene-Dataset \
  --repo-type dataset \
  --local-dir ./mira-scene-data \
  --include "manifests/*"

The archive upload is split into resumable commits. If you publish a local copy, use the repository helper's training target; it reports progress every 60 seconds and resumes through Hugging Face's .cache/.huggingface/ metadata:

python scripts_bak2/upload_hf.py training \
  --training-dir /path/to/mira_scene_hf_release_v1/release \
  --workers 4

The download contains compressed archives. Extract each subset into the same dataset root, preserving the archive paths:

cd ./mira-scene-data
for archive in objaverse_outpaint/shards/*.tar.gz; do
  tar -xzf "$archive"
done
for archive in 3dfront/shards/*.tar.gz; do
  tar -xzf "$archive"
done

Do not place the archives under an additional nested directory. The resulting dataset root should contain objaverse_outpaint/ and 3dfront/ directly.

Training configuration

Install Mira-Scene and UniDataset, then follow the training instructions in example_train/README.md. Set the public training configuration's dataset root to the directory containing objaverse_outpaint/ and 3dfront/. The published indexes can be used directly; no index conversion or preprocessing step is required.

The release includes an Objaverse loader index and public shard checksums for download verification. The manifests are optional and are not read by the training loader.

Source datasets and citation

The training data derives from Objaverse, 3D-FRONT, and 3D-FUTURE. Please follow each source dataset's applicable license, terms of use, and citation requirements.

License and source terms

This dataset combines Mira-Scene-generated annotations and rendered data with assets derived from Objaverse, 3D-FRONT, 3D-FUTURE, and BlendSwap. Each source dataset and asset retains its own license and terms of use. The MIT License for the Mira-Scene code does not apply to the dataset contents. Users are responsible for checking and complying with the terms of the relevant source datasets.

Downloads last month
1,341

Papers for Yang-Tian/Mira-Scene-Dataset