Obj Variant Ensemble: Advancing Point Cloud LLM Evaluation in Challenging Scenes with Subtly Distinguished Object

Qihang Cao1,2, Huangxun Chen2
1. Shanghai Jiao Tong University
2. Hong Kong University of Science and Technology (Guangzhou)

Comparison of 3D grounding benchmarks in complex scenes:

[Left] A scene from ScanNet/ScanRef, where the text description is insufficient for accurately locating a chair.

[Right] A scene from ObjVariantEnsemble, where a model successfully identifies targets given sufficiently detailed descriptions.

Problems of Current Datasets & Advantages of OVE

Problems with Existing Benchmarks

  • Small scale and insufficient annotations.
  • Lack of customizable challenge levels.

Larger & More Diverse Data

OVE significantly outperforms existing benchmarks in terms of data volume and annotation richness.

Dataset Comparison
Comparison of dataset size and annotations: OVE vs. existing benchmarks.

More Flexible Evaluation

OVE enables difficulty adjustments based on object similarity, spatial complexity, and scene variation.

Benchmark Summary
Four challenge levels: Loc, Loc+Shape, Loc+Color, and Loc+Class.

Abstract

3D scene understanding is an important task, and there has been a recent surge of research interest in aligning 3D representations of point clouds with text to empower embodied AI. However, due to the lack of comprehensive 3D benchmarks, the capabilities of 3D models in real-world scenes, particularly those that are challenging with subtly distinguished objects, remain insufficiently investigated. To facilitate a more thorough evaluation of 3D models’ capabilities, we propose a scheme, ObjVariantEnsemble, to systematically introduce more scenes with specified object classes, colors, shapes, quantities, and spatial relationships to meet model evaluation needs. More importantly, we intentionally construct scenes with similar objects to a certain degree and design an LLM-VLM-cooperated annotator to capture key distinctions as annotations. The resultant benchmark can better challenge 3D models, reveal their shortcomings in understanding, and potentially aid in the further development of 3D models.

Poster

Citation

If you use our dataset or code in your research, please cite our work using the following BibTeX entry:

@article{cao2024objvariantensemble,
  title={ObjVariantEnsemble: Advancing Point Cloud LLM Evaluation in Challenging Scenes with Subtly Distinguished Objects},
  author={Cao, Qihang and Chen, Huangxun},
  journal={arXiv preprint arXiv:2412.14837},
  year={2024}
}