publications

publications in reversed chronological order

2026

  1. lychsim.jpg
    LychSim: A Controllable and Interactive Simulation Framework for Vision Research
    Wufei Ma , Chloe Wang , Siyi Chen, Jiawei Peng , Patrick Li, and Alan Yuille
    arXiv preprint, 2026
    Dataset 3D Vision Vision-Lanugage
  2. mssr.png
    Pursuing Minimal Sufficiency in Spatial Reasoning
    Yejie Guo, Yunzhong Hou, Wufei Ma, Meng Tang , and Ming-Hsuan Yang
    In The Fourteenth International Conference on Learning Representations , 2026
    3D Vision Vision-Lanugage
  3. animal4d.png
    4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
    Shanshan Zhong, Jiawei Peng, Zehan Zheng, Zhongzhan Huang, Wufei Ma, Guofeng Zhang, Qihao Liu, Alan Yuille, and Jieneng Chen
    In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2026
    3D Vision

2025

  1. gar.png
    Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
    Qihao Liu, Luoxin Ye, Wufei Ma, Yu-Cheng Chou, and Alan Yuille
    arXiv preprint, 2025
    LLM
  2. spatialreasoner.png
    SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
    Wufei Ma, Yu-Cheng Chou, Qihao Liu , Xingrui Wang, Celso M de Melo, Jianwen Xie, and Alan Yuille
    In Advances in Neural Information Processing Systems , 2025
    3D Vision Vision-Lanugage
  3. 3dsrbench.png
    3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
    Wufei Ma , Haoyu Chen, Guofeng Zhang, Jieneng Chen, Celso M de Melo, and Alan Yuille
    In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , Oct 2025
    3D Vision Vision-Lanugage
  4. dinemo.jpg
    DINeMo: Learning Neural Mesh Models with no 3D Annotations
    Weijie Guo, Guofeng Zhang, Wufei Ma, and Alan Yuille
    3rd Workshop on Compositional 3D Vision at CVPR, Oct 2025
    3D Vision
  5. 3dvlm-teaser2-2.jpg
    SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
    Wufei Ma, Luoxin Ye, Celso M de Melo, Jieneng Chen, and Alan Yuille
    In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , Oct 2025
    (Highlight, 3.0%)
    3D Vision Vision-Lanugage
  6. pulsecheck.png
    Spatial457: A Diagnostic Benchmark for Comprehensive Spatial Reasoning of Large Multimodal Models
    Xingrui Wang, Wufei Ma , Tiezheng Zhang, Celso M de Melo, Jieneng Chen, and Alan Yuille
    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , Oct 2025
    (Highlight, 3.0%)
    Dataset 3D Vision Vision-Lanugage
  7. superclevr_phy.png
    Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
    Xingrui Wang, Wufei Ma , Angtian Wang , Shuo Chen, Adam Kortylewski, and Alan Yuille
    In The Thirteenth International Conference on Learning Representations , Oct 2025
    Dataset 3D Vision Vision-Lanugage

2024

  1. imagenet3d_logo.png
    ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
    In Advances in Neural Information Processing Systems , Oct 2024
    Dataset 3D Vision
  2. feint6k.png
    Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
    In European Conference on Computer Vision , Oct 2024
    (Strong Double Blind)
    Dataset Vision-Lanugage
  3. rcnet.png
    NOVUM: Neural Object Volumes for Robust Object Classification
    In European Conference on Computer Vision , Oct 2024
    3D Vision
  4. tmm_eg.png
    Uncertainty-Aware Deep Video Compression with Ensembles
    Wufei Ma , Jiahao Li , Bin Li, and Yan Lu
    IEEE Transactions on Multimedia, Oct 2024
    Compression
  5. 3ddst_compressed.png
    Generating Images with 3D Annotations Using Diffusion Models
    Wufei Ma*, Qihao Liu* , Jiahao Wang* , Angtian Wang, Xiaoding Yuan , Yi Zhang, Zihao Xiao, Guofeng Zhang , Beijia Lu, Ruxiao Duan, Yongrui Qi, Adam Kortylewski , Yaoyao Liu, and Alan Yuille
    In The Twelfth International Conference on Learning Representations , Oct 2024
    (Spotlight, 5%)
    Dataset 3D Vision
  6. deformable_nemo.png
    Neural textured deformable meshes for robust analysis-by-synthesis
    In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , Oct 2024
    3D Vision
  7. synthetic_nemo.png
    Robust Category-Level 3D Pose Estimation from Diffusion-Enhanced Synthetic Data
    In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , Oct 2024
    3D Vision
  8. oodcv_v2.png
    OOD-CV-v2: An Extended Benchmark for Robustness to Out-of-Distribution Shifts of Individual Nuisances in Natural Images
    Bingchen Zhao , Jiahao Wang, Wufei Ma, Artur Jesslen , Siwei Yang, Shaozuo Yu, Oliver Zendel, Christian Theobalt, Alan Yuille, and Adam Kortylewski
    IEEE Transactions on Pattern Analysis and Machine Intelligence, Oct 2024
    Dataset 3D Vision

2023

  1. 3dvqa.png
    3d-aware visual question answering about parts, poses and occlusions
    In Advances in Neural Information Processing Systems , Oct 2023
    Dataset 3D Vision Vision-Lanugage
  2. animal3d_sm.png
    Animal3D: A Comprehensive Dataset of 3D Animal Pose and Shape
    Jiacong Xu , Yi Zhang, Jiawei Peng, Wufei Ma, Artur Jesslen, Pengliang Ji, Qixin Hu , Jiehua Zhang, Qihao Liu , Jiahao Wang, and  others
    In Proceedings of the IEEE/CVF International Conference on Computer Vision , Oct 2023
    Dataset 3D Vision
  3. superclevr_eg.png
    Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning
    Zhuowan Li , Xingrui Wang, Elias Stengel-Eskin, Adam Kortylewski, Wufei Ma, Benjamin Van Durme, and Alan Yuille
    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , Oct 2023
    (Highlight)
    Dataset Vision-Lanugage

2022

  1. eccv22_6dpose_small.png
    Robust Category-Level 6D Pose Estimation with Coarse-to-Fine Rendering of Neural Features
    In European Conference on Computer Vision , Oct 2022
    3D Vision
  2. robin2.png
    OOD-CV: A Benchmark for Robustness to Out-of-Distribution Shifts of Individual Nuisances in Natural Images
    Bingchen Zhao, Shaozuo Yu, Wufei Ma , Mingxin Yu, Shenxiao Mei , Angtian Wang, Ju He, Alan Yuille, and Adam Kortylewski
    In European conference on computer vision , Oct 2022
    (Oral)
    Dataset 3D Vision

2021

  1. pub-3.png
    Making group decisions from natural language-based preferences
    Farhad Mohsin, Lei Luo, Wufei Ma, Inwon Kang , Zhibing Zhao , Ao Liu, Rohit Vaish, and Lirong Xia
    In Proceedings of the 8th International Workshop on Computational Social Choice (COMSOC) , Oct 2021
    Dataset Preference Learning
  2. material_perspective.png
    Adoption of Image-Driven Machine Learning for Microstructure Characterization and Materials Design: A Perspective
    Arun Baskaran, Elizabeth J Kautz, Aritra Chowdhary, Wufei Ma, Bulent Yener, and Daniel Lewis
    JOM, Oct 2021
    Microstructure

2020

  1. pub-2.png
    Image-driven discriminative and generative machine learning algorithms for establishing microstructure–processing relationships
    Wufei Ma, Elizabeth J Kautz, Arun Baskaran, Aritra Chowdhury, Vineet Joshi, Bulent Yener, and Daniel Lewis
    Journal of Applied Physics, Oct 2020
    Microstructure
  2. pub-1.png
    An image-driven machine learning approach to kinetic modeling of a discontinuous precipitation reaction
    Elizabeth Kautz, Wufei Ma, Saumyadeep Jana, Arun Devaraj, Vineet Joshi, Bülent Yener, and Daniel Lewis
    Materials Characterization, Oct 2020
    Microstructure