AggPose: Deep Aggregation Vision Transformer for Infant Pose Estimation Paper • 2205.05277 • Published Aug 10, 2022 • 1
ChronoVision: Temporal Reasoning via Latent State Reconstruction Paper • 2608.05631 • Published 11 days ago • 40
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs Paper • 2406.17126 • Published Jun 24, 2024
PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation Paper • 2601.07060 • Published Jan 11 • 1
TrialBench: Multi-Modal Artificial Intelligence-Ready Clinical Trial Datasets Paper • 2407.00631 • Published Jun 15, 2025
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions Paper • 2605.15764 • Published May 15 • 4
Evaluating Cognitive Age Alignment in Interactive AI Agents Paper • 2605.17894 • Published May 18 • 5
CogniRoute: Learning to Route Social Evidence in Omni-Modal Models Paper • 2606.20970 • Published Jun 18 • 4
World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible Paper • 2606.13652 • Published Jun 11 • 16
ReMix: Reinforcement routing for mixtures of LoRAs in LLM finetuning Paper • 2603.10160 • Published Mar 10 • 26
Toward Cognitive Supersensing in Multimodal Large Language Model Paper • 2602.01541 • Published Feb 2 • 16
PhysRig: Differentiable Physics-Based Skinning and Rigging Framework for Realistic Articulated Object Modeling Paper • 2506.20936 • Published Jun 26, 2025 • 12
Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation Paper • 2509.10687 • Published Sep 12, 2025 • 7
RGB-Only Supervised Camera Parameter Optimization in Dynamic Scenes Paper • 2509.15123 • Published Sep 18, 2025 • 5