DeepRead
Search
Search
Dark mode
Light mode
Explorer
Popular Tags
#VLM
#agentic-RL
#gui-agent
#web-agent
#LLM
#VLA
#computer-use
#task-planning
#manipulation
#imitation-learning
Tag: scene-understanding
85 items with this tag.
Aug 20, 2026
Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models
spatial-reasoning
VLM
scene-understanding
Aug 20, 2026
S2-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
spatial-reasoning
VLM
scene-understanding
Aug 20, 2026
DENALI: A Dataset Enabling Non-Line-of-Sight Spatial Reasoning with Low-Cost LiDARs
spatial-reasoning
scene-understanding
3D-representation
Aug 20, 2026
Curvature-Aware Captioning: Leveraging Geodesic Attention for 3D Scene Understanding
scene-understanding
spatial-reasoning
3D-representation
Aug 20, 2026
Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation
navigation
spatial-reasoning
scene-understanding
Aug 20, 2026
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
VLM
scene-understanding
Aug 20, 2026
Gemini Robotics ER 1.6: Enhanced Embodied Reasoning
embodied-reasoning
spatial-reasoning
scene-understanding
Aug 20, 2026
PV-Ground: Text-Guided Point-Voxel Interaction for 3D Visual Grounding
scene-understanding
3D-representation
spatial-reasoning
Aug 20, 2026
OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding
scene-understanding
3D-representation
semantic-map
Aug 20, 2026
AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object Detection
VLM
scene-understanding
LLM
Aug 20, 2026
MonoVLM: Monocular 3D Visual Grounding with Vision Language Models
VLM
spatial-reasoning
scene-understanding
Aug 20, 2026
Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
spatial-reasoning
scene-understanding
LLM
Aug 20, 2026
Affordance2Action: Task-Conditioned Scene-level Affordance Grounding for Real-Time Manipulation
manipulation
scene-understanding
spatial-reasoning
VLA
Aug 20, 2026
Lifting Unlabeled Internet-level Data for 3D Scene Understanding
scene-understanding
spatial-reasoning
VLN
Aug 20, 2026
VLM4RSDet: Collaborative Optimization with Vision-Language Model for Enhancing Remote Sensing Object Detection
VLM
scene-understanding
Aug 20, 2026
Unsupervised Multi-agent and Single-agent Perception from Cooperative Views
scene-understanding
3D-representation
Aug 20, 2026
StateVLM: A State-Aware Vision-Language Model for Robotic Affordance Reasoning
VLM
scene-understanding
embodied-reasoning
Aug 20, 2026
UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language Conditions
spatial-reasoning
scene-understanding
VLM
Aug 20, 2026
InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
spatial-reasoning
VLM
scene-understanding
OB
CN
OB/CN
Aug 20, 2026
Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Clouds
scene-understanding
3D-representation
spatial-reasoning
Aug 20, 2026
SyncMos: Scalable Motion Synchronisation for Multi-Agent Scene Interaction
embodied-reasoning
task-planning
scene-understanding
Aug 20, 2026
Hear you are: Teaching LLMs Spatial Reasoning with Vision and Spatial Sound
spatial-reasoning
VLM
scene-understanding
Aug 20, 2026
Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
3D-representation
scene-understanding
spatial-reasoning
Aug 20, 2026
GeoDiT: A Diffusion-based Vision-Language Model for Geospatial Understanding
VLM
spatial-reasoning
scene-understanding
Aug 20, 2026
First Logit Boosting: Visual Grounding Method to Mitigate Object Hallucination in Large Vision-Language Models
VLM
scene-understanding
Aug 20, 2026
EG-3DVG: Expression and Geometry Aware Grounding Decoder for 3D Visual Grounding
scene-understanding
spatial-reasoning
VLM
Aug 20, 2026
VLMDiff: Leveraging Vision-Language Models for Multi-Class Anomaly Detection with Diffusion
VLM
scene-understanding
Aug 20, 2026
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
VLN
spatial-memory
scene-understanding
Img
Aug 20, 2026
Empirical Evidence on Conversational Control of GUI in Semantic Automation
manipulation
scene-understanding
Aug 20, 2026
Geo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial Imagery
VLM
spatial-reasoning
scene-understanding
Aug 20, 2026
UIPro: Unleashing Superior Interaction Capability For GUI Agents
VLM
imitation-learning
scene-understanding
Aug 20, 2026
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
navigation
scene-understanding
spatial-memory
Aug 20, 2026
UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning
imitation-learning
RL
scene-understanding
Aug 20, 2026
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
VLM
scene-understanding
Aug 20, 2026
TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
VLM
navigation
scene-understanding
Aug 20, 2026
ScaleTrack: Scaling and back-tracking Automated GUI Agents
imitation-learning
scene-understanding
web-agent
Aug 20, 2026
OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics
mobile-manipulation
scene-understanding
manipulation
1-2
Aug 20, 2026
Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent
imitation-learning
RL
scene-understanding
Aug 20, 2026
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
navigation
imitation-learning
scene-understanding
Aug 20, 2026
MobileUse: A GUI Agent with Hierarchical Reflection for Autonomous Mobile Operation
navigation
imitation-learning
scene-understanding
Aug 20, 2026
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
VLM
scene-understanding
spatial-reasoning
Aug 20, 2026
Mobile-Agent-v3: Fundamental Agents for GUI Automation
RL
scene-understanding
web-agent
Aug 20, 2026
WebCanvas: Benchmarking Web Agents in Online Environments
imitation-learning
RL
scene-understanding
Aug 20, 2026
MEGA-GUI: Multi-stage Enhanced Grounding Agents for GUI Elements
VLM
scene-understanding
Aug 20, 2026
Synergy: A Next-Generation General-Purpose Agent for Open Agentic Web
RL
scene-understanding
Aug 20, 2026
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
VLM
scene-understanding
3D-representation
Aug 20, 2026
Towards Visual Grounding: A Survey
scene-understanding
Aug 20, 2026
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
VLM
imitation-learning
scene-understanding
Aug 20, 2026
LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents
imitation-learning
scene-understanding
web-agent
Aug 20, 2026
Ponder & Press: Advancing Visual GUI Agent towards General Computer Control
imitation-learning
scene-understanding
web-agent
Aug 20, 2026
GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
scene-understanding
task-planning
web-agent
Aug 20, 2026
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
VLM
scene-understanding
web-agent
Aug 20, 2026
GUIrilla: A Scalable Framework for Automated Desktop UI Exploration
VLM
navigation
scene-understanding
Aug 20, 2026
MobileFlow: A Multimodal LLM For Mobile GUI Agent
VLM
VLA
scene-understanding
Aug 20, 2026
Continual GUI Agents
RL
scene-understanding
web-agent
Aug 20, 2026
GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent
navigation
scene-understanding
web-agent
Aug 20, 2026
Foundations of GenIR
VLM
imitation-learning
scene-understanding
Aug 20, 2026
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
imitation-learning
scene-understanding
web-agent
Aug 20, 2026
Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
navigation
RL
scene-understanding
Aug 20, 2026
ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning
scene-understanding
3D-representation
spatial-reasoning
Aug 20, 2026
AUTO-Explorer: Automated Data Collection for GUI Agent
navigation
imitation-learning
scene-understanding
Aug 20, 2026
AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
scene-understanding
web-agent
Aug 20, 2026
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
scene-understanding
web-agent
Aug 20, 2026
Agent-Initiated Interaction in Phone UI Automation
VLM
scene-understanding
Aug 20, 2026
GarmentPile: Point-Level Visual Affordance Guided Retrieval and Adaptation for Cluttered Garments Manipulation
manipulation
scene-understanding
spatial-reasoning
Aug 20, 2026
Embodied Scene Understanding for Vision Language Models via MetaVQA
scene-understanding
spatial-reasoning
VLM
Aug 20, 2026
An Embodied Generalist Agent in 3D World
VLA
scene-understanding
spatial-reasoning
data
Aug 20, 2026
Learning Foresightful Dense Visual Affordance for Deformable Object Manipulation
manipulation
scene-understanding
Aug 20, 2026
ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning
scene-understanding
task-planning
SLAM
Aug 20, 2026
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
manipulation
task-planning
scene-understanding
Aug 20, 2026
HomeRobot: Open-Vocabulary Mobile Manipulation
mobile-manipulation
navigation
manipulation
scene-understanding
Aug 20, 2026
TidyBot: Personalized Robot Assistance with Large Language Models
LLM
mobile-manipulation
scene-understanding
instruction-following
Aug 20, 2026
OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
scene-understanding
RL
Aug 20, 2026
Dynamic Open-Vocabulary 3D Scene Graphs for Long-term Language-Guided Mobile Manipulation
mobile-manipulation
scene-understanding
semantic-map
task-planning
Aug 20, 2026
Collaborative Dynamic 3D Scene Graphs for Automated Driving
scene-understanding
3D-representation
semantic-map
Aug 20, 2026
Visual Language Maps for Robot Navigation
semantic-map
navigation
scene-understanding
Aug 20, 2026
Rich Screen Reader Experiences for Accessible Data Visualization
navigation
imitation-learning
scene-understanding
Aug 20, 2026
RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
manipulation
embodied-reasoning
scene-understanding
Aug 20, 2026
MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation
mobile-manipulation
navigation
scene-understanding
Aug 20, 2026
Investigating Compositional Challenges in Vision-Language Models for Visual Grounding
VLM
scene-understanding
Aug 20, 2026
Learning Precise Affordances from Egocentric Videos for Robotic Manipulation
manipulation
scene-understanding
spatial-reasoning
Aug 20, 2026
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
world-model
scene-understanding
3D-representation
Aug 20, 2026
GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
VLM
spatial-reasoning
scene-understanding
Aug 20, 2026
CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving
VLA
scene-understanding
embodied-reasoning
Aug 20, 2026
Web-CogReasoner: Towards Knowledge-Induced Cognitive Reasoning for Web Agents
VLM
imitation-learning
scene-understanding