DeepRead
Search
Search
Dark mode
Light mode
Explorer
Popular Tags
#VLM
#agentic-RL
#gui-agent
#web-agent
#LLM
#VLA
#computer-use
#task-planning
#manipulation
#imitation-learning
Tag: 3D-representation
45 items with this tag.
Aug 20, 2026
SAGE: Scalable Agentic 3D Scene Generation for Embodied AI
3D-representation
embodied-reasoning
mobile-manipulation
Aug 20, 2026
OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence
spatial-reasoning
VLM
3D-representation
Aug 20, 2026
Do as I Do: Dexterous Manipulation Data from Everyday Human Videos
imitation-learning
manipulation
3D-representation
Aug 20, 2026
DENALI: A Dataset Enabling Non-Line-of-Sight Spatial Reasoning with Low-Cost LiDARs
spatial-reasoning
scene-understanding
3D-representation
Aug 20, 2026
HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
world-model
3D-representation
VLM
Aug 20, 2026
Curvature-Aware Captioning: Leveraging Geodesic Attention for 3D Scene Understanding
scene-understanding
spatial-reasoning
3D-representation
Aug 20, 2026
Generative World Renderer
world-model
3D-representation
dataset
Aug 20, 2026
Generative World Renderer
world-model
3D-representation
Aug 20, 2026
AlayaWorld: Long-Horizon and Playable Video World Generation
world-model
3D-representation
Aug 20, 2026
PV-Ground: Text-Guided Point-Voxel Interaction for 3D Visual Grounding
scene-understanding
3D-representation
spatial-reasoning
Aug 20, 2026
OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding
scene-understanding
3D-representation
semantic-map
Aug 20, 2026
NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into Code
auto-research
3D-representation
LLM
Aug 20, 2026
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
VLA
manipulation
3D-representation
Aug 20, 2026
WRIVINDER: Towards Spatial Intelligence for Geo-locating Ground Images onto Satellite Imagery
spatial-reasoning
3D-representation
navigation
Aug 20, 2026
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
spatial-reasoning
VLM
3D-representation
Aug 20, 2026
Vinedresser3D: Towards Agentic Text-guided 3D Editing
3D-representation
VLM
task-planning
Aug 20, 2026
3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training
VLA
spatial-reasoning
manipulation
3D-representation
Aug 20, 2026
VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery
VLM
3D-representation
embodied-reasoning
Aug 20, 2026
Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images
3D-representation
spatial-reasoning
embodied-reasoning
Aug 20, 2026
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
VLM
3D-representation
spatial-reasoning
Aug 20, 2026
Unsupervised Multi-agent and Single-agent Perception from Cooperative Views
scene-understanding
3D-representation
Aug 20, 2026
Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Clouds
scene-understanding
3D-representation
spatial-reasoning
Aug 20, 2026
TopoMA: Topology-Guided Multi-Agent Dense RGB 3D Reconstruction via Distributed Inference
3D-representation
SLAM
spatial-memory
Aug 20, 2026
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
spatial-reasoning
VLM
3D-representation
Aug 20, 2026
HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
spatial-reasoning
VLM
3D-representation
Aug 20, 2026
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
spatial-reasoning
VLM
3D-representation
Aug 20, 2026
Grounded 3D-Aware Spatial Vision-Language Modeling
spatial-reasoning
VLM
3D-representation
Aug 20, 2026
Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
3D-representation
scene-understanding
spatial-reasoning
Aug 20, 2026
SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
spatial-reasoning
VLM
3D-representation
Aug 20, 2026
SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMs
spatial-reasoning
VLM
3D-representation
Aug 20, 2026
GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
VLA
manipulation
3D-representation
Aug 20, 2026
DyGeoVLN: Infusing Dynamic Geometry Foundation Model into Vision-Language Navigation
VLN
3D-representation
spatial-reasoning
Aug 20, 2026
GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors
imitation-learning
manipulation
VLA
3D-representation
Aug 20, 2026
G2 VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
spatial-reasoning
3D-representation
VLM
Aug 20, 2026
ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation
VLA
3D-representation
spatial-reasoning
manipulation
Aug 20, 2026
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
VLM
scene-understanding
3D-representation
Aug 20, 2026
ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning
scene-understanding
3D-representation
spatial-reasoning
Aug 20, 2026
ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping
manipulation
3D-representation
world-model
Aug 20, 2026
SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM
SLAM
3D-representation
1-2
Aug 20, 2026
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models
spatial-reasoning
VLM
3D-representation
Aug 20, 2026
VGGT: Visual Geometry Grounded Transformer
3D-representation
Aug 20, 2026
Collaborative Dynamic 3D Scene Graphs for Automated Driving
scene-understanding
3D-representation
semantic-map
Aug 20, 2026
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
world-model
scene-understanding
3D-representation
Aug 20, 2026
EmbodiedSplat: Personalized Real-to-Sim-to-Real Navigation with Gaussian Splats from a Mobile Device
navigation
3D-representation
Aug 20, 2026
OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving
world-model
3D-representation