DeepRead
Search
Search
Dark mode
Light mode
Explorer
Popular Tags
#VLM
#agentic-RL
#gui-agent
#LLM
#task-planning
#web-agent
#VLA
#manipulation
#computer-use
#world-model
Tag: task-planning
183 items with this tag.
Sep 30, 2026
Agent Harness 的组件归因:外置 state、fresh-context 执行与独立验证,哪一个在起作用
survey
task-planning
LLM
computer-use
Sep 30, 2026
TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents
agentic-RL
task-planning
LLM
Sep 30, 2026
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
LLM
task-planning
Sep 30, 2026
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
LLM
task-planning
Sep 30, 2026
SkillZip Pro: Execution-Aware Dynamic Compression of Progressively Loaded Skills for Self-Evolving Agents
task-planning
agentic-RL
Sep 30, 2026
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure
task-planning
agentic-RL
Sep 30, 2026
SkillJack: Persistent Skill Backdoors in Self-Evolving Agents
agentic-RL
task-planning
LLM
Sep 30, 2026
SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation
LLM
task-planning
Sep 30, 2026
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
LLM
task-planning
Sep 30, 2026
Web Agent Harness 设计:动作接口、执行循环与上下文预算
survey
web-agent
gui-agent
task-planning
LLM
Sep 30, 2026
Prime Agent: A Self-Improving RLM Harness
LLM
task-planning
agentic-RL
Sep 30, 2026
Practice: From Experience to Expertise in Self-Evolving Embodied Agents
task-planning
embodied-reasoning
agentic-RL
Sep 30, 2026
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
LLM
task-planning
agentic-RL
Sep 30, 2026
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks
gui-agent
computer-use
task-planning
Sep 30, 2026
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
computer-use
task-planning
gui-agent
Sep 30, 2026
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
task-planning
agentic-RL
LLM
Sep 30, 2026
Skills in Weights, Memory in Code: Hybrid Learning for Memory-Dependent Robot Manipulation
manipulation
VLA
task-planning
Sep 30, 2026
Modeling What Changes: Sparse, Residual World Models for Object-Centric Manipulation
world-model
manipulation
task-planning
Sep 30, 2026
Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents
task-planning
LLM
Sep 30, 2026
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
task-planning
LLM
Sep 30, 2026
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
auto-research
task-planning
LLM
Sep 30, 2026
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
computer-use
agentic-RL
task-planning
Sep 30, 2026
EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents
agentic-RL
task-planning
LLM
Sep 30, 2026
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
LLM
task-planning
Sep 30, 2026
DarwinX: Evolving Agent Harnesses Through Natural Selection
task-planning
gui-agent
web-agent
Sep 30, 2026
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
LLM
task-planning
Sep 30, 2026
Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
world-model
task-planning
manipulation
Sep 30, 2026
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
agentic-RL
task-planning
LLM
Sep 30, 2026
Omni Interaction Agent Technical Report
VLM
hci
task-planning
Sep 30, 2026
Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents
LLM
task-planning
Sep 30, 2026
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
LLM
task-planning
Sep 30, 2026
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces
LLM
task-planning
Sep 30, 2026
Apodex 1.1: Scaling Agentic Intelligence for Complex Work
agentic-RL
task-planning
LLM
Sep 30, 2026
Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents
task-planning
LLM
Sep 30, 2026
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization
agentic-RL
task-planning
LLM
Sep 30, 2026
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
LLM
task-planning
Sep 30, 2026
2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation
VLA
manipulation
task-planning
Sep 30, 2026
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence
task-planning
manipulation
VLA
Sep 30, 2026
World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
world-model
task-planning
manipulation
Sep 30, 2026
Is Progressive Disclosure All You Need for Long-Context Agents?
LLM
task-planning
Sep 30, 2026
Plover: Steering GUI Agents through Plan-Centric Interaction
gui-agent
task-planning
instruction-following
Sep 30, 2026
PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning
agentic-RL
task-planning
Sep 30, 2026
Quo Vadis, World Modeling? Towards Interactive World Proxies for Continually Improving Agents
world-model
agentic-RL
task-planning
Sep 30, 2026
Object-Centric Environment Modeling for Agentic Tasks
world-model
task-planning
Sep 30, 2026
A Task-State Representation for Long-Horizon Mobile GUI Agents
gui-agent
task-planning
Sep 30, 2026
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents
computer-use
gui-agent
task-planning
Sep 30, 2026
WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments
gui-agent
computer-use
task-planning
Sep 30, 2026
MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution
agentic-RL
LLM
task-planning
Sep 30, 2026
SKILL-KD: Contrastive Skill Distillation for LLM Agents
task-planning
LLM
agentic-RL
Sep 30, 2026
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
LLM
task-planning
Sep 30, 2026
The Tool Illusion: Rethinking Tool Use in Web Agents
web-agent
gui-agent
task-planning
Sep 30, 2026
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
web-agent
task-planning
LLM
Sep 30, 2026
SkillFlow: Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents
task-planning
LLM
Skills
Sep 30, 2026
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
computer-use
task-planning
LLM
Sep 30, 2026
MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems
agentic-RL
LLM
task-planning
Sep 30, 2026
Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember
agentic-RL
deep-research
task-planning
Sep 30, 2026
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
agentic-RL
LLM
task-planning
Sep 30, 2026
Skill0: In-Context Agentic Reinforcement Learning for Skill Internalization
agentic-RL
task-planning
LLM
Sep 30, 2026
SEE: Structure-aware Exploring & Exploiting for Long-horizon GUI Agent Trajectory Synthesis
gui-agent
task-planning
Sep 30, 2026
Self-Evolving Agents with Anytime-Valid Certificates
agentic-RL
task-planning
Sep 30, 2026
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
computer-use
task-planning
Sep 30, 2026
KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill
gui-agent
computer-use
task-planning
Sep 30, 2026
Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds
agentic-RL
LLM
task-planning
Sep 30, 2026
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
computer-use
task-planning
VLM
Sep 30, 2026
Harness Handbook: Making Evolving Agent Harnesses Readable, Navigable, and Editable
LLM
task-planning
Sep 30, 2026
Recursive Agent Harnesses
LLM
task-planning
Sep 30, 2026
Don't Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent Quality
LLM
task-planning
Sep 30, 2026
ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL
agentic-RL
task-planning
Sep 30, 2026
HarnessBank: Semantic Gene-Bank Search with Gated Verification for Agent-Harness Self-Evolution
LLM
task-planning
Sep 30, 2026
Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
VLA
embodied-reasoning
task-planning
Sep 30, 2026
Coarse-to-Control: Action-Token Planning for Vision-Language-Action Models
VLA
task-planning
manipulation
Sep 30, 2026
GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)
agentic-RL
task-planning
LLM
Sep 30, 2026
Beyond Sequential Tools: A Unified VLM Agent System for Photographic Post-Processing via Dynamic Multi-Expert Fusion
VLM
task-planning
instruction-following
Sep 30, 2026
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
auto-research
task-planning
LLM
Sep 30, 2026
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
LLM
task-planning
agentic-RL
Sep 30, 2026
Heterogeneous Scientific Foundation Model Collaboration
LLM
agentic-RL
task-planning
Sep 30, 2026
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
agentic-RL
task-planning
world-model
Sep 30, 2026
OpenRath: Session-Centered Runtime State for Agent Systems
LLM
task-planning
Sep 30, 2026
OctoT2I: A Self-Evolving Agentic Text-to-Image Router
task-planning
VLM
LLM
Sep 30, 2026
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
computer-use
gui-agent
task-planning
Sep 30, 2026
Agentic Retoucher for Text-To-Image Generation
VLM
task-planning
agentic-RL
Sep 30, 2026
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
task-planning
spatial-memory
embodied-reasoning
Sep 30, 2026
Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
gui-agent
computer-use
task-planning
Sep 30, 2026
MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management
gui-agent
task-planning
LLM
Sep 30, 2026
Dreaming when Necessary: Advancing World Action Models with Adaptive Multi-Modal Reasoning
VLA
world-model
embodied-reasoning
task-planning
Sep 30, 2026
Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time Scaling
VLM
task-planning
agentic-RL
Sep 30, 2026
Multi-Agent Computer Use
computer-use
gui-agent
task-planning
Sep 30, 2026
Vinedresser3D: Towards Agentic Text-guided 3D Editing
3D-representation
VLM
task-planning
Sep 30, 2026
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
video-LLM
video-understanding
task-planning
Sep 30, 2026
AGiLe: Learning Robust Long-Horizon Manipulation via Affordance-Grounded Bidirectional Latent Planning
manipulation
task-planning
spatial-reasoning
Sep 30, 2026
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
VLM
task-planning
spatial-reasoning
Sep 30, 2026
LensWalk: Agentic Video Understanding by Planning How You See in Videos
video-LLM
video-understanding
task-planning
Sep 30, 2026
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
computer-use
gui-agent
LLM
task-planning
Sep 30, 2026
Agentic Compilation: Mitigating the LLM Rerun Crisis for Minimized-Inference-Cost Web Automation
gui-agent
web-agent
task-planning
Sep 30, 2026
LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents
agentic-RL
task-planning
LLM
Sep 30, 2026
Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning
agentic-RL
VLA
task-planning
Sep 30, 2026
Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation
gui-agent
agentic-RL
task-planning
Sep 30, 2026
TeamBench: Evaluating Agent Coordination under Enforced Role Separation
task-planning
LLM
hci
Sep 30, 2026
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction
agentic-RL
task-planning
Sep 30, 2026
SEVerA: Verified Synthesis of Self-Evolving Agents
agentic-RL
task-planning
Sep 30, 2026
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
agentic-RL
task-planning
LLM
Sep 30, 2026
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning
agentic-RL
task-planning
Sep 30, 2026
RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
VLA
manipulation
task-planning
Sep 30, 2026
SyncMos: Scalable Motion Synchronisation for Multi-Agent Scene Interaction
embodied-reasoning
task-planning
scene-understanding
Sep 30, 2026
Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding
video-LLM
video-understanding
task-planning
Sep 30, 2026
Hybrid Self-evolving Structured Memory for GUI Agents
VLM
task-planning
web-agent
Sep 30, 2026
GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents
computer-use
gui-agent
task-planning
Sep 30, 2026
Web Agents Should Adopt the Plan-Then-Execute Paradigm
web-agent
gui-agent
task-planning
Sep 30, 2026
π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
gui-agent
task-planning
LLM
Sep 30, 2026
SKILL.nb: Selective Formalization and Gated Execution for Durable Agent Workflows
web-agent
task-planning
Sep 30, 2026
Are Online Skill and Memory Modules Always Worth Their Tokens? A Budget-Constrained Study of Web Agents
gui-agent
web-agent
task-planning
Sep 30, 2026
SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research
web-agent
task-planning
LLM
Sep 30, 2026
SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System
video-understanding
task-planning
video-LLM
Sep 30, 2026
FastContext: Training Efficient Repository Explorer for Coding Agents
LLM
task-planning
auto-research
Sep 30, 2026
MMSkills: Towards Multimodal Skills for General Visual Agents
gui-agent
task-planning
VLM
Sep 30, 2026
AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents
gui-agent
task-planning
Sep 30, 2026
Experience Transfer for Multimodal LLM Agents in Minecraft Game
embodied-reasoning
task-planning
VLM
Sep 30, 2026
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
computer-use
web-agent
task-planning
Sep 30, 2026
Harnessing LLM Agents with Skill Programs
agentic-RL
LLM
task-planning
Sep 30, 2026
GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents
agentic-RL
task-planning
Sep 30, 2026
Robix: A Unified Model for Robot Interaction, Reasoning and Planning
embodied-reasoning
task-planning
VLM
Sep 30, 2026
OmniEVA: Embodied Versatile PlAnner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning
embodied-reasoning
spatial-reasoning
task-planning
Sep 30, 2026
World-Model-Augmented Web Agents with Action Correction
web-agent
world-model
task-planning
Sep 30, 2026
Code as Agent Harness: Toward Executable, Verifiable, and Stateful Agent Systems
LLM
gui-agent
computer-use
task-planning
Sep 30, 2026
UniPlan: Vision-Language Task Planning for Mobile Manipulation with Unified PDDL Formulation
task-planning
mobile-manipulation
embodied-reasoning
Sep 30, 2026
RynnBrain: Open Embodied Foundation Models
VLA
spatial-reasoning
task-planning
Sep 30, 2026
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
agentic-RL
LLM
task-planning
Steps
Sep 30, 2026
MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
gui-agent
task-planning
Sep 30, 2026
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
auto-research
LLM
task-planning
Sep 30, 2026
Building Self-Evolving Agents via Experience-Driven Lifelong Learning: A Framework and Benchmark
agentic-RL
LLM
task-planning
Self-Motivat
Sep 30, 2026
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
agentic-RL
LLM
task-planning
Sep 30, 2026
Thinker: A vision-language foundation model for embodied intelligence
VLM
embodied-reasoning
task-planning
Sep 30, 2026
OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive Tasks
computer-use
gui-agent
task-planning
Sep 30, 2026
MAGNET: Towards Adaptive GUI Agents with Memory-Driven Knowledge Evolution
gui-agent
task-planning
Sep 30, 2026
BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
mobile-manipulation
task-planning
Sep 30, 2026
ActionEngine: From Reactive to Programmatic GUI Agents via State Machine Memory
gui-agent
web-agent
task-planning
Sep 30, 2026
CoAct-1: Computer-using Multi-agent System with Coding Actions
computer-use
gui-agent
task-planning
Sep 30, 2026
WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis
web-agent
world-model
task-planning
Sep 30, 2026
UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience
RL
task-planning
web-agent
Sep 30, 2026
PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
navigation
imitation-learning
task-planning
Sep 30, 2026
MobileDreamer: Generative Sketch World Model for GUI Agent
world-model
task-planning
web-agent
Sep 30, 2026
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
agentic-RL
LLM
task-planning
Sep 30, 2026
RoboBrain 2.0 Technical Report
spatial-reasoning
VLM
task-planning
Samples
Tunable
Sep 30, 2026
GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
scene-understanding
task-planning
web-agent
Sep 30, 2026
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
VLM
imitation-learning
task-planning
Sep 30, 2026
BEAP-Agent: Backtrackable Execution and Adaptive Planning for GUI Agents
navigation
task-planning
web-agent
Sep 30, 2026
History-Aware Reasoning for GUI Agents
imitation-learning
RL
task-planning
Sep 30, 2026
WebOperator: Action-Aware Tree Search for Autonomous Agents in Web Environment
web-agent
task-planning
Sep 30, 2026
GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration
navigation
imitation-learning
task-planning
Sep 30, 2026
GUI-ReWalk: Massive Data Generation for GUI Agent via Stochastic Exploration and Intent-Aware Reasoning
navigation
imitation-learning
task-planning
Sep 30, 2026
GUI-PRA: Process Reward Agent for GUI Tasks
RL
task-planning
Sep 30, 2026
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
VLM
imitation-learning
task-planning
Sep 30, 2026
V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents
VLM
task-planning
instruction-following
Sep 30, 2026
Towards a Science of Scaling Agent Systems
LLM
task-planning
web-agent
Sep 30, 2026
Augmenting the action space with conventions to improve multi-agent cooperation in Hanabi
manipulation
RL
task-planning
Sep 30, 2026
NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction
VLN
world-model
task-planning
Sep 30, 2026
Auto-scaling Continuous Memory for GUI Agent
VLM
task-planning
web-agent
Sep 30, 2026
AMAP Agentic Planning Technical Report
imitation-learning
task-planning
web-agent
Sep 30, 2026
MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments
gui-agent
task-planning
instruction-following
Sep 30, 2026
ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning
scene-understanding
task-planning
SLAM
Sep 30, 2026
Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method
VLN
navigation
task-planning
1/2
Sep 30, 2026
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
manipulation
task-planning
scene-understanding
Sep 30, 2026
Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence
web-agent
navigation
task-planning
Sep 30, 2026
AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management
gui-agent
task-planning
Sep 30, 2026
Audited Skill-Graph Self-Improvement for Agentic LLMs via Verifiable Rewards, Experience Synthesis, and Continual Memory
agentic-RL
task-planning
Sep 30, 2026
Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
web-agent
world-model
task-planning
Sep 30, 2026
Reasoning with Language Model is Planning with World Model
LLM
task-planning
world-model
Sep 30, 2026
PaLM-E: An Embodied Multimodal Language Model
VLM
embodied-reasoning
task-planning
Sep 30, 2026
Dynamic Open-Vocabulary 3D Scene Graphs for Long-term Language-Guided Mobile Manipulation
mobile-manipulation
scene-understanding
semantic-map
task-planning
Sep 30, 2026
Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
computer-use
gui-agent
task-planning
Sep 30, 2026
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
task-planning
instruction-following
mobile-manipulation
1/3
Sep 30, 2026
BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation
mobile-manipulation
embodied-reasoning
task-planning
Sep 30, 2026
Tree Search for Language Model Agents
web-agent
task-planning
Sep 30, 2026
VL-Nav: A Neuro-Symbolic Approach for Reasoning-based Vision-Language Navigation
VLN
navigation
task-planning
Sep 30, 2026
RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
embodied-reasoning
manipulation
task-planning
Sep 30, 2026
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
VLA
instruction-following
task-planning
manipulation
embodied-reasoning
Sep 30, 2026
HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation
VLA
manipulation
task-planning
Sep 30, 2026
Hierarchical Diffusion Policy for Kinematics-Aware Multi-Task Robotic Manipulation
diffusion-policy
manipulation
task-planning
Sep 30, 2026
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
embodied-reasoning
VLM
task-planning
Env
Tasks
Sep 30, 2026
EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence
embodied-reasoning
task-planning
agentic-RL
Sep 30, 2026
A Survey on Vision-Language-Action Models for Embodied AI
VLA
task-planning
manipulation
Sep 30, 2026
Branch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action Memory
web-agent
task-planning
Sep 30, 2026
UISim: An Interactive Image-Based UI Simulator for Dynamic Mobile Environments
navigation
task-planning