DeepRead
Search
Search
Dark mode
Light mode
Explorer
Popular Tags
#VLM
#agentic-RL
#gui-agent
#web-agent
#LLM
#VLA
#computer-use
#task-planning
#manipulation
#imitation-learning
Tag: RL
62 items with this tag.
Aug 20, 2026
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
agentic-RL
RL
LLM
Aug 20, 2026
Learning Robust Execution in Robotic Manipulation with Agentic Reinforcement Learning
manipulation
VLA
RL
Aug 20, 2026
QQWorld: Quantile-Quantile Matching for World Model Regularization
world-model
RL
Aug 20, 2026
World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
world-model
video-generation
RL
3D-consistency
Aug 20, 2026
WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
VLA
RL
world-model
Aug 20, 2026
N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
VLA
manipulation
RL
Aug 20, 2026
When Does Muon Help Agentic Reinforcement Learning?
agentic-RL
RL
Aug 20, 2026
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
agentic-RL
LLM
RL
Aug 20, 2026
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
VLM
RL
Aug 20, 2026
SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning for GUI Agents
RL
gui-agent
credit-assignment
long-horizon
Aug 20, 2026
RAGEN-2: Reasoning Collapse in Agentic RL
agentic-RL
RL
LLM
Aug 20, 2026
Dynamics-Aware Preference Optimization for Vision-Language Models
VLM
RL
Aug 20, 2026
Dual-Agent Reinforcement Learning for Adaptive and Cost-Aware Visual–Inertial Odometry
SLAM
navigation
RL
Aug 20, 2026
RLFTSim: Realistic and Controllable Multi-Agent Traffic Simulation via Reinforcement Learning Fine-Tuning
world-model
RL
Aug 20, 2026
PanoEnv: Exploring 3D Spatial Intelligence in Panoramic Environments with Reinforcement Learning
spatial-reasoning
VLM
RL
Aug 20, 2026
MangoBench: A Benchmark for Multi-Agent Goal-Conditioned Offline Reinforcement Learning
RL
manipulation
navigation
Aug 20, 2026
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning
agentic-RL
RL
LLM
Aug 20, 2026
Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning
RL
manipulation
Aug 20, 2026
Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime
agentic-RL
RL
Aug 20, 2026
TaskForce: Cooperative Multi-agent Reinforcement Learning for Multi-task Optimization
RL
Aug 20, 2026
Adaptive Milestone Reward for GUI Agents
agentic-RL
gui-agent
RL
Aug 20, 2026
RoboBrain 2.5: Depth in Sight, Time in Mind
spatial-reasoning
VLA
RL
Aug 20, 2026
UI-Venus Technical Report: Building High-performance UI Agents with RFT
VLM
navigation
RL
Aug 20, 2026
UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning
imitation-learning
RL
scene-understanding
Aug 20, 2026
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
navigation
imitation-learning
RL
Aug 20, 2026
Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
VLM
imitation-learning
RL
Aug 20, 2026
A Survey of Reinforcement Learning for Optimization in Automation
RL
cross-embodiment
Aug 20, 2026
A Survey on GUI Agents with Foundation Models Enhanced by Reinforcement Learning
VLA
RL
web-agent
Aug 20, 2026
Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent
imitation-learning
RL
scene-understanding
Aug 20, 2026
UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience
RL
task-planning
web-agent
Aug 20, 2026
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
navigation
RL
web-agent
Aug 20, 2026
Mobile-Agent-v3: Fundamental Agents for GUI Automation
RL
scene-understanding
web-agent
Aug 20, 2026
WebCanvas: Benchmarking Web Agents in Online Environments
imitation-learning
RL
scene-understanding
Aug 20, 2026
Synergy: A Next-Generation General-Purpose Agent for Open Agentic Web
RL
scene-understanding
Aug 20, 2026
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
VLM
imitation-learning
RL
Aug 20, 2026
LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
VLM
imitation-learning
RL
Aug 20, 2026
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
imitation-learning
RL
web-agent
Aug 20, 2026
GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training
VLM
imitation-learning
RL
Aug 20, 2026
History-Aware Reasoning for GUI Agents
imitation-learning
RL
task-planning
Aug 20, 2026
Continual GUI Agents
RL
scene-understanding
web-agent
Aug 20, 2026
GUI-PRA: Process Reward Agent for GUI Tasks
RL
task-planning
Aug 20, 2026
AI prediction leads people to forgo guaranteed rewards
RL
Aug 20, 2026
Improved GUI Grounding via Iterative Narrowing
VLM
imitation-learning
RL
Aug 20, 2026
Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
navigation
RL
scene-understanding
Aug 20, 2026
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
imitation-learning
RL
Aug 20, 2026
Augmenting the action space with conventions to improve multi-agent cooperation in Hanabi
manipulation
RL
task-planning
Aug 20, 2026
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
VLA
imitation-learning
RL
Aug 20, 2026
AOAD-MAT: Transformer-based multi-agent deep reinforcement learning model considering agents' order of action decisions
imitation-learning
RL
Aug 20, 2026
Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
agentic-RL
LLM
RL
Aug 20, 2026
AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning
VLM
imitation-learning
RL
Aug 20, 2026
Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment
VLM
imitation-learning
RL
Aug 20, 2026
A3: Android Agent Arena for Mobile GUI Agents with Essential-State Procedural Evaluation
imitation-learning
RL
web-agent
Aug 20, 2026
Free Process Rewards without Process Labels
agentic-RL
LLM
RL
Aug 20, 2026
Understanding World or Predicting Future? A Comprehensive Survey of World Models
world-model
VLA
RL
Aug 20, 2026
π*₀.₆: a VLA That Learns From Experience
VLA
RL
manipulation
flow-matching
Aug 20, 2026
Reward Hacking in Reinforcement Learning
agentic-RL
RL
LLM
Aug 20, 2026
OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
scene-understanding
RL
Aug 20, 2026
A Survey of Multi-Agent Deep Reinforcement Learning with Communication
RL
Aug 20, 2026
Digi-Q: Learning Q-Value Functions for Training Device-Control Agents
gui-agent
agentic-RL
RL
Aug 20, 2026
Robotic World Model: A Neural Network Simulator for Robust Policy Optimization in Robotics
world-model
RL
legged
Aug 20, 2026
Diffusion for World Modeling: Visual Details Matter in Atari
world-model
RL
diffusion-policy
Superhuman
Aug 20, 2026
White-Box AI Model: Next Frontier of Wireless Communications
imitation-learning
RL