DeepRead
Search
Search
Dark mode
Light mode
Explorer
Popular Tags
#VLM
#agentic-RL
#gui-agent
#LLM
#task-planning
#web-agent
#VLA
#manipulation
#computer-use
#world-model
Tag: LLM
202 items with this tag.
Sep 30, 2026
Self-Evolving and Self-Improving Agents: A Unified Survey of Evolution Targets, Feedback, Gating, and Safety
survey
self-evolving-agents
self-improvement
recursive-self-improvement
agentic-RL
LLM
misevolution
Sep 30, 2026
Agent Harness 的组件归因:外置 state、fresh-context 执行与独立验证,哪一个在起作用
survey
task-planning
LLM
computer-use
Sep 30, 2026
TTPO: Test-Time Policy Optimization
LLM
agentic-RL
Sep 30, 2026
TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents
agentic-RL
task-planning
LLM
Sep 30, 2026
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
LLM
task-planning
Sep 30, 2026
Stealing Reasoning Traces from Proprietary LLM APIs
LLM
Sep 30, 2026
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
LLM
task-planning
Sep 30, 2026
SkillJack: Persistent Skill Backdoors in Self-Evolving Agents
agentic-RL
task-planning
LLM
Sep 30, 2026
SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation
LLM
task-planning
Sep 30, 2026
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
LLM
task-planning
Sep 30, 2026
RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States
agentic-RL
LLM
Sep 30, 2026
Web Agent Harness 设计:动作接口、执行循环与上下文预算
survey
web-agent
gui-agent
task-planning
LLM
Sep 30, 2026
Prime Agent: A Self-Improving RLM Harness
LLM
task-planning
agentic-RL
Sep 30, 2026
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning
agentic-RL
LLM
Sep 30, 2026
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
agentic-RL
LLM
computer-use
Sep 30, 2026
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
LLM
task-planning
agentic-RL
Sep 30, 2026
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
agentic-RL
LLM
Sep 30, 2026
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
agentic-RL
LLM
Sep 30, 2026
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
task-planning
agentic-RL
LLM
Sep 30, 2026
Unlocking Lossless Speedups in LLMs via Discrete Diffusion
LLM
Sep 30, 2026
Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents
task-planning
LLM
Sep 30, 2026
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
task-planning
LLM
Sep 30, 2026
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
auto-research
task-planning
LLM
Sep 30, 2026
EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents
agentic-RL
task-planning
LLM
Sep 30, 2026
EnvHarness: Awakening Static Worlds for Agent Learning
agentic-RL
LLM
Sep 30, 2026
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
agentic-RL
LLM
Sep 30, 2026
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
LLM
task-planning
Sep 30, 2026
DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
world-model
LLM
Sep 30, 2026
Iris: Climbing to the Search Frontier
deep-research
agentic-RL
LLM
Sep 30, 2026
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
LLM
task-planning
Sep 30, 2026
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
agentic-RL
task-planning
LLM
Sep 30, 2026
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
agentic-RL
LLM
RL
Sep 30, 2026
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
agentic-RL
LLM
VLM
Sep 30, 2026
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
agentic-RL
LLM
Sep 30, 2026
Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents
LLM
task-planning
Sep 30, 2026
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
LLM
task-planning
Sep 30, 2026
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction
LLM
Sep 30, 2026
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
LLM
spatial-reasoning
Sep 30, 2026
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces
LLM
task-planning
Sep 30, 2026
Aspire: Can Models Self-Evolve from Vague Goals?
agentic-RL
auto-research
LLM
Sep 30, 2026
Apodex 1.1: Scaling Agentic Intelligence for Complex Work
agentic-RL
task-planning
LLM
Sep 30, 2026
Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents
task-planning
LLM
Sep 30, 2026
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization
agentic-RL
task-planning
LLM
Sep 30, 2026
Atria Dawn: The Dawn of Agentic Superintelligence
auto-research
LLM
hci
Sep 30, 2026
What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
agentic-RL
LLM
Sep 30, 2026
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
agentic-RL
LLM
rollouts
Sep 30, 2026
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
agentic-RL
LLM
Sep 30, 2026
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
LLM
task-planning
Sep 30, 2026
Zero-Mem: Zero-Token Memory Operations for LLM Agents
LLM
Sep 30, 2026
Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI
LLM
agentic-RL
Sep 30, 2026
Is Progressive Disclosure All You Need for Long-Context Agents?
LLM
task-planning
Sep 30, 2026
Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
web-agent
LLM
Sep 30, 2026
ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning
agentic-RL
LLM
Sep 30, 2026
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
agentic-RL
LLM
RL
Sep 30, 2026
How Benchmarks Mis-Score Computer-Use Agents
computer-use
gui-agent
web-agent
benchmark
evaluation
LLM
Sep 30, 2026
MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution
agentic-RL
LLM
task-planning
Sep 30, 2026
SKILL-KD: Contrastive Skill Distillation for LLM Agents
task-planning
LLM
agentic-RL
Sep 30, 2026
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
LLM
Sep 30, 2026
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
LLM
task-planning
Sep 30, 2026
Mental World Modeling
world-model
LLM
VLM
Sep 30, 2026
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
web-agent
task-planning
LLM
Sep 30, 2026
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
LLM
instruction-following
Sep 30, 2026
Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making
LLM
gui-agent
VLM
Sep 30, 2026
SkillFlow: Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents
task-planning
LLM
Skills
Sep 30, 2026
SWE-Explore: Benchmarking How Coding Agents Explore Repositories
gui-agent
LLM
Sep 30, 2026
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
computer-use
task-planning
LLM
Sep 30, 2026
MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems
agentic-RL
LLM
task-planning
Sep 30, 2026
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
agentic-RL
LLM
task-planning
Sep 30, 2026
Skill0: In-Context Agentic Reinforcement Learning for Skill Internalization
agentic-RL
task-planning
LLM
Sep 30, 2026
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
agentic-RL
LLM
Sep 30, 2026
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
agentic-RL
RL
LLM
Sep 30, 2026
Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training
agentic-RL
LLM
Sep 30, 2026
Recursive Multi-Agent Systems
LLM
agentic-RL
Sep 30, 2026
Kimi K3: Open Frontier Intelligence
LLM
agentic-RL
VLM
Sep 30, 2026
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
agentic-RL
LLM
Sep 30, 2026
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
auto-research
LLM
Sep 30, 2026
Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds
agentic-RL
LLM
task-planning
Sep 30, 2026
RAGEN-2: Reasoning Collapse in Agentic RL
agentic-RL
RL
LLM
Sep 30, 2026
Harness Handbook: Making Evolving Agent Harnesses Readable, Navigable, and Editable
LLM
task-planning
Sep 30, 2026
Does RL Expand the Capability Boundary of LLM Agents? A PASS@(k,T) Analysis
agentic-RL
deep-research
LLM
Sep 30, 2026
Recursive Agent Harnesses
LLM
task-planning
Sep 30, 2026
Don't Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent Quality
LLM
task-planning
Sep 30, 2026
HarnessBank: Semantic Gene-Bank Search with Gated Verification for Agent-Harness Self-Evolution
LLM
task-planning
Sep 30, 2026
Gemma 4 Technical Report
VLM
LLM
Sep 30, 2026
RAAS: LLM Agentic System Architecture Search with GRPO
agentic-RL
LLM
Sep 30, 2026
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
auto-research
agentic-RL
LLM
Sep 30, 2026
When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs
LLM
agentic-RL
Sep 30, 2026
QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents
agentic-RL
LLM
Sep 30, 2026
GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)
agentic-RL
task-planning
LLM
Sep 30, 2026
Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation
LLM
instruction-following
Sep 30, 2026
PrivacyAlign: Contextual Privacy Alignment for LLM Agents
computer-use
agentic-RL
LLM
Sep 30, 2026
Weak-to-Strong Generalization via Direct On-Policy Distillation
LLM
agentic-RL
Sep 30, 2026
PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agents
LLM
instruction-following
Sep 30, 2026
AI Agents Do Not Fail Alone: The Context Fails First
LLM
instruction-following
Sep 30, 2026
Paper2Figure: A Multi-Agent Collaborative System for Figure Generation Towards Academic Research Paper
auto-research
LLM
instruction-following
Sep 30, 2026
When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents
computer-use
LLM
agentic-RL
Sep 30, 2026
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
auto-research
task-planning
LLM
Sep 30, 2026
Fast-dVLM: Efficient Block-Diffusion VLM via Direct Conversion from Autoregressive VLM
VLM
LLM
Sep 30, 2026
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
LLM
task-planning
agentic-RL
Sep 30, 2026
Always-On Agents: A Survey of Persistent Memory, State, and Governance in LLM Agents
LLM
Sep 30, 2026
Heterogeneous Scientific Foundation Model Collaboration
LLM
agentic-RL
task-planning
Sep 30, 2026
How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks
LLM
Sep 30, 2026
Agents' Last Exam
computer-use
gui-agent
LLM
Sep 30, 2026
OpenRath: Session-Centered Runtime State for Agent Systems
LLM
task-planning
Sep 30, 2026
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
agentic-RL
LLM
Sep 30, 2026
OctoT2I: A Self-Evolving Agentic Text-to-Image Router
task-planning
VLM
LLM
Sep 30, 2026
Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
VLM
LLM
embodied-reasoning
Sep 30, 2026
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
auto-research
LLM
Sep 30, 2026
From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents
LLM
deep-research
Sep 30, 2026
Are We Ready For An Agent-Native Memory System?
LLM
web-agent
Sep 30, 2026
NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into Code
auto-research
3D-representation
LLM
Sep 30, 2026
AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object Detection
VLM
scene-understanding
LLM
Sep 30, 2026
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
LLM
Sep 30, 2026
Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection
VLM
LLM
Sep 30, 2026
MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video Synthesis
world-model
LLM
VLM
Sep 30, 2026
MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management
gui-agent
task-planning
LLM
Sep 30, 2026
Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
spatial-reasoning
scene-understanding
LLM
Sep 30, 2026
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
computer-use
agentic-RL
LLM
Sep 30, 2026
Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies
agentic-RL
LLM
Sep 30, 2026
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning
agentic-RL
RL
LLM
Sep 30, 2026
Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents
computer-use
agentic-RL
LLM
Sep 30, 2026
ViLoMem: Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
VLM
LLM
Sep 30, 2026
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
auto-research
LLM
web-agent
Sep 30, 2026
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
computer-use
gui-agent
LLM
task-planning
Sep 30, 2026
LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents
agentic-RL
task-planning
LLM
Sep 30, 2026
TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation
LLM
spatial-reasoning
Sep 30, 2026
Unlimited OCR Works
VLM
LLM
Sep 30, 2026
Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent
VLM
LLM
instruction-following
Sep 30, 2026
TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs
VLM
LLM
Sep 30, 2026
TeamBench: Evaluating Agent Coordination under Enforced Role Separation
task-planning
LLM
hci
Sep 30, 2026
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
agentic-RL
task-planning
LLM
Sep 30, 2026
Tackling Model Bias via Game-theoretic Multi-agent Collaboration Framework for Hateful Meme Classification
VLM
LLM
Sep 30, 2026
Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops
agentic-RL
LLM
Sep 30, 2026
PersonaVLM: Long-Term Personalized Multimodal LLMs
VLM
agentic-RL
LLM
Sep 30, 2026
GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning
VLM
LLM
Sep 30, 2026
SEAL: Synergistic Co-Evolution of Agents and Learning Environments
agentic-RL
LLM
Sep 30, 2026
Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps
LLM
Sep 30, 2026
π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
gui-agent
task-planning
LLM
Sep 30, 2026
EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery
auto-research
LLM
Sep 30, 2026
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
LLM
Sep 30, 2026
SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research
web-agent
task-planning
LLM
Sep 30, 2026
Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems
multi-agent
LLM
self-evolving-agents
Sep 30, 2026
Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients
VLM
LLM
Sep 30, 2026
FastContext: Training Efficient Repository Explorer for Coding Agents
LLM
task-planning
auto-research
Sep 30, 2026
Ares: Adaptive Reasoning Effort Selection for Efficient LLM Agents
LLM
web-agent
agentic-RL
Sep 30, 2026
Masking Stale Observations Helps Search Agents -- Until It Doesn't: A Regime Map and Its Mechanism
deep-research
LLM
Sep 30, 2026
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application
agentic-RL
LLM
world-model
Sep 30, 2026
AgentSwing: Adaptive Parallel Context Management Routing for Long-Horizon Web Agents
deep-research
LLM
Sep 30, 2026
Harnessing LLM Agents with Skill Programs
agentic-RL
LLM
task-planning
Sep 30, 2026
When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training
agentic-RL
LLM
Sep 30, 2026
When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents
computer-use
LLM
Sep 30, 2026
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
agentic-RL
LLM
Sep 30, 2026
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards
agentic-RL
LLM
Sep 30, 2026
Dockerless: Environment-Free Program Verifier for Coding Agents
agentic-RL
LLM
auto-research
Sep 30, 2026
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
LLM
gui-agent
VLM
Sep 30, 2026
Code as Agent Harness: Toward Executable, Verifiable, and Stateful Agent Systems
LLM
gui-agent
computer-use
task-planning
Sep 30, 2026
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
gui-agent
agentic-RL
LLM
Sep 30, 2026
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
agentic-RL
LLM
computer-use
Sep 30, 2026
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook
LLM
Sep 30, 2026
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
agentic-RL
LLM
Sep 30, 2026
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
agentic-RL
LLM
task-planning
Steps
Sep 30, 2026
Active Learners as Efficient PRP Rerankers
LLM
agentic-RL
Sep 30, 2026
Kimi K2.5: Visual Agentic Intelligence
VLM
agentic-RL
LLM
Sep 30, 2026
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
auto-research
LLM
task-planning
Sep 30, 2026
Building Self-Evolving Agents via Experience-Driven Lifelong Learning: A Framework and Benchmark
agentic-RL
LLM
task-planning
Self-Motivat
Sep 30, 2026
AI for Auto-Research: Roadmap & User Guide
auto-research
LLM
Sep 30, 2026
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
agentic-RL
LLM
task-planning
Sep 30, 2026
ACC: Compiling Agent Trajectories for Long-Context Training
agentic-RL
LLM
Sep 30, 2026
MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
agentic-RL
LLM
Sep 30, 2026
Agentic Test-Time Scaling for WebAgents
web-agent
gui-agent
LLM
Sep 30, 2026
Agent Lightning: Train ANY AI Agents with Reinforcement Learning
agentic-RL
LLM
Sep 30, 2026
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
agentic-RL
LLM
task-planning
Sep 30, 2026
Towards a Science of Scaling Agent Systems
LLM
task-planning
web-agent
Sep 30, 2026
Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
agentic-RL
LLM
RL
Sep 30, 2026
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
agentic-RL
LLM
Sep 30, 2026
MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
LLM
agentic-RL
Sep 30, 2026
GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators
agentic-RL
LLM
Sep 30, 2026
FoldAct: Efficient and Stable Context Folding for Long-Horizon Search Agents
deep-research
agentic-RL
LLM
Sep 30, 2026
WebArena: A Realistic Web Environment for Building Autonomous Agents
web-agent
LLM
computer-use
Sep 30, 2026
Free Process Rewards without Process Labels
agentic-RL
LLM
RL
Sep 30, 2026
TidyBot: Personalized Robot Assistance with Large Language Models
LLM
mobile-manipulation
scene-understanding
instruction-following
Sep 30, 2026
Reasoning with Language Model is Planning with World Model
LLM
task-planning
world-model
Sep 30, 2026
Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents
agentic-RL
LLM
auto-research
Sep 30, 2026
NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models
VLN
LLM
navigation
Sep 30, 2026
WebRollback: Enhancing Web Agents with Explicit Rollback Mechanisms
web-agent
LLM
Sep 30, 2026
Reward Hacking in Reinforcement Learning
agentic-RL
RL
LLM
Sep 30, 2026
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
VLM
LLM
Sep 30, 2026
Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?
agentic-RL
LLM
Sep 30, 2026
The Unreasonable Effectiveness of Scaling Agents for Computer Use
computer-use
gui-agent
LLM
Sep 30, 2026
Retrieval-Mediated Defense for Memory Misevolution
agentic-RL
LLM
research-idea
Sep 30, 2026
ATLaS: Agent Tuning via Learning Critical Steps
LLM
imitation-learning
Sep 30, 2026
Progress or Regress? Self-Improvement Reversal in Post-training
agentic-RL
LLM
Sep 30, 2026
Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
deep-research
agentic-RL
LLM
Sep 30, 2026
AI Agents That Matter
LLM
web-agent
Sep 30, 2026
Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
agentic-RL
LLM
auto-research
Sep 30, 2026
Counterfactual Probe Gating for Self-Evolution Steps
agentic-RL
LLM
research-idea
Sep 30, 2026
Scaling Long-Horizon LLM Agent via Context-Folding
deep-research
agentic-RL
LLM
Sep 30, 2026
Alignment Tipping Process: How Self-Evolution Pushes LLM Agents Off the Rails
agentic-RL
LLM
Sep 30, 2026
Toward Systems Foundations for Agentic Exploration
computer-use
LLM
Sep 30, 2026
Rho-1: Not All Tokens Are What You Need
LLM
Sep 30, 2026
Tree Search for LLM Agent Reinforcement Learning (Tree-GRPO)
agentic-RL
LLM
Sep 30, 2026
A Survey on Self-Evolution of Large Language Models
LLM
agentic-RL