DeepRead
Search
Search
Dark mode
Light mode
Explorer
Popular Tags
#VLM
#agentic-RL
#gui-agent
#web-agent
#LLM
#VLA
#computer-use
#task-planning
#manipulation
#imitation-learning
Tag: benchmark
9 items with this tag.
Aug 20, 2026
How Benchmarks Mis-Score Computer-Use Agents
computer-use
gui-agent
web-agent
benchmark
evaluation
LLM
Aug 20, 2026
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
web-agent
benchmark
multimodal
live-web
evaluation
Aug 20, 2026
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
web-agent
benchmark
multimodal
visual-grounding
evaluation
Aug 20, 2026
GAIA: a benchmark for General AI Assistants
web-agent
benchmark
deep-research
tool-use
evaluation
Aug 20, 2026
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
web-agent
safety
prompt-injection
benchmark
security
Aug 20, 2026
An Illusion of Progress? Assessing the Current State of Web Agents
web-agent
benchmark
evaluation
llm-as-judge
reliability
Aug 20, 2026
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
web-agent
benchmark
deep-research
information-seeking
evaluation
Aug 20, 2026
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
web-agent
benchmark
enterprise
knowledge-work
evaluation
Jun 24, 2026
GUI Environment 近期工作调研与 Agent-Facing Runtime 选题更新
report
gui-agent
computer-use
environment
benchmark
verifier
research-strategy