DeepRead
Search
Search
Dark mode
Light mode
Explorer
Popular Tags
#VLM
#agentic-RL
#gui-agent
#LLM
#task-planning
#web-agent
#VLA
#manipulation
#computer-use
#world-model
Tag: evaluation
7 items with this tag.
Sep 30, 2026
How Benchmarks Mis-Score Computer-Use Agents
computer-use
gui-agent
web-agent
benchmark
evaluation
LLM
Sep 30, 2026
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
web-agent
benchmark
multimodal
live-web
evaluation
Sep 30, 2026
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
web-agent
benchmark
multimodal
visual-grounding
evaluation
Sep 30, 2026
GAIA: a benchmark for General AI Assistants
web-agent
benchmark
deep-research
tool-use
evaluation
Sep 30, 2026
An Illusion of Progress? Assessing the Current State of Web Agents
web-agent
benchmark
evaluation
llm-as-judge
reliability
Sep 30, 2026
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
web-agent
benchmark
deep-research
information-seeking
evaluation
Sep 30, 2026
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
web-agent
benchmark
enterprise
knowledge-work
evaluation