DeepRead
Search
Search
Dark mode
Light mode
Explorer
Popular Tags
#VLM
#agentic-RL
#gui-agent
#web-agent
#LLM
#VLA
#computer-use
#task-planning
#manipulation
#imitation-learning
Tag: evaluation
7 items with this tag.
Aug 20, 2026
How Benchmarks Mis-Score Computer-Use Agents
computer-use
gui-agent
web-agent
benchmark
evaluation
LLM
Aug 20, 2026
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
web-agent
benchmark
multimodal
live-web
evaluation
Aug 20, 2026
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
web-agent
benchmark
multimodal
visual-grounding
evaluation
Aug 20, 2026
GAIA: a benchmark for General AI Assistants
web-agent
benchmark
deep-research
tool-use
evaluation
Aug 20, 2026
An Illusion of Progress? Assessing the Current State of Web Agents
web-agent
benchmark
evaluation
llm-as-judge
reliability
Aug 20, 2026
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
web-agent
benchmark
deep-research
information-seeking
evaluation
Aug 20, 2026
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
web-agent
benchmark
enterprise
knowledge-work
evaluation