DeepRead

Popular Tags
#VLM#agentic-RL#gui-agent#web-agent#LLM#VLA#computer-use#task-planning#manipulation#imitation-learning

Tag: benchmark

9 items with this tag.

  • Aug 20, 2026

    How Benchmarks Mis-Score Computer-Use Agents

    • computer-use
    • gui-agent
    • web-agent
    • benchmark
    • evaluation
    • LLM
  • Aug 20, 2026

    WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

    • web-agent
    • benchmark
    • multimodal
    • live-web
    • evaluation
  • Aug 20, 2026

    VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

    • web-agent
    • benchmark
    • multimodal
    • visual-grounding
    • evaluation
  • Aug 20, 2026

    GAIA: a benchmark for General AI Assistants

    • web-agent
    • benchmark
    • deep-research
    • tool-use
    • evaluation
  • Aug 20, 2026

    WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks

    • web-agent
    • safety
    • prompt-injection
    • benchmark
    • security
  • Aug 20, 2026

    An Illusion of Progress? Assessing the Current State of Web Agents

    • web-agent
    • benchmark
    • evaluation
    • llm-as-judge
    • reliability
  • Aug 20, 2026

    BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

    • web-agent
    • benchmark
    • deep-research
    • information-seeking
    • evaluation
  • Aug 20, 2026

    WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

    • web-agent
    • benchmark
    • enterprise
    • knowledge-work
    • evaluation
  • Jun 24, 2026

    GUI Environment 近期工作调研与 Agent-Facing Runtime 选题更新

    • report
    • gui-agent
    • computer-use
    • environment
    • benchmark
    • verifier
    • research-strategy

Created with Quartz v4.5.2 © 2026

  • GitHub