Added Knows: a Google Workspace benchmark for web agents (Docs, Sheets, Slides) - #397
Open
farhanishmam wants to merge 8 commits into
Open
farhanishmam wants to merge 8 commits into
farhanishmam wants to merge 8 commits into
Conversation
…s, Slides) Knows evaluates browser agents on long-horizon document authoring in Google Workspace. An agent drives a real browser to build a Doc, Sheet, or Slide deck from a natural-language goal, and the result is graded checkpoint-by-checkpoint via the Workspace APIs, yielding a fractional reward rather than binary success. 110 tasks across 22 families and 3 Workspace apps (25 docs / 45 sheets / 40 slides), exposed as a single `knows` split. As with TimeWarp, the task code lives in an external package (`pip install browsergym-knows`); this change is config wiring only: - `knows` action subset in core, adding file upload and direct mouse control on top of the usual bid actions, both needed to edit Workspace canvases - `knows` benchmark backend, action set args, and metadata CSV - lazy `import browsergym.knows` in `prepare_backend` and `_get_env_name` - `knows` extra, dev requirement, and expected task count in the tests Knows requires a Google account and a Google Cloud service account; see the benchmark README for setup. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
recursix
previously approved these changes
Jul 28, 2026
recursix
left a comment
Collaborator
There was a problem hiding this comment.
Super clean! Thanks :)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Installation and dependency paths currently reference an unavailable browsergym-knows package.
Review effort: Lite
Findings: 2
Open (3)
What changed in this PR
Adds the Knows Google Workspace benchmark integration for Docs, Sheets, and Slides.
Changes:
- Registers 110 benchmark tasks and configurations.
- Adds Knows-specific actions and backend preparation.
- Updates metadata, documentation, dependencies, and tests.
| File | Summary |
|---|---|
tests/experiments/test_benchmark.py |
Verifies the Knows benchmark task count. |
README.md |
Documents Knows setup and usage; installation references require correction. |
dev/requirements.txt |
Adds the Knows dependency, which is currently unavailable from PyPI. |
browsergym/experiments/src/browsergym/experiments/loop.py |
Adds lazy Knows environment registration. |
browsergym/experiments/src/browsergym/experiments/benchmark/utils.py |
Adds Knows backend preparation. |
browsergym/experiments/src/browsergym/experiments/benchmark/metadata/knows.csv |
Defines Knows task metadata. |
browsergym/experiments/src/browsergym/experiments/benchmark/configs.py |
Configures actions and benchmark execution. |
browsergym/experiments/src/browsergym/experiments/benchmark/base.py |
Adds Knows backend support. |
browsergym/experiments/pyproject.toml |
Adds Knows extras that currently reference an unavailable PyPI package. |
browsergym/core/src/browsergym/core/action/highlevel.py |
Defines the Knows action subset. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
The previous link targeted a private repository and returned 404 for anyone without access. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The other entries still use the pre-existing `browsergym-experiment` typo, which ServiceNow#403 fixes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The browsergym meta package pins browsergym-webarena-verified to the repo version, so make install failed after the 0.14.4 bump because that version is not on PyPI yet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator
Author
|
@recursix @chrish42 the earlier approval was dismissed when I pushed new commits, so this needs another review. Since the approval:
The other failing jobs were caused by |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


Summary
Adds Knows, a Google Workspace benchmark for web agents: 110 document-authoring tasks across Google Docs (25), Sheets (45), and Slides (40). The tasks ship in the
browsergym-knowspackage (source, project page). This PR registers the benchmark in BrowserGym the same way TimeWarp is registered.Changes
browsergym/core: newknowsaction subset. On top of the usual bid actions, the tasks need file upload and coordinate-based mouse actions to edit the Docs/Sheets/Slides canvases.browsergym/experiments:knowsbenchmark inconfigs.py(task list inmetadata/knows.csv, multi-tab,max_steps=120)knowsbackend inbase.pyandutils.pyloop.pybrowsergym/experiments/pyproject.toml: newknowsextra, also included inall.README.md: Knows added to the benchmark list, with install, setup, and usage sections.dev/requirements.txt,tests/experiments/test_benchmark.py: installsbrowsergym-knowsfor development, and the test checks the task count.Makefile:make installnow also installsbrowsergym/webarena_verifiedfrom source. This is unrelated to Knows. After the 0.14.4 version bump, theagentlabjob failed on every branch: thebrowsergymmeta package requiresbrowsergym-webarena-verified==0.14.4, and that version isn't on PyPI yet.Setup
Running the tasks requires a Google account and a Google Cloud service account. See the setup guide.
Testing
tests/experiments/test_benchmark.py(test_build_benchmarks,test_benchmark_subset*,test_dependency_graphs) passes locally on Python 3.12 with the uv version CI uses (0.9.17).