01Browser agents
An agent that works where you work.
The web is where the work happens. We tested Cook’s desktop browser on real websites: finding information, navigating complex pages, and carrying a task through to completion.
- 99.0%AsidePublished · 300 tasks
- 97.7%Browser UseReference benchmark panel
- 93.4%Cook BrowserDesktop · adjusted cohort
- 92.8%GPT-5.4Reference benchmark panel
- 84.0%Claude Opus 4.8Reference benchmark panel
- 70.0%ChatGPT AtlasReference benchmark panel
Cook: 268/287 after 13 impossible tasks were excluded; 268/300 (89.3%) across the full cohort. Merged baseline and targeted reruns. Aside: 297/300. Different runs and scoring policies; published reference, not a controlled head-to-head. Other competitors reproduce the reference benchmark panel; evaluation setups differ.
- 94.4%Cook BrowserDesktop · adjusted subset
- 93.0%AsidePublished · 100 tasks
- 89.5%Browser UseReference benchmark panel
- 80.0%Claude Fable 5Reference benchmark panel
- 68.0%GPT-5.5Reference benchmark panel
- 58.0%Gemini 3.6 FlashReference benchmark panel
Cook: 17/18 after two invalid frozen answers were excluded; 17/20 (85.0%) on the frozen subset. Aside: 93/100 on the broader benchmark. Different sample sizes; these are published reference results, not a controlled ranking. Other competitors reproduce the reference benchmark panel; evaluation setups differ.
- 88.8%AsidePublished · rubric items
- 86.25%Cook BrowserDesktop · adjusted points
- 70.0%Browser UseReference benchmark panel
- 60.8%WebWrightReference benchmark panel
- 44.5%Claude Opus 4.6Reference benchmark panel
- 33.5%GPT-5.4Reference benchmark panel
Cook: 34.5/40 rubric points after five unavailable tasks were excluded; strict score 39/47 (82.98%). Aside: 1,050/1,182 rubric items across 200 tasks. Cohorts and rubric aggregation differ; these are context, not equivalent evaluations. Other competitors reproduce the reference benchmark panel; evaluation setups differ.
Persistence is part of intelligence.
The browser research led to better long-running task persistence, explicit tab management, and more reliable handling of final answers and rich browser results.