Cook Research

Better at the work.
Built on evidence.

An AI teammate needs to do two things well: take action and remember the context. Explore the research behind Cook’s browser and memory.

03

Live-web benchmark suites

98.5%

Top-five memory recall

02

Research papers to download

Cook evaluations: August 2026. Memory recall measured on 476 answerable LongMemEval-S questions.

01Browser agents

An agent that works where you work.

The web is where the work happens. We tested Cook’s desktop browser on real websites: finding information, navigating complex pages, and carrying a task through to completion.

Task success · %Scale 0–100
  1. AsidePublished · 300 tasks
    99.0%
  2. Browser UseReference benchmark panel
    97.7%
  3. Cook BrowserDesktop · adjusted cohort
    93.4%
  4. GPT-5.4Reference benchmark panel
    92.8%
  5. Claude Opus 4.8Reference benchmark panel
    84.0%
  6. ChatGPT AtlasReference benchmark panel
    70.0%

Cook: 268/287 after 13 impossible tasks were excluded; 268/300 (89.3%) across the full cohort. Merged baseline and targeted reruns. Aside: 297/300. Different runs and scoring policies; published reference, not a controlled head-to-head. Other competitors reproduce the reference benchmark panel; evaluation setups differ.

Aside’s published benchmark results ↗

Persistence is part of intelligence.

The browser research led to better long-running task persistence, explicit tab management, and more reliable handling of final answers and rich browser results.

02Long-term memory

The right context. Even conversations later.

A useful teammate remembers what matters. We tested how reliably Cook finds the past exchange that answers a new question, even when it is buried in unrelated conversation history.

Questions with a relevant result in the top 5 · %Scale 0–100
  1. Cook MemoryAnswer-bearing excerpts
    98.5%
  2. agentmemoryPublished · gold sessions
    95.2%

Cook: 476 answerable questions, scored on answer-bearing excerpts. agentmemory: published 500-question session-retrieval evaluation. Different retrieval units and cohorts; reference comparison, not an identical-harness rerun. Retrieval is separate from final-answer accuracy.

agentmemory’s published methodology ↗

Finding the topic is only the beginning.

Combining keyword and semantic search, then ranking for the answer itself, lifted Cook’s top-five recall from 83.6% to 98.5% in the recorded evaluation.

The result matters.
So does the method.

Our papers document the tasks, the scoring rules, and the changes that improved the system. Browser results include both adjusted and full-cohort scores. Memory results separate finding evidence from answering correctly.

These are Cook-authored technical reports on recorded evaluations. External scores are attributed to their publishers; differences in task sets, retrieval units, and grading are explained beside the comparisons and in each paper.

Download chart data (JSON) ↓

Put the research to work.

Meet the AI teammate built to move your business forward.

Get started with Cook