FrontierMemory

by Letta

Benchmarking agent memory on verifiable tasks

Run the benchmark

Leaderboard

More harnesses coming soon
RankHarnessModelFrontierMemory score
1 Letta Dream Opus 5.5 (xhigh) 57.1%
2 Claude Managed Agents Opus 5.5 (high) 42.9%
3 OpenAI Agents SDK GPT-6.1 Sol (xhigh) 28.6%
— No memory — 14.3%

Pareto frontier

More harnesses coming soon

Varying the level of test-time compute per harness / model, while fixing the sleep-time compute (compute applied to memory generation).

FrontierMemory score
0%20%40%60%80%$0.20$0.50$1$2$5$10 lowmediumhighxhighLetta Dreammax lowmediumhighxhighClaude Managed Agentsmax lowmediumhighxhighOpenAI Agents SDKmax lowmediumhighxhighNo memorymax Average cost per task (log scale)

The first memory-native benchmark

Existing memory benchmarks measure recall of past facts. FrontierMemory is the first benchmark that measures the quality of memory. Agents are given access to trajectories of past experience and tasked to generate memory, which is then scored by how much it helps an agent solve downstream tasks.

About FrontierMemory

  1. Agents are given access to trajectories of agent experience and generate memory.

  2. Memory is evaluated based on how much it helps solve downstream tasks.

Task sources

We select challenging coding and knowledge-work tasks from public benchmarks, including HarborIndex, DeepSWE, and SkillsBench, both for source transcripts and for downstream evaluations. Some tasks are modified so solving them requires knowledge available only in the source transcripts.

Sample Tasks

More tasks coming soon
See all tasks →

Submit an agent harness or model

To submit a harness, it must have a continual learning or dreaming system it uses to create memories or skills between sessions. A model can also be scored on its own, inside a supported harness (such as Letta Code) which stays fixed as the control.