FrontierMemory
by LettaBenchmarking agent memory on verifiable tasks
Leaderboard
More harnesses coming soon| Rank | Harness | Model | FrontierMemory score | |
|---|---|---|---|---|
| 1 | Letta Dream | Opus 5.5 (xhigh) | 57.1% | |
| 2 | Claude Managed Agents | Opus 5.5 (high) | 42.9% | |
| 3 | OpenAI Agents SDK | GPT-6.1 Sol (xhigh) | 28.6% | |
| — | No memory | — | 14.3% | |
Pareto frontier
More harnesses coming soonVarying the level of test-time compute per harness / model, while fixing the sleep-time compute (compute applied to memory generation).
The first memory-native benchmark
Existing memory benchmarks measure recall of past facts. FrontierMemory is the first benchmark that measures the quality of memory. Agents are given access to trajectories of past experience and tasked to generate memory, which is then scored by how much it helps an agent solve downstream tasks.
About FrontierMemory
-
Agents are given access to trajectories of agent experience and generate memory.
-
Memory is evaluated based on how much it helps solve downstream tasks.
Task sources
We select challenging coding and knowledge-work tasks from public benchmarks, including HarborIndex, DeepSWE, and SkillsBench, both for source transcripts and for downstream evaluations. Some tasks are modified so solving them requires knowledge available only in the source transcripts.
Sample Tasks
More tasks coming soonAdd async initialization support to the Awilix library. Find the release tier in project metadata, then use the team’s prior policy to choose the version bump.
BigCodeBenchWord2Vec PipelineBuild a Word2Vec component. Earlier attempts demonstrate the missing interface, preprocessing, and output artifact contract.
SkillsBenchDAPT Intrusion DetectionCompute network statistics from a DAPT2020 packet capture. Prior packet analyses reveal counting, entropy, and threat-detection pitfalls to avoid.
Submit an agent harness or model
To submit a harness, it must have a continual learning or dreaming system it uses to create memories or skills between sessions. A model can also be scored on its own, inside a supported harness (such as Letta Code) which stays fixed as the control.
Sign up for leaderboard updates
Get notified when new harnesses and models are added to FrontierMemory.