Help build FrontierMemory

FrontierMemory grows one task bundle at a time. A task is worth adding when a fresh agent fails it without memory, and the prior trajectories contain what the agent needed: a hidden contract, a user's standing preference, a procedure, or a better solution to improve on.

What a task bundle contains

  • The exact instruction shown to the evaluation actor.
  • A reproducible environment and an executable verifier that scores the actor's artifact, not its explanation.
  • Source trajectories: the prior attempts a memory system may learn from, with their verifier results.
  • Ground truth kept out of the dream input: reference memory or skills, and an oracle solution.
  • Controlled baselines from the fixed actor with no memory and with the reference memory, so the task's floor and ceiling are known before anyone submits.

Where tasks come from today

The current set draws on Harbor-Index audited failures, SkillsBench, and DeepSWE-derived preference profiles. Every finalized bundle is reviewed by at least two people before it appears on the tasks page.

Get in touch

Task proposals and questions go to [email protected].