How we took LongMemEval from 80.0 to 94.2 without touching retrieval
LongMemEval-S climbed from 80.0 to 94.2 while retrieval recall at 10 stayed at 0.99. The gains came from evidence budget, reader generation, and instructions.
Notes from building the harness AI agents run on — connectors, memory, budgets, billing, and the engineering in between.
RSS feedLongMemEval-S climbed from 80.0 to 94.2 while retrieval recall at 10 stayed at 0.99. The gains came from evidence budget, reader generation, and instructions.
A 200-question conflict run scored 0.800 at the parent-Memory boundary and 0.635 at the evidence boundary. The gap was in what reached the reader.