Show HN: Agent Memory Leaderboard – first public results for AI memory systems

(agentmemoryleaderboard.ai)

3 points | by IreneAI 10 hours ago ago

1 comments

  • claudiusa 4 hours ago ago

    Are the published numbers single-run or averaged, and which model does the judging? With LLM-as-judge scoring I would expect a couple of points of run-to-run noise, which does not matter for the top spot but matters a lot for the middle of the table.