Evaluating long-term memory in LLM agents with the MERIT benchmark and using
Based on reporting by arXiv AI & Computer Science. Research, structure, and fact-checking by Groundwork.
for Large Language Model (LLM) agents, long-term memory is often touted as a key feature that enables them to learn and adapt over time. However, the actual value of long-term memory in these agents is not well understood. In this article, we'll explore into a recent study that sheds light on the effectiveness of long-term memory in tool-using LLM agents, with a focus on its cost-aware evaluation.
Long-term memory in LLM agents refers to the ability of the model to retain and recall information over an extended period. This is in contrast to short-term memory, which is typically limited to a single conversation or task. Long-term memory enables LLM agents to learn from their experiences and adapt to new situations, making them more effective in complex tasks.
Traditionally, long-term memory in LLM agents is evaluated using conversational recall benchmarks such as LoCoMo and LongMemEval. These benchmarks measure a model's ability to answer questions based on its dialogue history. However, these benchmarks have a significant limitation: they do not account for the actual impact of long-term memory on the model's performance.
To address the limitations of traditional benchmarks, researchers have developed a new evaluation framework called MERIT (Memory Evaluation for Realistic Instrumented Tasks). MERIT provides a comprehensive assessment of long-term memory in LLM agents, taking into account the model's performance, cost, and energy consumption.
The MERIT benchmark consists of episodic tool-use tasks in three domains: robotics, chemistry, and biology. These tasks are designed to simulate real-world scenarios where LLM agents need to use tools to complete tasks. The MERIT using is a software framework that enables researchers to easily implement and evaluate LLM agents on these tasks.
The MERIT benchmark has several key features that set it apart from traditional benchmarks:
The researchers conducted an extensive experiment using the MERIT benchmark and using. They evaluated the performance of three LLM models: GPT-4.1, Claude Haiku 4.5, and Claude Sonnet 5. The results showed that long-term memory significantly improves the performance of LLM agents, particularly in tasks that require the model to recall previously learned information.
The experiment revealed several key findings:
The MERIT benchmark and using provide a comprehensive evaluation framework for long-term memory in LLM agents. The experimental results demonstrate the significant impact of long-term memory on the performance of LLM agents, particularly in tasks that require the model to recall previously learned information. However, the results also highlight the limitations of traditional embedding retrieval and the importance of update-on-write stores in achieving high performance.
“The MERIT benchmark and using provide a significant advancement in the evaluation of long-term memory in LLM agents. However, further research is needed to fully understand the impact of long-term memory on the performance of LLM agents in real-world scenarios.”
The MERIT benchmark is a comprehensive evaluation framework for long-term memory in LLM agents, taking into account the model's performance, cost, and energy consumption.
The MERIT benchmark has several key features, including episodic tool-use tasks, a difficulty ladder, controlled memory corruption, and full token and dollar metering.
The experimental results show that long-term memory significantly improves the performance of LLM agents, particularly in tasks that require the model to recall previously learned information.
The key findings include memory lifts dependent-task success, embedding retrieval collapses unpredictably, agents act on a correctly retrieved value only 55% of the time, update-on-write stores remain at 0.70-1.00, and the hybrid is worse than the fact store alone.

Europe is moving forward with its ambitious Envision mission to Venus, despite a setback from NASA.

Prolonged laptop use has been linked to a range of health risks, including myopia, headaches, and eye strain.

Apple Watch listening features explained: data privacy, security, and how to use them safely.
Explore related evidence-based investigations, decision tools, and entity breakdowns:
Contextual evidence and verified documentation referenced in this research guide
Groundwork enforces a strict, independent verification standard. All claims and benchmark figures in this guide are cross-referenced against the primary documentation and regulatory registries listed below:
Sofia Reyes (2026). When Does Memory Help? A Cost-Aware: Tested & Reviewed. Groundwork. Retrieved from https://gworky.com/article/evaluating-long-term-memory-in-tool-using-llm-agents
Originally published at https://gworky.com/article/evaluating-long-term-memory-in-tool-using-llm-agents — Groundwork Evidence-Based Research.
Uncover forgotten seat licenses, redundant cloud services, and recurring overhead.
Evidence-Based • Free Open Access • Zero Guesswork
Connect your brand with over 50,000 monthly decision-makers seeking verified guidance in finance, health, and tech.
Compare estimated monthly cost across leading AI models based on your token usage.
tech
tech
tech
techEvaluated for performance, privacy protocols, and pricing transparency.
| Solution | Key Benchmark | Pricing | Verdict & Access |
|---|---|---|---|
NordVPNEditor Pick via Nord Security | Audited WireGuard no-logs protocol | $3.39/mo | |
ExpressVPN via Express Technologies | Lightway protocol, RAM-only servers | $6.67/mo | |
Cloudflare WARP+ via Cloudflare Inc. | Fast Argo edge routing | $4.99/mo | Reference Benchmark |
Smart Home & Digital Privacy Analyst
Smart home and digital privacy analyst focused on data ownership, device security, and power efficiency of AI utilities and gadgets.
This guide underwent secondary data verification to confirm primary source integrity, calculation formulas, and regulatory compliance before publication.