MIT CSAIL researchers introduced JAZ, an agent framework where the model treats its own history as code variables it can directly inspect and manipulate. On the StuLife benchmark, JAZ scored 70 % using the GPT-5.4 nano model, outperforming Letta’s 62 % at roughly half the cost. On the AppWorld benchmark for self-improvement, JAZ achieved 74 %, beating ACE by 4 percentage points while remaining cheaper. The core idea is that rather than equipping a language model with an external filing cabinet, JAZ exposes the agent’s history and prompt as variables in a code environment. The researchers released the framework and evaluation code on GitHub so other scientists can verify the results.
Source: Read the original article

