Memory & Sessions

Session persistence, PostgreSQL storage, and semantic long-term memory

Tutti has four kinds of memory:

  • Session memory — conversation history within a single session (short-term)
  • Semantic memory — facts the agent remembers about its own work, across sessions (long-term, agent-scoped)
  • User memory — per-fact entries attached to an end-user identifier, scoped across agents (long-term, user-scoped)
  • User model — one rolling LLM-summarised profile per end-user — summary, preferences, ongoing_projects — that complements the per-fact user_memory store

The four are independent: enable whichever fit the agent. semantic and the user_* stores share no rows; user_model summarises user_memory but does not replace it.

Session memory

When you call tutti.run(), the runtime creates a session automatically. The session stores the full message history.

const r1 = await tutti.run("assistant", "My name is Alice.");
// r1.session_id = "abc-123"

const r2 = await tutti.run("assistant", "What's my name?", r1.session_id);
// r2.output = "Your name is Alice."

Without a session_id, the agent starts fresh with no memory.

In-memory store (default)

Sessions live in memory. They’re lost when the process exits.

const tutti = new TuttiRuntime(score);
// sessions are in-memory by default

PostgreSQL store

Sessions persist across process restarts:

npx tutti-ai add postgres
DATABASE_URL=postgres://user:pass@localhost:5432/tutti
export default defineScore({
  provider: new AnthropicProvider(),
  memory: { provider: "postgres" },
  agents: { /* ... */ },
});

Use the async factory — it creates the tutti_sessions table on first run:

const tutti = await TuttiRuntime.create(score);

The table schema:

ColumnTypeDescription
idTEXT PRIMARY KEYSession UUID
agent_nameTEXTAgent that owns the session
messagesJSONBFull conversation history
created_atTIMESTAMPTZSession creation time
updated_atTIMESTAMPTZLast update time

:::tip For production, always use TuttiRuntime.create() instead of new TuttiRuntime() — it handles the async database initialization. :::

Semantic memory (long-term)

Semantic memory lets agents remember facts across sessions. When a user tells the coder agent “I prefer 2-space indentation”, the agent remembers this in the next session.

Enable it per agent

{
  coder: {
    name: "Coder",
    system_prompt: "You are a TypeScript developer.",
    memory: {
      semantic: {
        enabled: true,
        max_memories: 5,             // inject up to 5 relevant memories (default)
        inject_system: true,         // append to system prompt (default)
        curated_tools: true,         // expose remember/recall/forget tools to the agent (default)
        max_entries_per_agent: 1000, // LRU cap per agent (default)
      },
    },
    voices: [new FilesystemVoice()],
    permissions: ["filesystem"],
  },
}

How it works

  1. Before each LLM call, the runner searches semantic memory using the user’s input as a query
  2. The top N relevant memories are appended to the system prompt:
    Relevant context from previous sessions:
    - User prefers 2-space indentation
    - Project uses ESM modules
  3. The LLM sees these as context and can act on them

Storing memories from tools

Tools receive context.memory helpers when semantic memory is enabled:

execute: async (input, context) => {
  // Store a fact
  await context.memory?.remember("User prefers dark mode");

  // Search for relevant memories
  const prefs = await context.memory?.recall("UI preferences");
  // → [{ id: "abc", content: "User prefers dark mode" }]

  // Delete a memory
  await context.memory?.forget("abc");

  return { content: "Preferences updated." };
}

Agent-curated memory tools

When curated_tools is on (the default), the runtime exposes remember, recall, and forget as Tools the model itself can call across turns. Entries written by the model are tagged source: "agent", and a per-agent cap (max_entries_per_agent, default 1000) evicts the least-recently-used entry first when the cap is reached. Both surfaces — the context.memory helpers above and the agent-callable tools — share one enforcement pipeline, so the cap, LRU eviction, and memory:write / memory:read / memory:delete events fire exactly once per logical operation.

Subscribe to the events to observe what the agent is curating:

tutti.events.on("memory:write", (e) => {
  console.log(`agent ${e.agent_name} stored ${e.entry_id} (${e.source})`);
});

A two-turn end-to-end example lives at examples/curated-memory.ts.

Storing memories from your code

Access the semantic memory store directly on the runtime:

const tutti = new TuttiRuntime(score);

await tutti.semanticMemory.add({
  agent_name: "coder",
  content: "User prefers 2-space indentation",
  metadata: { source: "onboarding" },
});

const memories = await tutti.semanticMemory.search(
  "code style",
  "coder",
  5,
);

How search works (v1)

The InMemorySemanticStore uses keyword overlap scoring — no embeddings needed. It tokenises the query and each stored entry into word sets, scores by overlap ratio, and returns the top N.

This is simple and predictable. Future versions will support embedding-based search via custom SemanticMemoryStore implementations.

Memory lifecycle

MethodWhat it does
store.add({ agent_name, content, metadata })Store a new memory
store.search(query, agent_name, limit)Search by keyword overlap
store.delete(id)Delete one memory
store.clear(agent_name)Delete all memories for an agent

:::caution[Migrating from < v0.25.0] The agent-level semantic_memory field has moved to memory.semantic. Update your scores:

agents: {
  coder: {
-   semantic_memory: { enabled: true, max_memories: 5 },
+   memory: { semantic: { enabled: true, max_memories: 5 } },
  }
}

The shape of the inner object is unchanged. :::

User model (dialectic profile)

memory.user_model is a per-user rolling profile — summary, preferences, ongoing_projects — re-summarised by an LLM every N turns. It complements (does not replace) user_memory: per-fact entries continue to live in the UserMemoryStore exactly as before, and the consolidator reads from that store to produce a holistic profile that gets injected into the system prompt above the existing “What I remember about you:” block.

{
  assistant: {
    memory: {
      user_memory: { enabled: true, auto_extract: true },
      user_model: {
        enabled: true,
        every_n_turns: 20,             // consolidate every N turns (default 20)
        recent_memory_limit: 50,        // entries fed to the consolidator (default 50)
        // consolidation_model: "claude-haiku-4-20250514", // optional small/cheap model
      },
    },
  },
}

// Pass the user_id so the profile is scoped correctly:
await tutti.run("assistant", "Hi again", undefined, { user_id: "alice" });

Behaviour:

  • The runtime injects the profile into the system prompt at run start. Empty/bootstrap profiles do not inject — no dead weight on brand-new users.
  • Consolidation is fire-and-forget: failures are caught, logged, and never crash the run. The previous profile is preserved on JSON parse / schema-validation failure.
  • Empty profiles are skipped by the consolidator entirely — no LLM cost on a brand-new user_id.
  • The consolidation prompt explicitly forbids storing sensitive data (passwords, API keys, SSNs, payment-card numbers, government identifiers).
  • Subscribe to user_model:consolidated for telemetry / sync to a downstream system.
tutti.events.on("user_model:consolidated", (e) => {
  console.log(`profile updated for ${e.user_id} (turns since last: ${e.turns_since_last})`);
});

The UserModelStore interface is pluggable. InMemoryUserModelStore ships in v0.25.0; a Postgres backend will land behind the same interface (mirroring the UserMemoryStore factory pattern). Tests can swap in a mock via runtime.setUserModelStore(...).

A 25-turn end-to-end demonstration with a mock LLM provider lives at examples/user-model.ts — it shows the consolidator triggering at every-5-turn cadence, the profile evolving, and the final dump.

Multi-turn sessions

Each run() is one turn (which may involve multiple LLM calls if tools are used). Across multiple run() calls with the same session, the full history accumulates:

const r1 = await tutti.run("coder", "Create hello.ts");
const r2 = await tutti.run("coder", "Add tests for it", r1.session_id);
// r2 has full context of r1

Token usage

Usage is reported per run(), not per session:

result.usage.input_tokens   // tokens sent to the LLM
result.usage.output_tokens  // tokens received
result.turns                // agentic loop iterations

:::tip Long sessions accumulate large message histories, increasing input token counts. Start fresh sessions for unrelated tasks, and use semantic memory for facts that should persist. :::

Edit this page on GitHub →