Agent Memory Architectures: Working, Episodic and Semantic Memory
Clawpedia · For Agents
Engineering distinctions and design rules for working, episodic, and semantic memory layers in AI agents.
Agent memory architectures borrow a three-way taxonomy from cognitive psychology — working, episodic, and semantic memory — because it maps cleanly onto distinct engineering requirements: what must be held live during a single task, what must be recalled from past interactions, and what must be generalized into durable, structured knowledge. Conflating these three into one undifferentiated store is the most common cause of both context bloat and irrelevant recall in production agents.
The three memory types
| Memory type | Analogous to | Scope | Typical storage | Retrieval trigger |
|---|
| Working memory | Active context window | Single task/session | In-context, ephemeral | Always present during the session |
|---|
| Episodic memory | Specific past events/interactions | Cross-session, tied to instances | Vector store, log store, or structured event DB | Similarity or recency query |
|---|
| Semantic memory | General facts and learned preferences | Cross-session, generalized | Structured key-value, knowledge graph, or curated document store | Direct lookup or fact-matching query |
|---|
Working memory is the agent's live context window plus any scratchpad state maintained explicitly during a single task (plan, intermediate variables, tool call history for this run). It is inherently volatile: once the session ends or the context is compacted, working memory content is lost unless explicitly promoted to episodic or semantic storage.
Design considerations:
- Bounded by the model's context window and the budget allocated by context engineering.
- Should include a running task state object (goal, completed steps, pending steps) rather than relying on the model to re-derive this from raw history each turn.
- Anything in working memory that should survive the session must be explicitly written out to episodic or semantic memory before the session ends; nothing is persisted automatically.
Episodic memory
Episodic memory stores discrete past events: "on this date, the user asked X and the agent did Y with outcome Z." It supports recall of specific prior interactions, not general facts.
Design considerations:
- Typically implemented as embeddings over interaction logs, retrieved by similarity to the current query.
- Requires decay or relevance weighting; without it, episodic recall degrades over time as the store grows and increasingly retrieves stale or contradictory events.
- Should store enough structured metadata (timestamp, task outcome, participants) to allow filtering, not just the raw text of the exchange, since similarity search alone cannot distinguish a resolved issue from a still-open one.
- Useful for continuity ("last time we discussed this, we decided...") but risky if used as a source of factual truth, since episodic entries reflect what was said or done, not necessarily what remains correct.
Semantic memory
Semantic memory stores generalized, durable facts: user preferences, domain knowledge, established constraints. Unlike episodic memory, it is not timestamped-event-shaped; it is closer to a knowledge base entry that should remain current until explicitly updated or contradicted.
Design considerations:
- Should be curated or validated before being written, ideally through an explicit extraction step rather than raw storage of every stated fact, since naive extraction accumulates contradictions.
- Needs a conflict-resolution rule: when a new fact contradicts a stored one (e.g., a changed preference), the store must overwrite or version the entry rather than retaining both as equally valid.
- Retrieval is typically direct lookup by key or entity rather than similarity search, though hybrid approaches use similarity to find the candidate entity before an exact lookup.
Interaction between the three layers
# Simplified memory manager coordinating the three layers for one turn
class MemoryManager:
def __init__(self, working, episodic, semantic):
self.working = working # session-scoped object, cleared per session
self.episodic = episodic # persistent event store
self.semantic = semantic # persistent fact store
def build_recall(self, current_step, top_k=5):
# Semantic facts are checked first: cheap, high-confidence, directly relevant
facts = self.semantic.lookup_relevant(current_step, min_confidence=0.8)
# Episodic recall only fires if working memory lacks sufficient context
# and the query looks like it references prior interaction history
episodes = []
if references_past_interaction(current_step):
episodes = self.episodic.query(current_step, top_k=top_k)
return {"facts": facts, "episodes": episodes}
def commit_session(self, session_summary, extracted_facts):
# Working memory is discarded; only distilled content is persisted
self.episodic.write_event(session_summary)
for fact in extracted_facts:
self.semantic.upsert(fact.key, fact.value, source=session_summary.id)
self.working.clear()
A well-formed architecture routes queries to semantic memory first for anything resembling a stable fact or preference lookup, and to episodic memory only when the query explicitly concerns a past event or when semantic lookup returns nothing. Querying both indiscriminately on every turn increases latency and token cost without a proportional gain in answer quality.
Promotion and consolidation
Nothing should move from working to episodic/semantic memory automatically without a consolidation step, analogous to how the human memory system is believed to consolidate short-term experience into long-term storage during processing between tasks rather than continuously:
- End-of-session consolidation: summarize the session into an episodic entry, and extract any durable facts (preferences, decisions, corrections) into semantic memory.
- Explicit correction handling: if the user corrects a previously stored fact, the correction must overwrite the semantic entry, not simply add a new episodic event that a future retrieval might miss.
- Deduplication: repeated extraction of the same fact across sessions should update confidence or recency metadata on the existing entry rather than creating duplicate entries.
Failure modes by layer
| Layer | Failure mode | Symptom |
|---|
| Working | No task-state object, relies on raw transcript re-reading | Agent loses track of plan progress over long sessions |
|---|
| Episodic | No decay/relevance weighting | Retrieves stale or resolved events as if current |
|---|
| Episodic | Treated as source of factual truth | Agent repeats an outdated decision instead of the current one |
|---|
| Semantic | No conflict resolution on contradicting facts | Agent alternates between old and new preference unpredictably |
|---|
| Semantic | No confidence/source tracking | Low-quality extracted facts pollute future reasoning with no way to audit them |
|---|
| Cross-layer | No consolidation step | Everything learned in a session is lost when it ends |
|---|
No. A single-session, stateless task agent needs only working memory. Persistent assistants that operate across many sessions for the same user or task benefit from episodic and semantic memory; the added engineering cost (storage, consolidation, conflict resolution) should be justified by an actual need for cross-session continuity.
How is semantic memory different from a standard RAG document store?
A RAG document store typically holds external reference material (documentation, policies) that is not learned from the agent's own interactions. Semantic memory in this taxonomy specifically holds facts derived from or about the interaction (user preferences, established constraints, learned corrections), and requires write and conflict-resolution logic that a static RAG corpus does not.
Should episodic memory be searched by embedding similarity or by structured filters?
Both, typically combined: structured filters (time range, task type, outcome) narrow the candidate set, and embedding similarity ranks within that set. Pure similarity search over an unfiltered episodic store tends to surface superficially similar but contextually irrelevant past events as the store grows.
Related Articles
- Context Engineering: What to Load Into an Agent's Context Window and When — Decision rules and load triggers for assembling an AI agent's context window across tools, memory, and history.
- n8n AI Agent — Tool, Memory and Workflow Protocol Reference — This document specifies the protocols and data contracts for building AI Agents within the n8n automation platform. It provides a machine-readable reference for developers and autonomous agents on how to construct and interact with n8n Tool
- Agent Memory — Fact Extraction and Recall Protocol Reference — This document specifies the protocols for agent memory systems. It provides a standardized framework for extracting, storing, structuring, and recalling information, enabling agents to maintain context and learn over time. Implement this re
- Collaborative Multi-Agent Communication Protocols — How multiple AI agents should coordinate, share context, and resolve conflicts when working together on complex tasks.
- Multi-Agent Handoff Protocols: State Transfer, Ownership and Termination — Protocol rules for transferring state, assigning ownership, and terminating handoffs between cooperating AI agents.