Agent Memory and Persistence
Agent Memory는 AI Agent Architecture가 정의하는 orchestration 위에서, agent loop가 실패해도 진행 상태를 잃지 않게 하는 지속성(persistence) 계층이다. Agent Harness가 에이전트를 안정적으로 운영하기 위해 기대는 핵심 기능이기도 하다.
이 페이지가 답하는 질문:
- Agent가 장시간 실행 중 실패하면 어떻게 복구하는가?
- State/Thread/Checkpointer는 각각 무엇이 다른가?
- Human-in-the-loop는 memory와 어떻게 연결되는가?
핵심 모델: Durable Execution
- LangGraph는 AI agent와 응용을 위한 durable runtime을 표방한다.[1]
- 장시간 실행되는 agent가 실패하면 처음부터 재실행하는 비용과 시간이 크므로, LangGraph는 각 step마다 state를 저장하는 checkpointing으로 이 문제를 해결한다 — 재시작 시 정확히 중단 지점에서 resume한다.[1:1]
- 이 memory/persistence 기능은 단독으로 존재하지 않고, parallel 실행, streaming, human-in-the-loop, tracing/eval(LangSmith)와 함께 LangGraph를 써야 하는 이유로 묶인다.[1:2]
flowchart TD Invoke["invoke(initial state)"] Node["Node 실행"] State["State update"] CP["Checkpointer
영구 저장"] Next["다음 Super Step"] Fail["Node 실패?"] Resume["Checkpoint에서 resume"] Invoke --> Node --> State --> CP --> Next --> Node Node -->|실패| Fail --> Resume --> Node CP -.snapshot.-> Thread["Thread
(checkpoint history)"]
- 위 diagram에 등장하는 State·Node·Super Step·Checkpointer·Thread는 아래 실행 모델의 구성 요소다. 각각 어느 랩에서 정의되는지도 함께 표시한다.
- State·Node는 Nodes 랩에서 정의된다.[2]
- Super Step·Edge는 Edges 랩에서 정의된다.[3]
- Checkpointer·Thread는 Memory 랩에서 정의된다.[4]
| 용어 | 정의 |
|---|---|
| State | Graph가 운영하는 데이터. 모든 node가 공유. Python typed dict / dataclass / Pydantic model |
| Super Step | 활성 node들이 모두 완료하고 state에 값을 저장하는 단계. 병렬 실행 시 같은 super step 안의 node들은 같은 state 시점을 관찰 |
| Checkpointer | 각 super step 종료 시 state snapshot을 영구 저장하는 컴포넌트 |
| Thread | checkpoint들의 시계열 모음. 한 실행의 전체 state history |
| Node | state를 받아 update를 반환하는 함수. "node는 함수다" |
| Edge | node 간 제어 흐름. static edge(항상 이동)와 conditional edge(조건부 분기). Control은 edge를 따르지만 data는 따르지 않는다 — node는 parallel branch의 state 업데이트에도 접근 가능 |
- Thread를 구분 없이 사용하면 → 여러 사용자 세션이 같은 thread id로 충돌한다. Thread가 "한 실행의 전체 state history"라는 정의를 지키지 못하면 서로 다른 대화가 하나의 history로 뒤섞인다.
Checkpointer(short-term)와 Store(long-term): 지속성 범위 비교
- LangGraph가 구분하는 두 지속성 시스템은 범위가 다르다.
- checkpointer: 단일 thread에 속한 graph state를 checkpoint로 저장하는 short-term, thread-scoped memory.[5]
- store: graph state 밖에서 여러 thread에 걸쳐 application-defined data를 저장하는 long-term, cross-thread memory.[5:1]
| 저장 범위 | LangGraph 모델 | 적합한 정보 |
|---|---|---|
| 단일 thread | checkpointer가 저장한 graph state/checkpoint | 대화 history, 실행 재개, human-in-the-loop 진행 상태 |
| 여러 thread | store에 둔 application-defined memory | 사용자 선호, 장기 사실, 공유 knowledge |
- Checkpoint는 현재 graph execution을 재개하기 위한 state history이고, store는 그 밖에서 namespace로 관리하는 long-term data다.
- 대부분의 application은 둘 다 쓴다 — checkpointer는 현재 thread를, store는 여러 thread에 걸친 durable 정보를 함께 추적한다.[5:2]
"대화를 기억한다"는 요구만으로 모든 history를 long-term profile에 복사할 필요는 없다는 판단은 이 페이지의 해석이다.
- Checkpointer 없이 장시간 agent를 돌리면 → node 실패 시 전체를 처음부터 재실행해야 한다. checkpointer가 없다는 것은 위 표의 "단일 thread" 행이 아예 존재하지 않는다는 뜻이다.
- 영구 저장소 대신 in-memory dict만 쓰면 → 프로세스 재시작 시 checkpoint가 사라진다. checkpointer 자체는 있어도 저장 대상이 휘발성이면 short-term memory의 지속성 보장이 무너진다.
Reducer: 병렬 node의 state 충돌 해결
- 병렬 실행 시 여러 node가 같은 state key에 쓰면 기본적으로 마지막 값이 이전 값을 덮어쓴다.[3:1]
- Reducer 함수로 병합 정책을 정의할 수 있다 — 이름은 MapReduce에서 왔다.[3:2]
operator.add를 reducer로 쓰면 병렬 node들의 반환값이 리스트로 누적된다.[3:3]
from typing import Annotated
from operator import add
class State(TypedDict):
nlist: Annotated[list[str], add] # operator.add = 리스트 누적
- Custom reducer도 정의 가능하다 — state annotation의 두 번째 인자로 지정한다.[3:4]
Control은 edge를 따르지만 data는 따르지 않는다. 같은 super step의 병렬 node(B·C)는 같은 state를 보고 state는 super step이 끝날 때 갱신되므로, 다음 super step의 node(BB·CC)는 edge로 이어지지 않은 다른 branch가 쓴 state에도 접근 가능하다.[3:5] 이는 데이터 흐름을 제어 흐름과 분리해 유연한 parallel 패턴을 가능하게 하는, LangGraph의 핵심 설계 결정이라고 해석할 수 있다.
- 병렬 branch의 state 충돌을 reducer 없이 두면 → 마지막 값이 이전 값을 덮어쓴다. 위
operator.add예시처럼 명시적 reducer를 지정하지 않으면 이 기본 동작을 그대로 겪는다.
Human-in-the-loop: Suspend와 Resume
- Checkpointer가 state를 영구 저장하므로, agent 실행을 임의 시점에 suspend하고 human input을 기다릴 수 있다 — human이 응답하면 같은 thread에서 resume한다.[6]
- LLM 응답은 비결정적이므로 일부 응용은 human approval이 필요하며, LangGraph는 agent를 suspend하고 human input을 기다리는 human-in-the-loop 인터페이스를 구현한다.[1:3]
이는 "agent가 자율적으로 판단하되, 중요 결정 지점에서 사람의 승인을 구하는" bounded autonomy를 구현하는 핵심 메커니즘이라고 보인다.
- Human-in-the-loop을 프롬프트 수준 지시("요청해라")로만 구현하면 → 구조적 suspend/resume 보장이 없다. checkpointer 기반 suspend 없이는 human 응답을 기다리는 동안 state가 유실될 수 있다.
- LLM 응답 비결정성을 무시하면 → 한 번 실행한 결과를 "재현 가능하다"고 잘못 가정하게 된다. 바로 이 비결정성이 위에서 human approval을 요구하는 이유이기도 하다.
왜 durable runtime인가: 휘발성 RAM과의 대비
Agent Memory의 핵심은 state를 단기 기억(short-term memory)에만 두지 않고 영구 저장소로 지속(persist)하는 데 있다고 보인다. 위에서 정리한 내용을 다시 묶으면:
- 실패 복구 — node 실패 시 checkpoint에서 resume, 처음부터 재실행 없음
- Thread/history — 전체 실행 history 조회로 디버깅·재실행 가능
- Human-in-the-loop — suspend 후 human input 대기, resume 시 같은 thread에서 이어감
LangGraph가 "durable runtime"을 표방하는 이유는, 일반 프로그램의 메모리(RAM)가 휘발성인 것과 대비해 agent state를 휘발성에서 영구성으로 격상시켰다는 데 있다고 해석할 수 있다. 장시간 agent가 실패하면 그 시간과 비용이 통째로 손실되는 일반 프로그램과 달리, checkpoint가 있으면 마지막 checkpoint까지의 비용만 손실된다는 것이 이 모델의 핵심 이득으로 보인다.
관련
- AI Agent Architecture — Memory는 orchestration layer가 state를 유지·복구하는 메커니즘.
- Agent Harness — checkpointing은 harness가 제공하는 핵심 안정성 기능.
출처
테스트 질문
- Checkpointer가 없으면 장시간 agent 실패 시 어떤 문제가 발생하는가?
- "Control은 edge를 따르지만 data는 따르지 않는다"는 LangGraph 설계가 의미하는 바는 무엇인가?
- Reducer 함수가 병렬 node 실행에서 왜 필요한가?
- Human-in-the-loop이 checkpointing과 어떻게 연결되는가? suspend 없이 구현하면 어떤 보장을 잃는가?
- Thread가 checkpoint history라는 정의가 디버깅에 어떻게 도움이 되는가?
LangGraph Essentials-Python 대본.md (Intro 랩) — "LangGraph is a framework that provides a durable runtime for AI agents and applications.", "LangGraph saves the state of the application at each step. A restarted application can then resume exactly where it left off.", "LangGraph implements a human in the loop interface that suspends the agent to await human input.", "these features, parallel operation, streaming, checkpointing, human in the loop and support for tracing and evaluation are the reasons that you should learn to use LangGraph." ↩︎ ↩︎ ↩︎ ↩︎
LangGraph Essentials-Python 대본.md (Nodes 랩) — "State is simply data... It is supplied to the graph, is updated by the graph and returned to the user. The graphs themselves are stateless.", "State can be persisted across time and in particular across failures of nodes. So if, for example, a node were to fail during execution, it could be restarted, the state could be restored, and the function could be run again from the beginning." ↩︎
LangGraph Essentials-Python 대본.md (Edges 랩) — "Edges define control flow, but they do not control the data that nodes have access to.", "the value written last would overwrite the previous state... we can introduce the reducer function. The name comes from the general term MapReduce...", "We'll use operator dot... add as the reducer function. This will concatenate all list updates into state instead of overwriting the previous value.", "this reducer function is actually up for you to define, and you can even create your own custom reducer... annotated the state variables when defining the state.", "B to BB, C to CC" (L111), "nodes C and B execute within the same super step... they both observed the same state when they run... BB executes and CC executes the next super step... they both can see the state updates from nodes B and C" (L115-117), "state is shared across the graph and is provided to nodes at the start of a super step and is updated by those nodes at the end of a super step" (L195) ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎
LangGraph Essentials-Python 대본.md (Memory 랩) — "A checkpointer will store the state into more persistent storage at the end of each step, which is essentially taking a snapshot of the state at that point.", "A thread is the collection of those checkpoints over time. This is the entire history of the state at each execution step." ↩︎
langgraph-durable-execution.md, langgraph-memory.md — durable-execution.md의 checkpointer/store 비교표: "Scope | A single thread | Across threads", "Memory type | Short-term, thread-scoped memory | Long-term, cross-thread memory", "Most applications can use both: a checkpointer tracks the current thread, and a store tracks durable information across threads."; memory.md: "Short-term memory... tracks the ongoing conversation... State is persisted... using a checkpointer", "Long-term memory... is saved within custom 'namespaces'" ↩︎ ↩︎ ↩︎
LangGraph Essentials-Python 대본.md (Interrupt 랩) — "A human response can take a while, so ideally you'd like operation to suspend while you wait and then resume when the information is supplied. And this is exactly what interrupt is for." ↩︎