AI Agent Interaction Design
Agent quality는 model intelligence 하나가 아니라 goal, prompt, orchestration, tool, state, safety, evaluation이 feedback을 교환하는 interaction contract에서 나온다고 보인다. 한 layer의 약점을 다른 layer prompt로 가릴수록 demo는 작동해도 production failure를 진단하기 어려워진다. 이 페이지는 이 vault의 ai-agents 도메인 concept page들이 각자 다루는 layer를 하나의 loop로 묶는다.
전체 loop
flowchart TB
G[Goal] --> P[Prompt and context]
P --> O[Orchestration]
O --> M[Model decision]
M --> T[Tool proposal]
T --> S[Safety and authorization]
S --> X[Execution]
X --> V[Validated observation]
V --> O
O <--> K[Persistent state]
O --> R[Response and trace]
R --> E[Evaluation]
E --> H[Human steering]
H --> G각 layer가 어느 source에 대응하는가
| Layer | Source가 다루는 것 | Wiki concept page |
|---|---|---|
| Prompt | role, context, instruction, criteria[1] | [[Wiki/ai-agents/concepts/Prompt Engineering |
| Architecture | model, orchestration, tool loop[2] | [[Wiki/ai-agents/concepts/AI Agent Architecture |
| Harness | context, observability, rule[3] | [[Wiki/ai-agents/concepts/Agent Harness |
| Memory | state, checkpoint, thread[4] | [[Wiki/ai-agents/concepts/Agent Memory and Persistence |
| Safety | mandatory moderator 사례[5] | [[Wiki/ai-agents/concepts/Agent Safety Boundary |
| Evaluation | rubric, fixed scenario[6] | [[Wiki/ai-agents/concepts/Agent Evaluation Loop |
Source마다 적용 범위가 다르다
- Whitepaper는 agent를 가진 도구로 세계를 관찰하고 행동해 목표를 이루려는 application으로 넓게 정의한다.[7] Agent Patterns 노트는 Anthropic 분류상 workflow와 agent를 모두 agent system으로 보면서, 미리 정해진 코드 경로를 따르는 workflow와 작업 과정·도구 사용을 동적으로 제어하는 agent를 정의로 구분한다.[8] Label보다 control 위임 정도가 중요하다는 것은 이 두 source를 나란히 놓았을 때 나오는 결론이다.
- Harness source는 repository engineering(Codex 사례)[3:1]과 product SDK(Deep Agents류)[9]라는 다른 용례를 사용한다.
- Safety source는 단일 application(Leafy) 사례이므로 임상 안전성이나 보편 topology를 주장하지 않는다.[5:1]
- Evaluation source는 이 프로젝트가 쓰는 Gemini API가 logprob를 제공하지 않을 뿐 아니라 GPT-4, Claude API도 logprob를 완전히 노출하지 않으므로, 상용 LLM API 환경 전반에서 PPL 직접 측정은 현실적으로 불가능하다고 본다.[6:1]
Layer별로 신뢰를 쌓는 순서
이 순서는 위 표의 layer들을 이 페이지가 하나의 절차로 배열한 것이며, 특정 source가 이 순서 자체를 제시하지는 않는다.
- Goal을 acceptance와 금지 조건으로 바꾼다.
- Deterministic rule은 prompt가 아니라 code, schema, policy에 둔다.
- Tool proposal과 execution을 분리한다.
- State는 thread와 checkpoint identity를 가진다.
- Irreversible action 전에 safety boundary를 둔다.
- Observation을 untrusted input으로 검증한다.
- Response와 trace를 scenario별로 평가한다.
- Human은 failure를 올바른 layer 수정으로 연결한다.
한 layer로 다른 layer의 문제를 덮을 때
- Prompt로 permission 문제를 해결하면 bypass가 남는다.
- 강한 model로 state identity 문제를 해결하면 cross-thread contamination이 남는다.
- Retry를 늘리면 duplicate write가 생길 수 있다.
- 실행 후 moderation이면 unsafe action에는 늦다.
- Judge score만 최적화하면 grounding과 trace가 악화될 수 있다.
아직 근거가 부족한 부분
- Multi-agent coordination, distributed consensus, long-term semantic memory quality를 다루는 source가 이 도메인에는 부족하다.[10]
- Formal capability security model은 없다.
- Product API는 version drift가 있으므로, 이 페이지는 특정 SDK 표기보다 boundary 중심으로 유지한다.
테스트 질문
- Prompt, architecture, harness는 어떤 실패를 해결하는가?
- Memory, safety, evaluation은 loop 어디에 연결되는가?
관련
출처
Prompt Engineering.md — "역할과 맥락 지정", "명확하고 구체적 지시", "예시 제공(Few-shot)", "맥락과 입력 정보 제공", "모델이 생성한 결과를 평가할 기준을 프롬프트에 넣음" ↩︎
Agent Architecture.md — "모델 (Model): 추론을 담당하는 LLM", "오케스트레이션 레이어 (Orchestration Layer): 판단-행동 루프를 조율하는 시스템", "도구 (Tools): 에이전트가 외부 세계와 상호작용하는 수단" ↩︎
Harness Engineering.md — "Codex에는 1,000페이지의 설명서가 아니라 맵을 제공해야 한다.", "에이전트 관점에서 실행 중 컨텍스트 안에서 접근할 수 없는 것은 사실상 존재하지 않는다.", "전용 린터와 CI 작업은 지식 베이스가 최신 상태이고, 교차 링크되어 있으며, 올바르게 구성되어 있는지 검증한다." ↩︎ ↩︎
LangGraph Essentials-Python 대본.md — "A checkpointer will store the state into more persistent storage at the end of each step...", "A thread is the collection of those checkpoints over time." (상세 quote는 Agent Memory and Persistence 참고) ↩︎
multi-agent-safety-moderator.md — "에이전트 응답이 SafetyModerator를 거치지 않고 사용자에게 도달하는 경로가 존재하지 않는다." ↩︎ ↩︎
06-에이전트-성능-평가.md — (L18) "Gemini API는 logprob를 제공하지 않는다.", (L28) "GPT-4, Claude API도 logprob를 완전히 노출하지 않는다. 상용 LLM API 환경에서 PPL 직접 측정은 현실적으로 불가능하다.", 4개 차원 judge rubric과 고정 시나리오 표 ↩︎ ↩︎
22365_19_Agents_v8.pdf — "What is an agent?" 절(PDF 3쪽, 추출 텍스트의 공백 정규화 기준): "In its most fundamental form, a Generative AI agent can be defined as an application that attempts to achieve a goal by observing the world and acting upon it using the tools that it has at its disposal." ↩︎
Agent Patterns.md — (L18) "Anthropic에서는 위 2개 다 에이전트 시스템이라고 분류한다.", (L19-20) "워크플로우란, LLM과 여러 도구들이 미리 정해진 코드 경로에 따라 작동할 수 있도록 하는 시스템이다.", "에이전트는 LLM들이 자신의 작업 과정과 도구 사용 방식을 동적으로 제어할 수 있는 시스템이다." ↩︎
deepagents-agent-loop-termination.md — (L434) "Deep Agents is an" ... "agent harness" ... "It is the same core tool calling loop as other agent frameworks, but with built-in capabilities that make agents reliable for real tasks" ↩︎
출처 매핑 미확인 — multi-agent coordination, distributed consensus, long-term semantic memory quality에 대한 이 도메인의 평가는 기존 7개 source 중 어느 것도 직접 다루지 않는다. 해당 주제를 다루는 새 source를 ingest해야 확인할 수 있다. ↩︎