Agent Tool Use
Tool calling은 model이 tool과 arguments를 제안하는 interface일 뿐이다. Tool use는 그보다 넓어서 selection, authorization, execution, observation validation, retry, stopping까지 포함한다. ReAct는 inference 중 observation을 reasoning에 되먹이는 실행 패턴이고, Toolformer는 API call example을 self-supervised하게 학습하는 training method다 — 둘은 같은 layer에 있지 않다.
실행 흐름: 제안과 실행은 분리된다
sequenceDiagram
participant O as Orchestrator
participant M as Model
participant X as Executor
O->>M: Context and tool schema
M-->>O: Tool and arguments
O->>O: Validate and authorize
O->>X: Execute
X-->>O: Raw result
O->>O: Validate and record
O->>M: Observation- Whitepaper Figure 9는 model이 arguments를 채우고 client가 external API를 실행하는 lifecycle을 보여준다.[1] function calling이 자동 실행과 같은 뜻이 아니라는 것은 이 sequence에서 나온다 — client가 returned arguments로 API를 호출하는 단계가 별도로 있다.
Whitepaper의 도구 분류
| 유형 | 역할 | control |
|---|---|---|
| Extensions | API와 agent의 standardized bridge | agent-side infrastructure |
| Functions | model이 function과 arguments 제안 | client가 실제 실행 |
| Data Stores | indexed data retrieval | context observation 반환 |
- Whitepaper Figure 8은 Extensions의 agent-side control과 Functions의 client-side control을 구분한다.[1:1]
- 이 taxonomy는 Google whitepaper의 ecosystem-specific 구분이며, 다른 vendor의 tool 분류와 그대로 대응하지 않는다. 기존에 이 페이지에 있던 Plugins와 Actions 항목은 cited source가 직접 정의하지 않아 제거했다.
ReAct와 Toolformer는 같은 층위가 아니다
| 기준 | ReAct | Toolformer |
|---|---|---|
| 시점 | inference prompting | training data와 fine-tuning |
| 단위 | Thought, Action, Observation | API call과 result |
| feedback | environment observation | future-token loss filtering |
| 아닌 것 | architecture 전체 | runtime planner |
- ReAct는 observation을 다음 reasoning에 반영한다.[2]
- Toolformer는 loss improvement가 충분한 call만 training set에 남긴다.[3]
Production에서 신뢰성을 좌우하는 것
- Production reliability는 model selection보다 schema, authorization, idempotency, observation validation에 크게 의존한다.
- Tool description은 affordance이자 security surface다. 겹치는 tool은 ambiguity와 privilege를 키운다.
- Observation은 external input이므로 provenance, freshness, schema를 검증한다.
- Read와 irreversible write는 같은 retry policy를 사용할 수 없다.
이 네 가지는 위 실행 흐름과 도구 분류로부터 이 페이지가 도출한 운영 원칙이며, 세 source가 직접 명시한 production checklist는 아니다.
흔히 겪는 실패
- Tool hallucination, ambiguous selection, premature execution, observation poisoning, duplicate write, context flooding, stopping rule 부재가 주요 실패다.
테스트 질문
- Function calling에서 model과 client 책임은 어떻게 나뉘는가?
- ReAct와 Toolformer는 어느 lifecycle 단계에서 tool use를 다루는가?
출처
22365_19_Agents_v8.pdf (Figure 8, Figure 9) — 이 PDF는 이번 변환 세션에서 poppler 부재로 재추출하지 못했고, 2026-07-11 conformance 기록 시점에 검증된 기존 claim을 그대로 유지한다. ↩︎ ↩︎
ReAct_2210.03629.pdf — 이 PDF는 이번 변환 세션에서 poppler 부재로 재추출하지 못했고, 2026-07-11 conformance 기록 시점에 검증된 기존 claim을 그대로 유지한다. 상세 mechanism은 ReAct 참고. ↩︎
Toolformer_2302.04761.pdf — 이 PDF는 이번 변환 세션에서 poppler 부재로 재추출하지 못했고, 2026-07-11 conformance 기록 시점에 검증된 기존 claim을 그대로 유지한다. 상세 mechanism은 Toolformer 참고. ↩︎