ReAct (Reasoning and Acting)

ReAct는 언어 모델이 추론(Reasoning)과 행동(Acting)을 어떻게 결합해야 복잡한 문제를 풀면서 할루시네이션을 줄일 수 있는가라는 질문에 답하는 논문이다. 상위 개념은 Agent Architecture와 Prompt Engineering이고, 하위 개념은 Thought-Action-Observation Loop이며, 이웃 개념은 Chain-of-Thought(CoT)와 Reflexion이다.

추론과 행동의 교차 생성

ReAct는 에이전트가 외부 환경과 상호작용하기 위해 사용하는 추론과 행동의 교차 생성(Interleaved Generation) 패러다임이다.

논문은 사람에게서 "acting"과 "reasoning"의 타이트한 결합이 새 작업을 빨리 배우고, 처음 보는 상황이나 불확실한 정보 아래에서도 견고한 의사결정이나 추론을 하게 한다는 점을 동기로 든다.[1] 모델 쪽 결과로는 ReAct가 한 개에서 여섯 개의 in-context 예시만으로 새 task instance에 강한 일반화를 보인다고 보고한다.[1:1]

논문이 제시하는 근거

백엔드 시스템에서 ReAct가 갖는 의미

ReAct는 단순한 텍스트 생성기가 아니라, 외부 API와 연동되는 분산 워크플로 시스템(Distributed Workflow System)의 코어 엔진 역할을 한다고 해석할 수 있다. ReAct 루프는 상태 머신(State Machine)처럼 작동하며, 각 단계의 Observation이 다음 Thought의 입력으로 주입되면서 동적인 계획 수정(Dynamic Replanning)을 가능하게 한다.

동시에 성능과 유연성은 상충한다. ReAct는 외부 검색에 크게 의존하기 때문에 검색된 정보가 유용하지 않을 경우(non-informative search) 오히려 추론 흐름이 꼬여버릴 수 있다. 이를 방지하기 위해 내부 지식이 확고한 경우 CoT로, 그렇지 않은 경우 ReAct로 전환하는 혼합 방식(ReAct + CoT-SC)이 실무적으로 권장된다.

Note

아래 API 추상화 설계와 실패 모드 대응은 ReAct 논문이 직접 제시한 architecture가 아니라, 논문의 Thought-Action-Observation loop를 backend system에 적용할 때의 설계 해석이다.

에이전트 친화적 API 추상화

사람이 사용하는 GUI나 복잡한 웹 네비게이션을 그대로 모델에게 노출하는 것은 비효율적이다. ReAct 논문이 Wikipedia 탐색을 오직 Search, Lookup, Finish 3가지의 명확한 시맨틱으로 제한한 것처럼,[1:6] 백엔드 개발자는 LLM이 오해 없이 파싱하고 예측할 수 있는 수준으로 도구(API)의 추상화 계층을 설계해야 한다.

실패 패턴과 운영 대응

논문에서 분석된 ReAct의 주요 에러 패턴은 백엔드 시스템의 모니터링·복구(retry) 전략과 직결된다.[1:7]

ReAct 루프가 도는 순서

  1. 사용자 요청 입력: 예) "X가 Y보다 큰가?"
  2. Thought 생성: "X의 크기를 찾은 후 Y의 크기를 찾아야 한다. 먼저 X를 검색하자."
  3. Action 실행: Search[X]
  4. Observation 획득: 외부 API가 반환한 X의 크기 정보.
  5. Thought 갱신: "X의 크기는 알았다. 이제 Y를 검색하자."
  6. Action 실행: Search[Y]
  7. Observation 획득: 외부 API가 반환한 Y의 크기 정보.
  8. Thought 결론: "X와 Y의 크기를 비교해보면..."
  9. Action 종료: Finish[결과 반환]
sequenceDiagram
    participant LLM
    participant Env as External Environment (API)
    
    rect rgb(30, 40, 50)
    note right of LLM: ReAct Loop
    LLM->>LLM: 1. Thought (추론: 현재 상태 분석 및 계획)
    LLM->>Env: 2. Action (행동: 도구 호출)
    Env-->>LLM: 3. Observation (관찰: 도구 실행 결과)
    end
    
    LLM->>LLM: 4. Thought (결론 도출)
    LLM->>Env: 5. Action (Finish: 결과 반환)

루프 종료: LLM의 신호와 코드의 상한

ReAct 루프에는 종료 주체가 둘 있다. 평소에는 LLM이 끝낼 때를 정하고, LLM이 끝내지 못할 때는 코드가 상한으로 끊는다.

HotpotQA 예시로 보는 재검색

다른 개념과의 관계

관련

출처

테스트 질문


  1. ReAct_2210.03629.pdf (Introduction, 실험 결과, 에러 분석) — 이 PDF는 이번 변환 세션에서 poppler 부재로 재추출하지 못했고, 2026-07-11 conformance 기록 시점에 검증된 기존 claim(Introduction의 CoT/Act-only 한계 서술, ALFWorld/WebShop 벤치마크 수치, Wikipedia action space Search/Lookup/Finish, Reasoning Error 47%·Search Error 23% 에러 분류, HotpotQA 예시)을 그대로 유지한다. 2026-10-06 텍스트 추출본 대조 인용: "This tight synergy between “acting” and “reasoning” allows humans to learn new tasks quickly and perform robust decision making or reasoning, even under previously unseen circumstances or facing information uncertainties." (p.1 Introduction), "ReAct shows strong generalization to new task instances while learning solely from one to six in-context examples" (p.4), "when ReAct fails to return an answer within given steps, back off to CoT-SC. We set 7 and 5 steps for HotpotQA and FEVER respectively" (p.5), Table 2 "Reasoning error Wrong reasoning trace (including failing to recover from repetitive steps)" (p.6), "one frequent error pattern specific to ReAct, in which the model repetitively generates the previous thoughts and actions, and we categorize it as part of “reasoning error”" (p.6) ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  2. deepagents-agent-loop-termination.md — LangChain Agents 도입부: "An agent is a model calling tools in a loop until a given task is complete.", "Deep Agents builds on create_agent"; Deep Agents overview: "It is the same core tool calling loop as other agent frameworks"; Fault tolerance: "Without limits, a confused agent can burn through your LLM API budget in minutes by looping on the same tool call or making hundreds of model calls. Set caps on both model calls and tool executions per run", "Use run_limit to cap calls within a single invocation (resets each turn). Use thread_limit to cap calls across an entire conversation (requires a checkpointer)."; factory.py: "# 3. If the model hasn't called any tools, exit the loop / # this is the classic exit condition for an agent loop"; graph.py: return create_agent( ... "recursion_limit": 9_999. ↩︎ ↩︎ ↩︎ ↩︎