Typed Decision Model

"Agent 안의 작은 판단(라우팅, 분류, 점수, 통과 여부)을 매번 LLM에게 문장으로 물어야 하는가?" 이 페이지는 그 질문에 답한다. Typed decision model은 판단 대상인 상태(state)와 미리 형태를 정한 질문을 받아, 글 대신 타입이 정해진 값과 확률 분포를 돌려주는 모델 계층이다. 2026-09-15 TypeSafe AI가 Jev를 "System One model"로 공개하면서 널리 알려졌지만, 이 페이지는 제품 소식이 아니라 판단 단계를 코드·결정 모델·LLM 사이에 어떻게 나눌지를 다룬다. 상위 맥락은 AI Workflow vs Agent의 control boundary이며, 이 계층은 control을 코드에 남긴 채 그 안의 흐릿한 조건문만 모델에 맡기는 쪽으로 볼 수 있다.

무엇을 주고 무엇을 받나

질문 형태 묻는 것 돌려받는 것
Choice 보기 목록 중 하나를 고른다 choice, probabilities, confidence
Score 정해진 루브릭 단계 위에 상태를 놓는다 score, probabilities, confidence
Noul 이 문장이 참인가 noul (0–1)

코드·결정 모델·LLM의 역할 분담

구성 요소 잘하는 일 맡기면 안 되는 일
Deterministic code 산술, 권한, 한도, 날짜, 상태 전이, 부작용 규칙으로 유지할 수 없는 모호한 의미 분류
Typed decision model 주어진 상태에 대한 제한된 선택, 점수, 예/아니오 판단 사용자용 문장, 정확한 계산, 열린 계획, 최종 승인
Language model 설명, 종합, 초안, 대화, 열린 추론 결과가 큰 부작용에 대한 무감독 권한
flowchart LR
    R[요청과 상태] --> P{코드: 금지 규칙·한도}
    P -->|금지| X[거절]
    P -->|허용| D[결정 모델: Choice·Score·Noul]
    D --> G{confidence 게이트}
    G -->|높음| H[선택된 handler: 코드 또는 LLM]
    G -->|낮음| U[사람 검토 또는 큰 모델]
    H --> A[코드: 권한 확인·실행·기록]
단계 결정 모델이 줄 수 있는 것 코드가 계속 가져야 하는 것
Route 의도, 복잡도, 위험 등급 허용된 목적지, quota, provider 상태
Approve 추천 또는 의미상 정책 일치 여부 승인 권한, 한도, 레코드 버전, 최종 commit
Rank 루브릭 점수, 쌍별 관련도 안정 정렬, 동점 규칙, 필수 포함 항목
Retry or stop 결과가 불완전하거나 과제를 벗어났는지 재시도 상한, 멱등성, timeout 상태
Escalate 모호함, 민감도, 예외 가능성 escalation 정책, 접근 제어

confidence는 판단의 확신도이지 권한이 아니다

action = response.answers["action"]
if action.confidence < 0.5:
    route_to_human(user_message)          # 모델이 모른다고 할 때는 추측하지 않는다
elif action.choice == "check_balance":
    show_balance(account_id)              # 틀려도 되돌릴 수 있는 읽기 작업
elif action.choice == "approve_transfer":
    if action.confidence > 0.9:
        confirm_then_execute(account_id)  # 위험한 작업은 더 높은 임계값 + 확인
    else:
        ask_user_to_confirm(account_id)

큰 질문 하나보다 원자적 질문 여러 개

실패 유형

실패 유형 대신 할 일
글자 그대로 읽기 정확한 조건과 보기별 기준을 쓴다
수학과 숫자 산술은 코드에 둔다
날짜·시간 비교 구성 요소를 추출하고 비교는 코드에서 한다
간접 참조 단계를 줄이고 관련 상태를 이름으로 가리킨다
관계없는 정보가 많은 큰 상태 먼저 걸러 필요한 필드만 보낸다
적대적 내용 기준을 명확히 쓰고 배포 전 경계 사례를 시험한다
지시문과 기준의 모순 기준을 지시문의 연장으로 맞춘다
상식적 구조 불변식 한 판단은 한 방식으로 묻고 항등식은 코드에서 강제한다
생성 생성 모델을 쓴다

성능은 어느 정도인가

2026-09 기준 스냅샷

아래 수치는 jev-1.13 공개 2주 이내의 측정이며, jev-latest alias가 바뀌면 달라질 수 있다. 버전별 세부 수치는 원본 요약 페이지와 Raw source에 남기고, 이 절에는 설계 판단에 필요한 결론만 둔다.

한국어 입력

새로운 기술인가

조각 선행 연구
보기별 label likelihood로 점수 매기기 GPT-3 §2.4, MMLU §4.1, lm-evaluation-harness multiple_choice
label을 토큰 하나에 대응시키기 Verbalizers (Schick & Schütze)
label 확률에 사전 편향이 있음 Calibrate Before Use (Zhao et al., 2021)
온도 하나로 과신을 고침 Guo et al. (2017), Kadavath et al. (2022)
확신이 없으면 큰 모델로 넘기기 Selective classification (2017), FrugalGPT

도입 전에 할 일

관련

테스트 질문

출처


  1. TypeSafe docs: introduction, confidence, jaggedness — Introduction: "Jev evaluates typed questions against a state and returns structured results directly. No text generation, no parsing.", 세 primitive 표(Choice/Score/Noul), "All three question types can be mixed in a single API call. Every question is evaluated in parallel and in isolation against the same state in one go.", "If the question you want to ask would require extended reasoning or weighs multiple independent factors, decompose it."; Confidence: "confidence is a statistic computed from the probability distribution", "(Noul answers don’t carry one.)", 세 구간(High/Medium/Low)과 "Thresholds scale with risk" 예시 코드; Jaggedness: 9개 failure mode 표(L143-151: "Keep the arithmetic in code", "Extract components; compare in code", "Ask each decision one way; enforce identities in code", "Write the exact condition, criteria for each available options", "Use a generative model"), (L206) "Extraction is a judgment, so give it to the model. Arithmetic is not, so keep it in code.", "count in code", refund 0.72 / not_refund 0.47 / Sum 1.19, "State is data, and jev-1.13 does not treat it as hostile by default." ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  2. DEV independent evidence review — "Jev sits level with mid-price LLMs and 6.5 to 11.5 points behind the frontier in the cleanest comparison.", "Binary and few-class decisions, fast, with answers that always fit the schema.", Janardhan 표 Jev ECE 0.161, "The direction of the error changes with the data.", "one temperature fitted on 50 to a few hundred labels fixes most of the error", "Speed runs from 0.5x, slower than a local Gemma, to 12.1x faster; cost runs from 0.6x, dearer, to 478x cheaper", "Most of the cost gap comes from output being free.", "A router given option names with no descriptions sent all 40 hard tasks to the cheap model, at a median confidence of 0.96. One-line option descriptions fixed 37 of 40.", "every answer on KoBBQ bias questions is wrong by design, and Jev gave them 0.79 confidence", "Questions batched in one request cannot see each other's answers.", 피싱(L221): "Luce trained on 1,000 labels reached 97.4% against Jev's 62.6% on the same benchmark, though on different items.", "a two-line regex scored 91.6%, five narrow Jev questions combined by logistic regression reached 95.0% on a held-out half, and Haiku 4.5 asked the same five questions reached 93.2%, a difference too small to be significant (p = 0.063)", 선행 연구 표, "The visible contribution is the packaging", "Reinforcement learning for calibrated decisions has no paper, patent or method description.", "Any open model can be read this way". ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  3. TypeSafe launch post — "Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots.", "This is where the claims of 193.6x faster, 444.6x cheaper on our home page comes from, and we expect that these are on the higher end of real world gains.", "they were made by individuals on our model capabilities team, so some bias could exist." ↩︎ ↩︎

  4. Wavect Jev review — Deterministic code / Jev / Language model 역할 표, "Hard policy belongs in code, before probabilistic routing.", "the handler remains responsible for its own permissions, validation and output quality", Route/Approve/Rank/Retry or stop/Escalate 표, "Typed is not correct.", "Confidence is not authorization. A high value cannot grant access, approve a payment or bypass a risk limit.", "State is an attack surface.", "Closed sets need an escape route.", "Start with one frequent, reversible decision that already has labeled outcomes.", 7단계 pilot 목록. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  5. arXiv 2609.37647 — Abstract: "It beats Qwen on 27 of 37 datasets, with none of Qwen’s nine leads outside the bootstrap intervals, and Gemma on all 37.", "Jev’s choice probabilities are well calibrated and support selective prediction. Binary probabilities rank well but are poorly placed relative to a fixed 0.5 threshold; thresholds tuned on training data raise micro-F" ... "on UNFAIR-ToS from 0.50 to 0.75."; §4: "Pooled over 22 Choice datasets and 279,925 answers, the ECE is 0.028."; (L202) "TypeSafe AI states that English is Jev’s primary language", "On Belebele, accuracy is 97.2% for English, and the median over 122 language varieties is 91.1%: 71 varieties reach at least 90%, while 10 fall below 70%, down to 37.4% for Nigerian Fulfulde."; (L255) "Performance drops sharply for low-resource languages". ↩︎ ↩︎ ↩︎ ↩︎

  6. Jev Korean sample check — Belebele 한국어 96/100·영어 97/100, PAWS-X 한국어 76/100·영어 80/100, "Korean content scored within a point or two whether the instructions were Korean or English. Translating your instructions buys nothing.", "moving from English to Korean costs Luna 3 points on Belebele and 11 on PAWS-X ... The same switch costs Jev 1 and 4 points", KorMedMCQA Jev 80 / Luna 88 "the one result in this check that survives its own uncertainty", "100 questions give roughly ±8 points of uncertainty". ↩︎ ↩︎ ↩︎