o2o-negosium-original/agent/tools/train_full_autonomy.py
jwkim 51b3820cc9 [feat] agent: 완전 자율 협상 모드 (AUTONOMY_MODE) + LLM 멘트
판정 룰(앵커타결/와일드존/3라운드결렬)과 카드 선택을 학습 정책으로 대체:
수락/역제안 금액/압박 화법/결렬 전부 행동 30개(수락1+결렬1+역제안 6단x화법4+압박4)에서 선택.

- autonomy_actions: 행동 공간·특징 인코딩 (학습/서빙 공유)
- autonomy_store: numpy 서빙 + 행동 봉투 7개(수락<=목표가 / 역제안 단조 / 역제시·결렬은
  설득 2회 후 해금 / 마무리국면 압박 금지 / 첫 역제안 앵커 이하 / 최종제안 1회 보장)
- chat_engine: 자율 스텝(역제안/최종제안/압박1~4), 최종제안 금액=목표가, 턴캡 12
- ment_generator: Gemini 멘트 생성 + 가드(숫자 화이트리스트·목표가 비공개·금지어·문장완결,
  실패시 템플릿 폴백, 6초 컷), 인터넷최저가 근거 인용(수집됨+제시가 초과시만), 대화 기억
- chat_service: 자율 행동 experience_logs 로깅(AUT|종류|위치|전략), 대화기억 ctx 관리
- context loader/CRUD: 인터넷최저가·견적기간·협력사 이력 로드 (v3 상태 21차원)
- train_full_autonomy: 시뮬 15k ep — 앵커율 0.8~6% 정합, 협력사 현실화(컷반발·반복짜증·
  양보 상호성), 관측성 마스크(마감 40% 미관측·15% 전부미상 — 서빙 중립값 분포 정합),
  보상 수정(목표가 초과 타결=결렬 취급)
- 서빙 v3.5 (v3.3 목표가 즉시지르기 퇴화, v3.4 소액지형 첫턴 통보 퇴화 — 게이트 반려 이력 보관)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 16:57:38 +09:00

366 lines
20 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

"""train_full_autonomy — 행동 룰 0개, 완전 자율 협상 에이전트 (v2 시뮬 프로토타입).
기존 시스템의 룰(앵커 이하 강제타결 / 3라운드 강제결렬 / 와일드카드 존 / 카드 카탈로그)을
전부 제거하고, 모든 결정을 에이전트 행동으로 이관한다:
행동 공간 (action-as-feature, ScoreNet 이 후보 열거 채점):
ACCEPT 현재 제시가로 타결 ← '앵커 이하 강제타결' 룰 대체
WALK 협상 결렬 선언 ← '3라운드 강제결렬' 룰 대체
COUNTER(C) "C원이면 수락" 역제안 ← '와일드카드 1%' 룰 대체 (금액도 학습)
PRESS(strategy) 설득 압박(카드의 일반화) ← 카드 카탈로그 대체 (전략만 남음)
룰이 사라진 자리는 보상이 채운다(유일한 스펙):
R = W×R_price + (1W)×R_end λ×round (기존 RewardCalculator 그대로)
협상이 끝나는 길: 에이전트의 ACCEPT/WALK, 협력사의 COUNTER 수락, 협력사의 인내심 소진(이탈).
마지막 것은 시스템 룰이 아니라 상대방 특성이다.
베이스라인 = 현행 룰 시스템을 같은 환경에서 재현(앵커타결/1%클로징/3라운드결렬 + 압박).
실행: APP_ENV=local PYTHONUTF8=1 python -m tools.train_full_autonomy
"""
import os
import random
from typing import Optional, Tuple
import numpy as np
import torch
from eval_harness.feature_buyer import AFFINITY, STRATEGY_PROFILE, SupplierProfile, sample_supplier
from negotiation.policies.autonomy_actions import (
ACTION_DIM, ACTIONS, COUNTER_GRID, EXTRA_STATE_DIM, Action, extra_state,
internet_gap_feat, settle_norm as extra_settle)
from negotiation.policies.feature_dqn_policy import FeatureDQNPolicy
from negotiation.qtable.domain.model.snapshot import NegotiationOutcome, NegotiationSnapshot
from negotiation.qtable.domain.service.feature_builder import (
STATE_FEATURE_DIM, TENANT_FEATURE_DIM, build_state_features, build_tenant_features)
from negotiation.qtable.domain.service.reward_calculator import RewardCalculator
from tenancy.config_loader import TenantConfigLoader
from tools.train_feature_dqn import pref_config, sample_tenant_pref
_HERE = os.path.dirname(os.path.abspath(__file__))
CKPT_PATH = os.path.join(_HERE, "..", "artifacts", "full_autonomy.pt")
TARGET = 10000.0
# 앵커율(v3.1): 실운영 기하 정합 — 앵커가 = 목표가×(1a), a ∈ [0.8%, 6%] 를 에피소드마다 샘플링.
# (기존 고정 20% 폭은 실제(≈1%)와 지형이 달라, 실서비스에서 압박/역제안 밸런스가 어긋났다.)
ANCHOR_RATE_RANGE = (0.008, 0.06)
# 행동 공간(Action/ACTIONS/COUNTER_GRID/ACTION_DIM)은 negotiation.policies.autonomy_actions 공유
# — 서빙(autonomy_store, numpy 전용)과 학습이 같은 인코딩을 쓴다.
# ---- 협력사 모델 (상대 반응: 역제안 수락/재제안 포함) ---------------------------------
class AutonomousBuyer:
"""FeatureBuyer 확장: 역제안(C)에 반응한다. 이탈은 '인내심' — 시스템 룰이 아닌 상대 특성."""
def __init__(self, sup: SupplierProfile, seed: int):
self.sup = sup
self.rng = np.random.default_rng(seed)
# 기질 t ∈ [0,1]: 0=터프(하한 높고 안 물러섬) ↔ 1=수월. 관측 가능한 이력·최저가가
# 이 숨은 기질과 상관되게 생성된다 → 에이전트가 이력/최저가 특징을 읽을 '이유'가 생긴다.
# 하한은 '우리 앵커'가 아니라 협력사 사정(≈목표가 기준)으로 정해진다(v3.1) —
# 하한 > 목표가(≈35%)면 애초에 성사 불가능한 협상이고, 그걸 빨리 알아채고 끊는 것도 실력이다.
t = float(self.rng.uniform(0.0, 1.0))
self.floor = TARGET * float(np.clip(1.12 - 0.24 * t + self.rng.normal(0, 0.02), 0.85, 1.18))
self.patience = int(self.rng.integers(4, 9)) + (1 if t > 0.7 else 0)
# 첫 제시가: 목표가의 105~150% — 실운영(기존 공급가가 목표가를 26%+ 상회) 분포를 덮는다.
# 좁게(110~125%) 학습하면 큰 갭 상황에서 정책이 분포 밖 일반화(대형컷 역제안)를 한다.
self.price = TARGET * float(self.rng.uniform(1.05, 1.50))
# 하한가가 첫 제시가보다 높을 수 없다(자기 하한 밑으로 부르고 시작하는 판매자는 없음).
# 이 보정이 없으면 on_press 의 max(floor,·)가 가격을 '역주행'시키는 비현실이 생긴다.
self.floor = min(self.floor, self.price * 0.98)
self._last_c: Optional[float] = None # 직전 역제안 (같은 숫자 반복 짜증 모델링)
# ---- 관측 가능 부가정보 (v3 특징 소스 — 기질과 상관, 노이즈 있음) ----
self.hist_n = int(self.rng.integers(0, 6)) # 과거 협상 횟수 (0=신규)
if self.hist_n:
self.hist_success = float(np.clip(0.25 + 0.6 * t + self.rng.normal(0, 0.10), 0.0, 1.0))
self.hist_settle_ratio = float(np.clip(1.18 - 0.28 * t + self.rng.normal(0, 0.04), 0.80, 1.30))
else:
self.hist_success = self.hist_settle_ratio = None
# 인터넷최저가: 숨은 하한가의 노이즈 관측치. 60% 확률로만 수집돼 있음(현실: 미수집 흔함).
self.internet_lowest = (self.floor * float(self.rng.uniform(0.98, 1.08))
if self.rng.random() < 0.6 else None)
def _powers(self, strategy: int) -> Tuple[float, float]:
conc, acc = STRATEGY_PROFILE.get(strategy, (0.5, 0.5))
m = AFFINITY[self.sup.segment].get(strategy, 0.5)
scale = 0.35 + 0.85 * m
return conc * scale, acc * scale
def on_press(self, strategy: int, turn: int) -> Tuple[bool, float]:
"""(이탈여부, 새 제시가). 압박이 안 먹히는 세그먼트면 이탈 위험이 실재한다."""
c_pow, a_pow = self._powers(strategy)
walk_p = 0.04 + 0.30 * (1.0 - a_pow) * (turn / self.patience)
if self.rng.random() < walk_p:
return True, self.price
concession = (self.price - self.floor) * (0.10 + 0.55 * c_pow)
self.price = max(self.floor, self.price - concession)
return False, self.price
def on_counter(self, c: float, strategy: int, turn: int) -> Tuple[str, float]:
"""역제안 C 반응: 'accept'(C로 타결) | 'walk' | 'counter'(새 제시가).
현실화(v2): 현 제시가 대비 인하 요구폭(cut)이 클수록 수락률이 급감하고 이탈 위험이 커진다
— 초기 버전에서 에이전트가 't1 원샷 로우볼'로 시뮬 허점을 착취하던 것을 막는다.
압박으로 가격을 충분히 끌어내린 뒤 작은 컷으로 클로징해야 통하는 구조.
"""
_, a_pow = self._powers(strategy or 3)
cut = max(0.0, (self.price - c) / max(self.price, 1.0)) # 인하 요구폭 (현 제시가 대비)
prev_c = self._last_c
repeated = prev_c is not None and abs(c - prev_c) < 1e-6 # 같은 숫자 반복
self._last_c = c
# 양보 상호성(v3.3): 직전 제안보다 올려 부르면(성의 있는 양보) 호의적으로 반응한다.
# 이 신호가 있어야 '상대가 내리면 우리도 조금 올리는' tit-for-tat 이 학습으로 나온다.
warm = 0.0
if prev_c is not None and c > prev_c + 1e-9:
warm = float(np.clip((c - prev_c) / max(self.price - self.floor, 1.0), 0.0, 0.35))
if c >= self.floor:
margin = (c - self.floor) / max(self.floor, 1.0)
p_acc = float(np.clip(0.20 + 0.9 * margin / 0.08, 0.0, 0.95)) * (0.75 + 0.35 * a_pow)
p_acc *= float(np.clip(1.0 - (cut - 0.05) / 0.20, 0.0, 1.0)) # 컷 5% 초과부터 반발, 25%면 수락 0
if repeated:
p_acc *= 0.25 # 이미 거절한 숫자를 또 내밀면 설득력 급감
p_acc *= 1.0 + warm
if self.rng.random() < min(p_acc, 0.97):
return "accept", c
# 모욕적 요구(하한 미달·과도한 원샷 컷·앵무새 반복) → 이탈 위험
low = max(0.0, (self.floor - c) / max(self.floor, 1.0))
p_walk = min(0.5, 2.0 * low) + 0.35 * max(0.0, cut - 0.20) / 0.20 + (0.15 if repeated else 0.0)
if self.rng.random() < min(p_walk * (1.0 - warm), 0.7):
return "walk", self.price
self.price = max(self.floor, c + (self.price - c) * float(self.rng.uniform(0.30, 0.60) + warm))
return "counter", self.price
# ---- 에피소드 실행 (룰 없음 — 종료는 행동 또는 상대 특성으로만) ------------------------
def make_snapshot(sup, price, turn, p0, anchor, outcome=NegotiationOutcome.ONGOING):
return NegotiationSnapshot(
revenue_amount=sup.revenue_amount, distribution_code=sup.distribution_code,
partner_count=sup.partner_count, acceptance_ratio=max(0.0, (p0 - price) / p0),
input_price=price, anchor_price=anchor, target_price=TARGET,
round_number=turn, outcome=outcome)
MIN_PRESS = int(os.getenv("AUTONOMY_MIN_PRESS", "2")) # 역제시 해금에 필요한 최소 설득 횟수
def available_actions(price: float, last_counter_q: Optional[float] = None,
counter_locked: bool = False) -> list:
"""행동 봉투 (serving autonomy_store 와 동일해야 한다):
① 목표가 초과 제시가는 '수락' 제외 — 매입 승인 범위(목표가 초과 수락 착취 방지)
② 직전 역제안 미만 금액의 역제안 제외 — 단조 양보 원칙(제안 철회 금지;
양보 '속도'는 정책이 배우고, 후퇴 '금지'만 구조로 보장)
③ counter_locked: 설득 MIN_PRESS 회 전에는 역제시 잠금 — 옛 제품 의미론
(일반 카드=설득, 역제시=와일드카드 성격의 마무리 수단) 복원
④ 마무리 국면(제시가 ≤ 목표가×1.005): 압박 제외 — 푼돈 차이에서 재검토 요청 반복 방지
⑤ 첫 역제안은 앵커 이하(q ≤ 0)만 — 낮게 개시해 사다리를 다 쓰며 올라간다"""
near_target = price <= TARGET * 1.005
return [a for a in ACTIONS
if not (a.kind == "accept" and price > TARGET)
and not (a.kind == "counter" and counter_locked and not near_target)
and not (a.kind == "walk" and counter_locked and not near_target)
and not (a.kind == "press" and near_target)
and not (a.kind == "counter" and last_counter_q is None and a.counter_q > 1e-9)
and not (a.kind == "counter" and last_counter_q is not None
and a.counter_q < last_counter_q - 1e-9)]
def action_feats(price: float, anchor: float, last_counter_q: Optional[float] = None,
counter_locked: bool = False):
"""현 제시가 기준 (가용 행동 리스트, 특징 [K, ACTION_DIM]). counter 는 컷폭 포함."""
span = max(TARGET - anchor, 1.0)
pos = (price - anchor) / span
acts = available_actions(price, last_counter_q, counter_locked)
rows = []
for a in acts:
cut = 0.0
if a.kind == "counter":
c = anchor + a.counter_q * span
cut = max(0.0, (price - c) / max(price, 1.0))
rows.append(a.feat(pos, cut))
return acts, np.stack(rows)
def run_episode(policy_fn, sup, rc: RewardCalculator, tf: np.ndarray, seed: int,
learner: Optional[FeatureDQNPolicy] = None, trace: Optional[list] = None):
"""policy_fn(state_feat, price_pos) → Action. learner 지정 시 replay 저장+학습."""
buyer = AutonomousBuyer(sup, seed)
p0 = buyer.price
env_rng = np.random.default_rng(seed + 7)
# 앵커율 샘플링(v3.1): 실운영처럼 앵커가 목표가 바로 아래(0.8~6%) — 좁은 스팬 지형에서 학습.
anchor = TARGET * (1.0 - float(env_rng.uniform(*ANCHOR_RATE_RANGE)))
span = max(TARGET - anchor, 1.0)
turn, settled, walked = 0, None, False
last_kind, last_q = "", 0.0 # 직전 역제안 기억 (같은 숫자 반복 방지의 학습 근거)
press_n = 0 # 설득 횟수 — 역제시 해금(MIN_PRESS) 카운터
# 견적 마감(환경 사실): 마감 도달 시 협상은 미타결 종료된다 — 룰이 아니라 세상의 시계.
deadline_turns = int(env_rng.integers(3, 11))
# 관측성 마스크(v3.5): 실서빙은 마감·이력·최저가가 '없는' 세션이 흔하고 로더가 중립값
# (0.5/0)을 대입한다. 시뮬이 항상 다 아는 세계만 학습하면 그 중립 상태가 분포 밖이 된다
# — v3.4 가 라이브 소액 지형에서 첫 턴 결렬로 퇴화한 원인 추정. 세계(마감 종료·상대 특성)는
# 그대로 두고 관측만 가린다: 마감은 40% 미관측(0.5 고정), 15% 는 전부 미상(신규 견적의 전형).
deadline_known = env_rng.random() < 0.6
blind = env_rng.random() < 0.15
if blind:
deadline_known = False
# 협력사 이력·최저가 특징 (에피소드 내 불변)
known_hist = buyer.hist_n and not blind
fixed_extra = dict(
hist_n=min(buyer.hist_n, 5) / 5.0 if not blind else 0.0,
hist_success=buyer.hist_success if known_hist else 0.5,
hist_settle=extra_settle(buyer.hist_settle_ratio) if known_hist else 0.5,
internet_gap=internet_gap_feat(buyer.internet_lowest or 0.0, anchor) if not blind else 0.0,
)
pending = None # (state_feat, action_feat) — 최종 결과 시점만 채점, 중간 r=0
while True:
turn += 1
price = buyer.price
deadline_remain = (max(0.0, (deadline_turns - turn + 1) / deadline_turns)
if deadline_known else 0.5) # 미관측 → 서빙 로더와 동일한 중립값
sf = np.concatenate([build_state_features(make_snapshot(sup, price, turn, p0, anchor)), tf,
extra_state(last_kind, last_q, deadline=deadline_remain, **fixed_extra)])
lcq = last_q if last_kind == "counter" else None
locked = lcq is None and press_n < MIN_PRESS
act = policy_fn(sf, price, anchor, lcq, locked)
if trace is not None:
trace.append((turn, int(price), act))
if act.kind == "accept":
settled = price
elif act.kind == "walk":
walked = True
elif act.kind == "counter":
c = anchor + act.counter_q * span
resp, val = buyer.on_counter(c, act.strategy, turn)
last_kind, last_q = "counter", act.counter_q # 역제안 기억 갱신
if resp == "accept":
settled = c
elif resp == "walk":
walked = True
else: # press
press_n += 1
left, _ = buyer.on_press(act.strategy, turn)
walked = walked or left
if not settled and not walked and turn >= buyer.patience:
walked = True # 인내심 소진(상대 특성) — 시스템 룰 아님
if not settled and not walked and turn >= deadline_turns:
walked = True # 견적 마감 도달(환경 사실) — 미타결 종료
done = settled is not None or walked
final_price = settled if settled is not None else buyer.price
# 성사 보너스는 목표가 이하 타결에만 — v3.1 이 '비싸게라도 성사'로 착취한 보상 구멍의
# 원인 차단(봉투 ① 의 마스크와 이중 방어: 유인 자체를 올바르게). 초과 타결 = 결렬 취급.
outcome = (NegotiationOutcome.SUCCESS if settled is not None and settled <= TARGET
else NegotiationOutcome.FAILURE if done else NegotiationOutcome.ONGOING)
r = rc.calculate(make_snapshot(sup, final_price, turn, p0, anchor, outcome)).total if done else 0.0
if learner is not None:
pos = (price - anchor) / span
cut = 0.0
if act.kind == "counter":
cut = max(0.0, (price - (anchor + act.counter_q * span)) / max(price, 1.0))
af = act.feat(pos, cut)
if pending:
nxt_lcq = last_q if last_kind == "counter" else None
learner.remember(*pending, 0.0, sf,
action_feats(price, anchor, nxt_lcq,
nxt_lcq is None and press_n < MIN_PRESS)[1], False)
pending = (sf, af)
if done:
learner.remember(sf, af, r, None, None, True)
learner.train_step()
if done:
return settled, turn, r
# ---- 정책들 ------------------------------------------------------------------
def dqn_policy(policy: FeatureDQNPolicy):
def f(sf, price, anchor, last_counter_q=None, counter_locked=False):
acts, feats = action_feats(price, anchor, last_counter_q, counter_locked)
i, _, _ = policy.select(sf, feats)
return acts[i]
return f
class RuleBaseline:
"""현행 시스템 룰 재현: 앵커 이하 수락 / 존내 1% 클로징 / 3회 압박 후 결렬."""
def __init__(self):
self.presses, self.closed = 0, False
def __call__(self, sf, price, anchor, last_counter_q=None, counter_locked=False) -> Action:
span = max(TARGET - anchor, 1.0)
if price <= anchor:
return Action("accept")
if price <= anchor * 1.02 and not self.closed:
self.closed = True
return Action("counter", (price * 0.99 - anchor) / span, 3)
if self.presses < 3:
self.presses += 1
return Action("press", 0.0, 3)
return Action("walk")
# ---- 학습/평가 ----------------------------------------------------------------
def evaluate(name, make_policy_fn, base_cfg, tcfg_state, episodes=3000, seed0=777):
rc = RewardCalculator(pref_config(base_cfg, 0.5), tcfg_state)
tf = build_tenant_features(pref_config(base_cfg, 0.5))
rng = np.random.default_rng(seed0)
rewards, settles, rounds = [], [], []
for i in range(episodes):
sup = sample_supplier(rng)
settled, turn, r = run_episode(make_policy_fn(), sup, rc, tf, seed0 * 91 + i)
rewards.append(r)
rounds.append(turn)
if settled is not None:
settles.append(settled / TARGET)
sr = len(settles) / episodes
print(f"{name:<14} 보상 {np.mean(rewards):.4f} ±{np.std(rewards)/np.sqrt(episodes):.4f}"
f" 성사율 {sr:.3f} 타결가/목표 {np.mean(settles):.3f} 평균라운드 {np.mean(rounds):.2f}")
return dict(reward=float(np.mean(rewards)), success=sr,
settle_ratio=float(np.mean(settles)) if settles else None, rounds=float(np.mean(rounds)))
def main(episodes=15000, seed=42):
random.seed(seed); np.random.seed(seed); torch.manual_seed(seed)
tcfg = TenantConfigLoader().load("ktcommerce")
policy = FeatureDQNPolicy(state_dim=STATE_FEATURE_DIM + TENANT_FEATURE_DIM + EXTRA_STATE_DIM,
card_dim=ACTION_DIM, eps_decay=5000, gamma=0.97)
rng = np.random.default_rng(seed)
print(f"=== 완전 자율 학습 {episodes}ep (행동 {len(ACTIONS)}개, 룰 0개) ===")
recent = []
for ep in range(1, episodes + 1):
sup = sample_supplier(rng)
rcfg, tf = sample_tenant_pref(rng, tcfg.reward)
rc = RewardCalculator(rcfg, tcfg.state)
_, _, r = run_episode(dqn_policy(policy), sup, rc, tf, seed * 131 + ep, learner=policy)
recent.append(r)
if ep % 3000 == 0:
print(f" ep {ep:>6} eps={policy.eps():.3f} 최근3000 평균보상={np.mean(recent[-3000:]):.4f}")
policy.save(CKPT_PATH)
print("\n=== 평가 3000ep (중립 성향 p=0.5, 동일 협력사 분포) ===")
policy.greedy = True
evaluate("룰시스템(현행)", lambda: RuleBaseline(), tcfg.reward, tcfg.state)
evaluate("완전자율 DQN", lambda: dqn_policy(policy), tcfg.reward, tcfg.state)
# 궤적 예시 — 에이전트가 룰 없이 뭘 하는지 눈으로
print("\n=== 궤적 예시 (완전자율) ===")
rc = RewardCalculator(pref_config(tcfg.reward, 0.5), tcfg.state)
tf = build_tenant_features(pref_config(tcfg.reward, 0.5))
rng2 = np.random.default_rng(7)
for k in range(3):
sup = sample_supplier(rng2)
trace = []
settled, turn, r = run_episode(dqn_policy(policy), sup, rc, tf, 5000 + k, trace=trace)
seg = "·".join(sup.segment)
print(f"[{seg}] " + "".join(
f"t{t}:{p:,}{a.kind}{'' if a.kind in ('accept', 'walk') else f'({a.counter_q:.2f},전략{a.strategy})' if a.kind == 'counter' else f'(전략{a.strategy})'}"
for t, p, a in trace) + f"{'타결 ' + format(int(settled), ',') + '' if settled else '결렬'} (r={r:.3f})")
if __name__ == "__main__":
main()