가게별 첫 화면 문구를 만드는 곳이 없어 모든 숙박이 공통 문구 10개를 돌려 썼다. - COPY 잡에 catchphrase 단계(prepare→generate→save→catchphrase→faq_fill), 숙박이 아니면 건너뜀 - prompts/catchphrase.py·gemini_text.generate_catchphrases: 캐치프레이즈 1 + 일반 20·계절 3×4·월 12·날씨 2×4. 문구마다 길이·상호명·ground_check(근거 없는 시설·수치) 검사 - fact tagline·catchphrases(lodging 스키마, allow_llm). 순환 문구는 사장님이 줄 단위로 고치게 글로 저장 (catchphrase_text: `봄 | ` · `9월 | ` · `비 | ` 머리) - site_payload: narrative.tagline·catchphrases 로 넘기고 이용정보 표(facts)에서는 뺀다 - 빌더 생성 화면 단계 문구, DATA_MODEL·GENERATION_FLOW 테스트 6건 추가. 관련 20개 파일 386 passed · 실패 102건은 main 과 목록이 같다(기존 실패) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
415 lines
17 KiB
Python
415 lines
17 KiB
Python
"""소개문·메타설명·FAQ 생성 — 겹들을 엮어 결과를 만드는 자리."""
|
|
import hashlib
|
|
import json
|
|
from dataclasses import dataclass, field
|
|
from typing import Optional
|
|
|
|
import httpx
|
|
|
|
from common.enums import PlaceCategory
|
|
from common.logger import LOG
|
|
from services.grounding.copy import FactInput, faq_polarity_ok, ground_check
|
|
from services.llm import provider
|
|
from services.llm.errors import LlmError
|
|
from services.llm.errors import LlmInvalidOutput as GeminiInvalidOutput
|
|
from services.llm.errors import LlmNotConfigured as GeminiNotConfigured
|
|
from services.prompts import catchphrase as catchphrase_prompt
|
|
from services.prompts.copy import RESPONSE_SCHEMA, build_prompt
|
|
|
|
|
|
def is_configured() -> bool:
|
|
"""호출측(copy_service.py, place_service.py 등)은 이 겹만 안다 — 어느 공급자가 활성인지는 몰라도 된다."""
|
|
return provider.active().is_configured()
|
|
|
|
|
|
@dataclass
|
|
class GeneratedFaq:
|
|
question: str
|
|
answer: str
|
|
fact_keys: list[str] = field(default_factory=list)
|
|
|
|
|
|
@dataclass
|
|
class GeneratedCopy:
|
|
"""생성 결과."""
|
|
|
|
intro: Optional[str] = None
|
|
intro_fact_keys: list[str] = field(default_factory=list)
|
|
meta_description: Optional[str] = None
|
|
faqs: list[GeneratedFaq] = field(default_factory=list)
|
|
rejected: list[tuple[str, str]] = field(default_factory=list)
|
|
source: str = "" # "openai:gpt-5.6-luna" 형식 — copy_steps.py 가 fact 출처 표기에 쓴다
|
|
|
|
|
|
def _unit_facts(unit_summaries: Optional[list[dict]]) -> list[FactInput]:
|
|
"""객실·프로그램 요약을 근거 fact 로 펼친다."""
|
|
out: list[FactInput] = []
|
|
for unit in unit_summaries or []:
|
|
name = str(unit.get("name") or "").strip()
|
|
labels = unit.get("labels") or {}
|
|
if name:
|
|
out.append(FactInput(key=f"unit:{name}", label="객실·프로그램명", value=name))
|
|
for key, value in (unit.get("facts") or {}).items():
|
|
if value is None or str(value).strip() == "":
|
|
continue
|
|
spec = labels.get(key) or {}
|
|
out.append(FactInput(
|
|
key=f"{name}:{key}" if name else key,
|
|
label=spec.get("label") or key,
|
|
value=str(value),
|
|
unit=spec.get("unit"),
|
|
))
|
|
return out
|
|
|
|
|
|
def _valid_keys(claimed: list, allowed: set[str]) -> list[str]:
|
|
"""모델이 적어준 근거 key 중 실제로 존재하는 것만 남긴다(없는 key 를 지어내기도 한다)."""
|
|
return [k for k in (claimed or []) if isinstance(k, str) and k in allowed]
|
|
|
|
|
|
async def generate_copy(
|
|
place_name: str,
|
|
category: PlaceCategory,
|
|
facts: list[FactInput],
|
|
*,
|
|
unit_summaries: Optional[list[dict]] = None,
|
|
records: Optional[list[str]] = None,
|
|
suggested_questions: Optional[list[str]] = None,
|
|
max_faqs: int = 8,
|
|
model: Optional[str] = None,
|
|
max_retries: int = 2,
|
|
client: Optional[httpx.AsyncClient] = None,
|
|
) -> GeneratedCopy:
|
|
"""확보된 fact 만으로 소개문·메타설명·FAQ 를 만든다."""
|
|
llm = provider.active()
|
|
if not llm.is_configured():
|
|
raise GeminiNotConfigured(f"{llm.__name__.rsplit('.', 1)[-1].upper()}_API_KEY 가 설정되지 않았다")
|
|
model = model or llm.DEFAULT_MODEL
|
|
# 사업장 fact 가 없어도 객실·메뉴 근거가 있으면 쓴다.
|
|
unit_grounding = _unit_facts(unit_summaries)
|
|
if not facts and not unit_grounding:
|
|
LOG.i(f"[llm-text] '{place_name}' 근거 fact 0건 — 생성하지 않는다(호출 없음)")
|
|
return GeneratedCopy(rejected=[("(전체)", "근거 fact 가 없다 — 생성하지 않았다")])
|
|
|
|
# 검증에 쓸 근거 = 넘겨받은 fact + 객실 요약 + 상호명(상호에 숫자가 있어도 근거로 본다)
|
|
grounding = list(facts) + unit_grounding
|
|
grounding.append(FactInput(key="place_name", label="상호명", value=place_name))
|
|
allowed_keys = {f.key for f in facts} | {f.key for f in grounding}
|
|
|
|
prompt = build_prompt(place_name, category, facts, max_faqs, unit_grounding, records, suggested_questions)
|
|
|
|
owns_client = client is None
|
|
client = client or httpx.AsyncClient(timeout=httpx.Timeout(120.0, connect=10.0))
|
|
try:
|
|
llm_result = await llm.generate(
|
|
client, model, prompt=prompt, response_schema=RESPONSE_SCHEMA, temperature=0.2, max_retries=max_retries,
|
|
)
|
|
finally:
|
|
if owns_client:
|
|
await client.aclose()
|
|
|
|
parsed = llm_result.json
|
|
usage = llm_result.usage
|
|
|
|
result = GeneratedCopy(source=f"{llm.__name__.rsplit('.', 1)[-1]}:{model}")
|
|
|
|
# ── 소개문 ──
|
|
intro = (parsed.get("intro") or "").strip()
|
|
if intro:
|
|
ok, reasons = ground_check(intro, grounding)
|
|
if ok:
|
|
result.intro = intro
|
|
result.intro_fact_keys = _valid_keys(parsed.get("intro_fact_keys"), allowed_keys)
|
|
else:
|
|
result.rejected.append((intro, " / ".join(reasons)))
|
|
|
|
# ── 메타 설명 ──
|
|
meta_desc = (parsed.get("meta_description") or "").strip()
|
|
if meta_desc:
|
|
ok, reasons = ground_check(meta_desc, grounding)
|
|
if ok:
|
|
result.meta_description = meta_desc
|
|
else:
|
|
result.rejected.append((meta_desc, " / ".join(reasons)))
|
|
|
|
# ── FAQ ── 항목마다 따로 검사한다.
|
|
for item in (parsed.get("faqs") or [])[:max_faqs]:
|
|
question = (item.get("question") or "").strip()
|
|
answer = (item.get("answer") or "").strip()
|
|
if not question or not answer:
|
|
continue
|
|
keys = _valid_keys(item.get("fact_keys"), allowed_keys)
|
|
if not keys:
|
|
# 근거를 못 대는 FAQ 는 버린다 — 사실인지 확인할 방법이 없다.
|
|
result.rejected.append((question, "근거 fact_keys 가 없다"))
|
|
continue
|
|
ok, reasons = ground_check(f"{question} {answer}", grounding)
|
|
# 질문은 주장이 아니라 값-반대 판정에서 빠진다.
|
|
polar_ok, polar_reasons = faq_polarity_ok(question, answer, grounding)
|
|
if not ok or not polar_ok:
|
|
result.rejected.append((question, " / ".join(reasons + polar_reasons)))
|
|
continue
|
|
result.faqs.append(GeneratedFaq(question=question, answer=answer, fact_keys=keys))
|
|
|
|
LOG.i(
|
|
f"[llm-text] '{place_name}' 생성 — 소개문 {'O' if result.intro else 'X'} · "
|
|
f"메타 {'O' if result.meta_description else 'X'} · FAQ {len(result.faqs)}건 · "
|
|
f"반려 {len(result.rejected)}건 · model={model} · "
|
|
f"tokens in={usage.input_tokens} out={usage.output_tokens} · 약 ${llm.price(model, usage)}"
|
|
)
|
|
return result
|
|
|
|
|
|
# ── 첫 화면 문구(generate_catchphrases) ─────────────────────────────────
|
|
@dataclass
|
|
class GeneratedCatchphrases:
|
|
tagline: Optional[str] = None
|
|
items: list[dict] = field(default_factory=list)
|
|
rejected: list[tuple[str, str]] = field(default_factory=list)
|
|
source: str = ""
|
|
|
|
|
|
async def generate_catchphrases(
|
|
place_name: str,
|
|
facts: list[FactInput],
|
|
*,
|
|
records: Optional[list[str]] = None,
|
|
model: Optional[str] = None,
|
|
max_retries: int = 2,
|
|
client: Optional[httpx.AsyncClient] = None,
|
|
) -> GeneratedCatchphrases:
|
|
"""고정 캐치프레이즈 하나와 순환 문구 모음 — 문구마다 소개문과 같은 근거 검사를 거친다."""
|
|
llm = provider.active()
|
|
if not llm.is_configured():
|
|
raise GeminiNotConfigured(f"{llm.__name__.rsplit('.', 1)[-1].upper()}_API_KEY 가 설정되지 않았다")
|
|
model = model or llm.DEFAULT_MODEL
|
|
grounding = list(facts) + [FactInput(key="place_name", label="상호명", value=place_name)]
|
|
prompt = catchphrase_prompt.build_prompt(place_name, facts, records)
|
|
|
|
owns_client = client is None
|
|
client = client or httpx.AsyncClient(timeout=httpx.Timeout(120.0, connect=10.0))
|
|
try:
|
|
llm_result = await llm.generate(
|
|
client, model, prompt=prompt, response_schema=catchphrase_prompt.RESPONSE_SCHEMA,
|
|
temperature=0.7, max_retries=max_retries,
|
|
)
|
|
finally:
|
|
if owns_client:
|
|
await client.aclose()
|
|
|
|
parsed = llm_result.json or {}
|
|
result = GeneratedCatchphrases(source=f"{llm.__name__.rsplit('.', 1)[-1]}:{model}")
|
|
seen: set[str] = set()
|
|
|
|
def accept(text: object) -> Optional[str]:
|
|
line = str(text or "").strip()
|
|
if not line or line in seen:
|
|
return None
|
|
if not (catchphrase_prompt.MIN_LEN <= len(line) <= catchphrase_prompt.MAX_LEN):
|
|
result.rejected.append((line, f"길이 {len(line)}자"))
|
|
return None
|
|
if place_name and place_name in line:
|
|
result.rejected.append((line, "상호명 포함"))
|
|
return None
|
|
ok, reasons = ground_check(line, grounding)
|
|
if not ok:
|
|
result.rejected.append((line, " / ".join(reasons)))
|
|
return None
|
|
seen.add(line)
|
|
return line
|
|
|
|
result.tagline = accept(parsed.get("tagline"))
|
|
for text in (parsed.get("general") or [])[: catchphrase_prompt.COUNT_GENERAL]:
|
|
if line := accept(text):
|
|
result.items.append({"text": line, "kind": "general"})
|
|
for row in parsed.get("season") or []:
|
|
season = (row or {}).get("season")
|
|
if season in catchphrase_prompt.SEASONS and (line := accept(row.get("text"))):
|
|
result.items.append({"text": line, "kind": "season", "season": season})
|
|
for row in parsed.get("month") or []:
|
|
month = (row or {}).get("month")
|
|
if isinstance(month, int) and 1 <= month <= 12 and (line := accept(row.get("text"))):
|
|
result.items.append({"text": line, "kind": "month", "month": month})
|
|
for row in parsed.get("weather") or []:
|
|
weather = (row or {}).get("weather")
|
|
if weather in catchphrase_prompt.WEATHERS and (line := accept(row.get("text"))):
|
|
result.items.append({"text": line, "kind": "weather", "weather": weather})
|
|
|
|
usage = llm_result.usage
|
|
LOG.i(
|
|
f"[llm-text] '{place_name}' 첫 화면 문구 — 캐치프레이즈 {'O' if result.tagline else 'X'} · "
|
|
f"순환 {len(result.items)}개 · 반려 {len(result.rejected)}건 · model={model} · "
|
|
f"tokens in={usage.input_tokens} out={usage.output_tokens} · 약 ${llm.price(model, usage)}"
|
|
)
|
|
return result
|
|
|
|
|
|
# ── 요약(summarize_text) ──────────────────────────────────────────────────
|
|
_SUMMARY_CACHE: dict[str, str] = {}
|
|
_SUMMARY_CACHE_MAX = 500
|
|
_SUMMARY_PROMPT = (
|
|
"다음 숙소 소개에서 핵심 특징 1~2개만 골라 한국어 한 문장, 공백 포함 60~80자로 요약해줘. "
|
|
"원문에 없는 사실이나 과장 표현을 추가하지 말고, 선택한 사실의 조건과 부정 표현을 유지해. "
|
|
"반복되는 상호명, 인사말, 홍보 수식어는 생략해. "
|
|
"요약문만 출력하고 다른 말은 붙이지 마.\n\n"
|
|
)
|
|
|
|
|
|
async def summarize_text(
|
|
text: str,
|
|
*,
|
|
model: Optional[str] = None,
|
|
max_retries: int = 2,
|
|
client: Optional[httpx.AsyncClient] = None,
|
|
) -> Optional[str]:
|
|
"""캔버스 미리보기용 축약문."""
|
|
stripped = text.strip()
|
|
if not stripped:
|
|
return None
|
|
llm = provider.active()
|
|
if not llm.is_configured():
|
|
return None
|
|
model = model or llm.DEFAULT_MODEL
|
|
|
|
# 길이 기준을 바꾼 뒤 이전 길이의 요약을 재사용하지 않도록 프롬프트도 키에 넣는다.
|
|
cache_key = hashlib.sha256((_SUMMARY_PROMPT + stripped).encode("utf-8")).hexdigest()
|
|
cached = _SUMMARY_CACHE.get(cache_key)
|
|
if cached is not None:
|
|
return cached
|
|
|
|
owns_client = client is None
|
|
client = client or httpx.AsyncClient(timeout=httpx.Timeout(60.0, connect=10.0))
|
|
try:
|
|
result = await llm.generate(client, model, prompt=_SUMMARY_PROMPT + stripped, temperature=0.2, max_retries=max_retries)
|
|
summary = result.text.strip()
|
|
except LlmError as ex:
|
|
LOG.w(f"[llm-text] 요약 실패: {ex}")
|
|
return None
|
|
finally:
|
|
if owns_client:
|
|
await client.aclose()
|
|
|
|
if not summary:
|
|
return None
|
|
usage = result.usage
|
|
LOG.i(
|
|
f"[llm-text] 요약 {len(stripped)}자 → {len(summary)}자 · model={model} · "
|
|
f"tokens in={usage.input_tokens} out={usage.output_tokens} · 약 ${llm.price(model, usage)}"
|
|
)
|
|
if len(_SUMMARY_CACHE) >= _SUMMARY_CACHE_MAX:
|
|
_SUMMARY_CACHE.clear() # 간단한 캐시 상한 — 관리 도구 트래픽 규모에는 LRU 가 과하다.
|
|
_SUMMARY_CACHE[cache_key] = summary
|
|
return summary
|
|
|
|
|
|
@dataclass
|
|
class GeneratedSong:
|
|
"""가사 생성 결과."""
|
|
|
|
title: str
|
|
lyrics: str
|
|
style: str
|
|
|
|
|
|
async def generate_song(
|
|
place_name: str,
|
|
category: PlaceCategory,
|
|
*,
|
|
region: str,
|
|
grounding: list[str],
|
|
intro: str = "",
|
|
model: Optional[str] = None,
|
|
max_retries: int = 2,
|
|
client: Optional[httpx.AsyncClient] = None,
|
|
) -> GeneratedSong:
|
|
"""이 업소의 노래 가사를 쓴다."""
|
|
llm = provider.active()
|
|
if not llm.is_configured():
|
|
raise GeminiNotConfigured("API 키가 설정되지 않았다")
|
|
if not grounding and not (intro or "").strip():
|
|
raise GeminiInvalidOutput("가사를 쓸 재료가 없다 — 확인된 fact 도 소개문도 없다")
|
|
model = model or llm.DEFAULT_MODEL
|
|
|
|
from common.category_schema import get_schema
|
|
from services.prompts.song import RESPONSE_SCHEMA as SONG_SCHEMA, build_prompt as build_song_prompt
|
|
|
|
prompt = build_song_prompt(place_name, get_schema(category).label, region, grounding, intro)
|
|
|
|
owns_client = client is None
|
|
client = client or httpx.AsyncClient(timeout=httpx.Timeout(120.0, connect=10.0))
|
|
try:
|
|
llm_result = await llm.generate(
|
|
client, model, prompt=prompt, response_schema=SONG_SCHEMA, temperature=0.9, max_retries=max_retries,
|
|
)
|
|
finally:
|
|
if owns_client:
|
|
await client.aclose()
|
|
|
|
parsed = llm_result.json or {}
|
|
title = (parsed.get("title") or "").strip()
|
|
lyrics = (parsed.get("lyrics") or "").strip()
|
|
style = (parsed.get("style") or "").strip()
|
|
if not lyrics:
|
|
raise GeminiInvalidOutput("가사가 비어 있다")
|
|
|
|
usage = llm_result.usage
|
|
LOG.i(
|
|
f"[llm-text] '{place_name}' 가사 — '{title}' ({style}) · {len(lyrics)}자 · "
|
|
f"tokens in={usage.input_tokens} out={usage.output_tokens} · 약 ${llm.price(model, usage)}"
|
|
)
|
|
# 제목이 비면 상호를 쓴다 — 빈 제목은 플레이어에서 빈 줄로 보인다.
|
|
return GeneratedSong(title=title or place_name, lyrics=lyrics, style=style or "acoustic ballad")
|
|
|
|
|
|
async def generate_social_post(place_name, facts, link_url, provider=2, *, client=None,
|
|
region='', category=''):
|
|
"""실제 게시 문자열을 검증한다."""
|
|
from services.prompts import social
|
|
from services.external.social import adapter, weighted_length, URL
|
|
from services.llm import provider as llm_provider
|
|
import unicodedata
|
|
|
|
if not facts:
|
|
raise GeminiInvalidOutput('NO_GROUNDED_FACTS')
|
|
llm = llm_provider.active()
|
|
if not llm.is_configured():
|
|
raise GeminiNotConfigured(f"{llm.__name__.rsplit('.', 1)[-1].upper()}_API_KEY 가 설정되지 않았다")
|
|
limit = adapter(provider).weighted_limit()
|
|
allowed = {f.key for f in facts}
|
|
feedback = ''
|
|
owns = client is None
|
|
client = client or httpx.AsyncClient(timeout=45)
|
|
try:
|
|
for _ in range(3):
|
|
prompt = social.build_prompt(
|
|
place_name, facts, limit - weighted_length('\n\n' + link_url, provider), feedback,
|
|
region=region, category=category)
|
|
result = await llm.generate(
|
|
client, llm.DEFAULT_MODEL, prompt=prompt,
|
|
response_schema=social.RESPONSE_SCHEMA, temperature=0.2, max_retries=0,
|
|
)
|
|
try:
|
|
parsed = result.json if result.json is not None else json.loads(result.text)
|
|
body = unicodedata.normalize('NFC', parsed['body'].strip())
|
|
keys = parsed['fact_keys']
|
|
text = body + '\n\n' + link_url
|
|
# ★ 지역·업종은 fact 가 아니라 검증된 place 값이다. 근거에 얹지 않으면
|
|
# 본문에 쓴 순간 "근거 없는 주장" 으로 반려되고 3회 재시도를 태우고 실패한다 —
|
|
# 상호명을 이렇게 다루는 그 방식 그대로다.
|
|
grounds = facts + [FactInput(key='name', label='상호명', value=place_name)]
|
|
if region:
|
|
grounds += [FactInput(key='region', label='지역', value=region)]
|
|
if category:
|
|
grounds += [FactInput(key='category', label='업종', value=category)]
|
|
ok, _ = ground_check(body, grounds)
|
|
if (body and isinstance(keys, list) and keys and all(k in allowed for k in keys)
|
|
and ok and not URL.search(body) and weighted_length(text, provider) <= limit):
|
|
return text
|
|
except (ValueError, KeyError, TypeError):
|
|
pass
|
|
feedback = '이전 응답은 길이 또는 근거 검증에 실패했다. 더 짧게, 제공된 사실만으로 다시 써라.'
|
|
raise GeminiInvalidOutput('SOCIAL_INVALID_OUTPUT')
|
|
finally:
|
|
if owns:
|
|
await client.aclose()
|