가게별 첫 화면 문구를 만드는 곳이 없어 모든 숙박이 공통 문구 10개를 돌려 썼다. - COPY 잡에 catchphrase 단계(prepare→generate→save→catchphrase→faq_fill), 숙박이 아니면 건너뜀 - prompts/catchphrase.py·gemini_text.generate_catchphrases: 캐치프레이즈 1 + 일반 20·계절 3×4·월 12·날씨 2×4. 문구마다 길이·상호명·ground_check(근거 없는 시설·수치) 검사 - fact tagline·catchphrases(lodging 스키마, allow_llm). 순환 문구는 사장님이 줄 단위로 고치게 글로 저장 (catchphrase_text: `봄 | ` · `9월 | ` · `비 | ` 머리) - site_payload: narrative.tagline·catchphrases 로 넘기고 이용정보 표(facts)에서는 뺀다 - 빌더 생성 화면 단계 문구, DATA_MODEL·GENERATION_FLOW 테스트 6건 추가. 관련 20개 파일 386 passed · 실패 102건은 main 과 목록이 같다(기존 실패) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
328 lines
13 KiB
Python
328 lines
13 KiB
Python
"""소개문·FAQ 생성 — COPY 잡이 하는 일."""
|
|
import uuid
|
|
from dataclasses import dataclass
|
|
|
|
from common.category_schema import CategorySchema, CategorySchemaError, get_schema
|
|
from common.database.db_session_manager import DB_SESSION_MNG
|
|
from common.faq_catalog import FaqCatalog, find_catalog
|
|
from common.database.model.models import place_facts, place_faqs, place_channels, places, place_units
|
|
from common.enums import (
|
|
DBWRType,
|
|
ErrorType,
|
|
FactStatus,
|
|
PlaceCategory,
|
|
SourceType,
|
|
)
|
|
from common.logger import LOG
|
|
from common.models.gmodel import UserInfo
|
|
from common.utils.gtime import GTime
|
|
from config.server_configs import external_api_config
|
|
from crud.fact_crud import FactCRUD
|
|
from crud.faq_crud import FaqCRUD
|
|
from crud.place_crud import PlaceCRUD
|
|
from router.v1.fact.protocol import Req_UpsertFact
|
|
from services import catchphrase_text, faq_fill, place_research
|
|
from services.external import gemini_text
|
|
from services.llm import provider
|
|
from services.fact_service import FactService
|
|
from common.job_errors import PermanentJobError
|
|
|
|
_fact_crud = FactCRUD()
|
|
_faq_crud = FaqCRUD()
|
|
_place_crud = PlaceCRUD()
|
|
|
|
|
|
class CopyAborted(PermanentJobError):
|
|
"""재시도해도 소용없는 중단 — 잡의 last_error 로 남는다."""
|
|
|
|
|
|
@dataclass
|
|
class CopyInputs:
|
|
place: places
|
|
schema: CategorySchema
|
|
grounded: list[gemini_text.FactInput]
|
|
records: list[str]
|
|
unit_summaries: list[dict]
|
|
catalog: FaqCatalog | None
|
|
known_fact_keys: set[str]
|
|
|
|
@property
|
|
def ungrounded(self) -> bool:
|
|
return not self.grounded and not self.unit_summaries
|
|
|
|
|
|
async def prepare_copy(place_id: str, owner_user_id: str) -> CopyInputs:
|
|
|
|
err, place = await DB_SESSION_MNG.execute_lambda(
|
|
places.DBType(),
|
|
DBWRType.DB_READ.value,
|
|
lambda s: _place_crud.get_place(s, uuid.UUID(owner_user_id), uuid.UUID(place_id)),
|
|
)
|
|
if err != ErrorType.SUCCESS or place is None:
|
|
raise CopyAborted(f"사업장을 찾을 수 없다: {place_id}")
|
|
|
|
try:
|
|
schema = get_schema(PlaceCategory(place.category))
|
|
except (CategorySchemaError, ValueError) as ex:
|
|
raise CopyAborted(f"지원하지 않는 업종: {place.category}") from ex
|
|
|
|
# 노출 가능한 fact 만 근거로 준다.
|
|
pid = uuid.UUID(place_id)
|
|
f_err, fact_rows = await DB_SESSION_MNG.execute_lambda(
|
|
place_facts.DBType(),
|
|
DBWRType.DB_READ.value,
|
|
lambda s: _fact_crud.list_facts(s, pid, None, None, True, True),
|
|
)
|
|
if f_err != ErrorType.SUCCESS:
|
|
raise CopyAborted(f"fact 조회 실패: {f_err.name}")
|
|
|
|
grounded = [
|
|
gemini_text.FactInput(
|
|
key=r.key,
|
|
label=(schema.get(r.key).label if schema.get(r.key) else r.key),
|
|
value=r.value,
|
|
unit=r.unit,
|
|
)
|
|
for r in fact_rows
|
|
if r.unit_id is None and (r.value or "").strip()
|
|
]
|
|
# 수집 원문도 근거로 넘긴다 — fact 가 아니라 place_channels.raw 에 박제된 글이다.
|
|
records: list[str] = []
|
|
l_err, link_rows = await DB_SESSION_MNG.execute_lambda(
|
|
place_channels.DBType(),
|
|
DBWRType.DB_READ.value,
|
|
lambda s: _place_crud.list_links(s, pid, False),
|
|
)
|
|
if l_err == ErrorType.SUCCESS:
|
|
for link in (link_rows or []):
|
|
raw = link.raw if isinstance(link.raw, dict) else {}
|
|
if link.confirmed_at is None and raw.get("kind") != place_research.RAW_KIND:
|
|
continue
|
|
text = (raw.get("text") or "").strip()
|
|
if text:
|
|
# fact 목록이 아니라 records 로 넘긴다.
|
|
records.append(text[:4000])
|
|
# ground_check 는 여전히 이 글을 근거로 인정해야 한다 — 근거 목록에도 남긴다.
|
|
grounded.append(gemini_text.FactInput(
|
|
key=f"source:{link.link_id}", label="수집 원문", value=text[:4000],
|
|
))
|
|
|
|
# 객실·메뉴 요약도 근거로 넘긴다 — "최대 4명" 같은 수치가 통과하려면 근거에 있어야 한다.
|
|
u_err, unit_rows = await DB_SESSION_MNG.execute_lambda(
|
|
place_units.DBType(),
|
|
DBWRType.DB_READ.value,
|
|
lambda s: _place_crud.list_units(s, pid),
|
|
)
|
|
unit_summaries = []
|
|
if u_err == ErrorType.SUCCESS:
|
|
by_unit: dict = {}
|
|
for r in fact_rows:
|
|
if r.unit_id and (r.value or "").strip():
|
|
by_unit.setdefault(str(r.unit_id), {})[r.key] = r.value
|
|
unit_summaries = [
|
|
{
|
|
"name": u.name,
|
|
"facts": by_unit.get(str(u.unit_id), {}),
|
|
# 스키마 라벨·단위를 같이 넘긴다 — 이게 없으면 프롬프트에 'weekday_price' 라는 날 key 가 그대로 실려 모델이 그 낱말로 문장을 쓴다.
|
|
"labels": {
|
|
key: {
|
|
"label": schema.get(key).label if schema.get(key) else key,
|
|
"unit": schema.get(key).unit if schema.get(key) else None,
|
|
}
|
|
for key in by_unit.get(str(u.unit_id), {})
|
|
},
|
|
}
|
|
for u in unit_rows
|
|
if by_unit.get(str(u.unit_id))
|
|
]
|
|
|
|
# FAQ 채우기에 쓸 업종 카탈로그.
|
|
catalog = find_catalog(place.category, place.external_category)
|
|
# 사업장·객실 fact 를 가리지 않는다 — "기준 인원" 은 객실 fact 로 답한다.
|
|
known_fact_keys = {r.key for r in fact_rows if (r.value or "").strip()}
|
|
|
|
return CopyInputs(place, schema, grounded, records, unit_summaries, catalog, known_fact_keys)
|
|
|
|
|
|
async def generate_copy(inputs: CopyInputs) -> gemini_text.GeneratedCopy:
|
|
active_provider = provider.active()
|
|
model = (
|
|
external_api_config.openai_text_model if active_provider.__name__.endswith("openai")
|
|
else external_api_config.gemini_text_model
|
|
)
|
|
try:
|
|
return await gemini_text.generate_copy(
|
|
inputs.place.name,
|
|
PlaceCategory(inputs.place.category),
|
|
inputs.grounded,
|
|
unit_summaries=inputs.unit_summaries or None,
|
|
records=inputs.records or None,
|
|
suggested_questions=faq_fill.suggested_questions(inputs.catalog, inputs.known_fact_keys) if inputs.catalog else None,
|
|
max_faqs=faq_fill.FAQ_TARGET,
|
|
model=model,
|
|
)
|
|
except gemini_text.GeminiNotConfigured as ex:
|
|
raise CopyAborted(str(ex)) from ex
|
|
|
|
|
|
async def save_copy(inputs: CopyInputs, copy: gemini_text.GeneratedCopy | None) -> dict:
|
|
pid = inputs.place.place_id
|
|
place_id = str(pid)
|
|
schema = inputs.schema
|
|
now = GTime.UTC()
|
|
stat = {
|
|
"place_id": place_id,
|
|
"grounded_facts": len(inputs.grounded),
|
|
"intro": False,
|
|
"meta": False,
|
|
"faqs": 0,
|
|
"faq_fill": 0, # 목표 수를 채운 문의 안내 문항 수
|
|
# 반려된 문장을 그대로 남긴다 — 소개문이 왜 안 나왔는지 운영자가 알아야 한다.
|
|
"rejected": [list(r) for r in (copy.rejected or [])][:20] if copy else [],
|
|
}
|
|
|
|
if copy is None:
|
|
# 키만 없는 경우에는 기존 생성물을 보존한다.
|
|
if inputs.ungrounded and inputs.catalog is not None:
|
|
await DB_SESSION_MNG.execute_lambda_claim(
|
|
place_faqs.DBType(), lambda s: _faq_crud.expire_generated(s, pid, now),
|
|
)
|
|
return stat
|
|
|
|
actor = UserInfo(
|
|
# 잡이 쓰는 신원.
|
|
user_id=str(inputs.place.owner_user_id),
|
|
id="generator",
|
|
role=1,
|
|
)
|
|
service = FactService(_fact_crud, _place_crud)
|
|
|
|
# 소개문·메타는 fact 로 들어간다 — FactService 가 allow_llm 을 다시 확인하고(뒷문 없음), LLM 출처라 후보가 아니라 노출값으로 앉힌다(upsert_fact 의 LLM 분기).
|
|
for key, text_value in (("intro", copy.intro), ("meta_description", copy.meta_description)):
|
|
if not (text_value or "").strip():
|
|
continue
|
|
if not (schema.get(key) and schema.get(key).allow_llm):
|
|
# 이 업종 스키마가 LLM 작성을 허용하지 않는 필드다.
|
|
continue
|
|
res = await service.upsert_fact(
|
|
actor, place_id,
|
|
Req_UpsertFact(
|
|
key=key, value=text_value.strip(),
|
|
source_type=SourceType.LLM, source_url=f"llm:{copy.source or external_api_config.gemini_text_model}",
|
|
),
|
|
)
|
|
if res.result.success:
|
|
stat["intro" if key == "intro" else "meta"] = True
|
|
else:
|
|
stat["rejected"].append([key, res.result.desc])
|
|
|
|
# 확인 안 된 기존 생성 FAQ 는 내리고 새로 넣는다.
|
|
await DB_SESSION_MNG.execute_lambda_claim(
|
|
place_faqs.DBType(),
|
|
lambda s: _faq_crud.expire_generated(s, pid, now),
|
|
)
|
|
for order, faq in enumerate(copy.faqs or []):
|
|
if not faq.fact_keys:
|
|
# 근거 없는 FAQ 는 저장하지 않는다.
|
|
stat["rejected"].append([faq.question, "근거 fact 없음"])
|
|
continue
|
|
row = place_faqs(
|
|
place_id=pid,
|
|
question=faq.question,
|
|
answer=faq.answer,
|
|
source_fact_ids=list(faq.fact_keys),
|
|
generated_by=SourceType.LLM.value,
|
|
# 바로 노출한다.
|
|
status=FactStatus.VERIFIED.value,
|
|
sort_order=order,
|
|
)
|
|
run_err = await DB_SESSION_MNG.execute_lambda_run(
|
|
[place_faqs.DBType()],
|
|
[lambda s, r=row: _faq_crud.add_faq(s, r)],
|
|
)
|
|
if run_err == ErrorType.SUCCESS:
|
|
stat["faqs"] += 1
|
|
|
|
LOG.i(f"[copy] place={place_id} 소개문 {'O' if stat['intro'] else 'X'} · FAQ {stat['faqs']}건 · "
|
|
f"반려 {len(stat['rejected'])}건 (근거 fact {len(inputs.grounded)}개)")
|
|
return stat
|
|
|
|
|
|
async def generate_catchphrases(inputs: CopyInputs) -> gemini_text.GeneratedCatchphrases:
|
|
active_provider = provider.active()
|
|
model = (
|
|
external_api_config.openai_text_model if active_provider.__name__.endswith("openai")
|
|
else external_api_config.gemini_text_model
|
|
)
|
|
try:
|
|
return await gemini_text.generate_catchphrases(
|
|
inputs.place.name, inputs.grounded, records=inputs.records or None, model=model,
|
|
)
|
|
except gemini_text.GeminiNotConfigured as ex:
|
|
raise CopyAborted(str(ex)) from ex
|
|
|
|
|
|
async def save_catchphrases(inputs: CopyInputs, generated: gemini_text.GeneratedCatchphrases) -> dict:
|
|
place_id = str(inputs.place.place_id)
|
|
stat = {"tagline": False, "catchphrases": 0, "rejected": [list(r) for r in generated.rejected][:20]}
|
|
actor = UserInfo(user_id=str(inputs.place.owner_user_id), id="generator", role=1)
|
|
service = FactService(_fact_crud, _place_crud)
|
|
values = (("tagline", generated.tagline), ("catchphrases", catchphrase_text.encode(generated.items)))
|
|
for key, value in values:
|
|
spec = inputs.schema.get(key)
|
|
if not (value or "").strip() or not (spec and spec.allow_llm):
|
|
continue
|
|
res = await service.upsert_fact(
|
|
actor, place_id,
|
|
Req_UpsertFact(
|
|
key=key, value=value.strip(),
|
|
source_type=SourceType.LLM, source_url=f"llm:{generated.source or external_api_config.gemini_text_model}",
|
|
),
|
|
)
|
|
if not res.result.success:
|
|
stat["rejected"].append([key, res.result.desc])
|
|
elif key == "tagline":
|
|
stat["tagline"] = True
|
|
else:
|
|
stat["catchphrases"] = len(generated.items)
|
|
LOG.i(f"[copy] place={place_id} 캐치프레이즈 {'O' if stat['tagline'] else 'X'} · 순환 {stat['catchphrases']}개")
|
|
return stat
|
|
|
|
|
|
async def fill_faqs(pid: uuid.UUID, catalog: FaqCatalog, known_fact_keys: set[str], phone: str | None) -> int:
|
|
"""노출 중인 FAQ 가 목표 수에 모자란 만큼 문의 안내 문항을 넣는다."""
|
|
l_err, rows = await DB_SESSION_MNG.execute_lambda(
|
|
place_faqs.DBType(),
|
|
DBWRType.DB_READ.value,
|
|
lambda s: _faq_crud.list_faqs(s, pid, True),
|
|
)
|
|
if l_err != ErrorType.SUCCESS:
|
|
LOG.e_no_callstack(f"[copy] FAQ 채우기 건너뜀 — 목록 조회 실패 place={pid} {l_err.name}")
|
|
return 0
|
|
|
|
picks = faq_fill.pick_fill_faqs(
|
|
catalog,
|
|
[faq_fill.ExistingFaq(r.question, r.source_fact_ids) for r in rows],
|
|
known_fact_keys,
|
|
phone,
|
|
)
|
|
next_order = max((r.sort_order for r in rows), default=-1) + 1
|
|
added = 0
|
|
for offset, pick in enumerate(picks):
|
|
row = place_faqs(
|
|
place_id=pid,
|
|
question=pick.question,
|
|
answer=pick.answer,
|
|
source_fact_ids=None,
|
|
generated_by=SourceType.TEMPLATE.value,
|
|
status=FactStatus.VERIFIED.value,
|
|
sort_order=next_order + offset,
|
|
)
|
|
run_err = await DB_SESSION_MNG.execute_lambda_run(
|
|
[place_faqs.DBType()],
|
|
[lambda s, r=row: _faq_crud.add_faq(s, r)],
|
|
)
|
|
if run_err == ErrorType.SUCCESS:
|
|
added += 1
|
|
return added
|