[feat] solution/backend: 발행 때 온톨로지 키워드 생성을 켠다 — 호텔 판정 · 업종별 유형어 게이트

generate:false 로 불러 새 숙소가 사전 밖 단어를 받지 못했다. 생성을 켜면 호텔이 펜션 사전의
`군산 독채펜션` 을 받는데, 메타 게이트가 '펜션' 을 업종어로 보고 자료 없이 통과시켰다.

- site_ontology: generate+sync — 처음 보는 숙소만 저쪽이 기다리게 하고 publish 상한 60초
- seo_keywords: 외부 분류(없으면 상호)에 '호텔' 이면 stay.hotel, 자기 유형어만 자료 없이 통과
- snapshot: place.external_category 를 실음 — 호텔 판정 근거
- .env.example · ARCHITECTURE · DEVLOG: ONTOLOGY_LLM_PROVIDER, 발행 흐름, 실측

test_seo_keywords 20 passed(호텔 4건 추가). 전체 1041 passed / 47 failed —
47건은 이 변경 전 main 과 같은 목록

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This commit is contained in:
hbyang 2026-10-01 13:49:37 +09:00
parent 847943a8bc
commit 94c4116689
7 changed files with 86 additions and 20 deletions

View File

@ -47,6 +47,8 @@ SUNO_CALLBACK_URL=https://example.com/api/suno/callback
# ★ 워커가 부르는 주소다. compose 로 띄우면 컨테이너 안에서 보는 주소(http://host.docker.internal:3100), # ★ 워커가 부르는 주소다. compose 로 띄우면 컨테이너 안에서 보는 주소(http://host.docker.internal:3100),
# 백엔드를 네이티브로 돌리면 http://127.0.0.1:3100 # 백엔드를 네이티브로 돌리면 http://127.0.0.1:3100
SITE_ONTOLOGY_URL= SITE_ONTOLOGY_URL=
# 새 숙소 키워드를 OpenAI 로 만든다(OPENAI_API_KEY 를 함께 쓴다). mock 이면 생성분이 매칭에 나가지 않는다.
ONTOLOGY_LLM_PROVIDER=mock
# 프리렌더가 절대 굽지 않는 슬러그(쉼표 구분). 손으로 만든 목업(/s/stay·stay2·stay3·stay4·stay5) # 프리렌더가 절대 굽지 않는 슬러그(쉼표 구분). 손으로 만든 목업(/s/stay·stay2·stay3·stay4·stay5)
# 이름과 같은 슬러그로 실제 발행이 생기면 그 payload 로 목업을 덮어 구워버린다 — 비우지 않는다. # 이름과 같은 슬러그로 실제 발행이 생기면 그 payload 로 목업을 덮어 구워버린다 — 비우지 않는다.

View File

@ -42,7 +42,7 @@ React 를 렌더해야 하고, 그때부터 디자인 수정에 백엔드 배포
BUILD 잡 (worker) ─ services/build_service.py:99 run_build() BUILD 잡 (worker) ─ services/build_service.py:99 run_build()
├ build_snapshot → 원본 데이터(JSONB) ├ build_snapshot → 원본 데이터(JSONB)
├ seo_keywords.fetch() → SiteOntology 추천을 이 가게 자료로 거른 키워드 → snapshot["seo"] ├ seo_keywords.fetch() → SiteOntology 추천을 이 가게 자료로 거른 키워드 → snapshot["seo"]
│ (숙박만 · 설정 없거나 실패하면 생략 · 발행은 계속) │ (숙박만 · 처음 보는 숙소는 OpenAI 키워드 생성을 기다림 · 실패하면 생략 · 발행은 계속)
│ → site_versions 행 insert │ → site_versions 행 insert
├ 1차 게이트 (DB 사실 기준) → publish_gate.evaluate() ├ 1차 게이트 (DB 사실 기준) → publish_gate.evaluate()
├ site_payload.emit_payload() → out/payloads/<slug>.json ★ 백엔드의 유일한 산출물 ├ site_payload.emit_payload() → out/payloads/<slug>.json ★ 백엔드의 유일한 산출물

View File

@ -4,6 +4,25 @@
**나중에 같은 실수를 막아 주는 것**(결정의 이유·밟은 함정·실측값)만 남긴다. **나중에 같은 실수를 막아 주는 것**(결정의 이유·밟은 함정·실측값)만 남긴다.
2026-09-29에 요약본으로 다시 썼다. 원문 전체는 git 히스토리(이 파일의 09-29 이전 버전)에 있다. 2026-09-29에 요약본으로 다시 썼다. 원문 전체는 git 히스토리(이 파일의 09-29 이전 버전)에 있다.
## 2026-10-01 — 온톨로지: 새 숙소 키워드를 OpenAI 로 쌓는다 · 호텔 업종 · 숙소 검색
- 발행이 `generate:false` 로 불러 사전이 고정이었다 — 사전에 없는 숙소는 새 단어를 못 받았다.
이제 **처음 보는 숙소는 첫 발행에서 생성을 기다린다**(실측 7초, 키워드 15·질문 5). 재빌드·재발행은
30일 안이면 부르지 않는다 — 재빌드마다 부르는 창구라 조건이 없으면 저장마다 요금이 나간다.
- 생성분이 매칭에 안 나오던 이유는 `MATCH_SOURCES` 가 `llm` 을 뺀 것. 업종 뿌리(`stay`)로 거르고
brand 는 그 업체 것만 받게 해서 넣었다. **mock 생성분은 `source='mock'`** — compose 기본이 mock 이라
키 없이 띄운 환경마다 가짜 단어가 발행본에 나갈 뻔했다.
- 호텔(`stay.hotel`): 외부 분류(`places.external_category`, 없으면 상호)에 "호텔" 이면. 매칭은 펜션 사전도
받되 메타 게이트가 **자기 유형어만** 자료 없이 통과시킨다 — 호텔에 `군산 독채펜션` 이 붙지 않는다.
- `/v1/lodgings/search`: ★ 가까운 키워드를 사전 전체(7천)에서 뽑으면 상위가 아무 숙소에도 안 붙은
데이터셋 단어로 차서 `선유도 독채` 가 0건이었다 → **숙소에 연결된 키워드만** 비교한다.
★ e5-small 은 지역이 틀려도 0.85 가 나온다(`여수 호텔` → 군산 호텔 0.848) — 지역은 임계값이 아니라
**지역 이름**으로 가른다.
- ★ 로컬에서 임베딩 모델 내려받기가 Node 에서 `ECONNRESET`(terminated)으로 끊겨 3분 걸리고 실패했다.
curl 로는 받힌다. 실패한 로딩 promise 가 프로세스에 남아 **재시작 전까지 모든 임베딩이 500**이다.
- 실키 검증(가짜 숙소 2곳): 메타에 `군산 시내 호텔`·`군산 오션뷰 호텔` 이 실렸고 미확인 `주차`·`애견동반`
은 안 나왔다. 백엔드 `1041 passed / 47 failed` — 47건은 이 변경 전 main 과 같은 목록.
## 2026-09-30 — 템플릿 검수 · 코랄 · 미니멀 · 솔숲 추가 ## 2026-09-30 — 템플릿 검수 · 코랄 · 미니멀 · 솔숲 추가
- 한국 펜션 사이트(코랄트리 · 바다동화)를 참고해 `coral` · `minimal`, 디자인 스킬 시안에서 `pine`(솔숲). - 한국 펜션 사이트(코랄트리 · 바다동화)를 참고해 `coral` · `minimal`, 디자인 스킬 시안에서 `pine`(솔숲).

View File

@ -7,6 +7,8 @@ from config.server_configs import external_api_config
# 로컬 임베딩 모델이라 호출당 1~2초다. # 로컬 임베딩 모델이라 호출당 1~2초다.
TIMEOUT_SEC = 15.0 TIMEOUT_SEC = 15.0
# 처음 보는 업체는 publish 가 OpenAI 키워드 생성을 기다린다(키워드 15개 + 질문 5개).
PUBLISH_TIMEOUT_SEC = 60.0
DEFAULT_LIMIT = 40 DEFAULT_LIMIT = 40
PUBLISH_PATH = "/v1/merchants/publish" PUBLISH_PATH = "/v1/merchants/publish"
@ -25,8 +27,8 @@ def is_configured() -> bool:
return bool(base_url()) return bool(base_url())
async def _post(client: httpx.AsyncClient, path: str, body: dict) -> dict: async def _post(client: httpx.AsyncClient, path: str, body: dict, timeout: float = TIMEOUT_SEC) -> dict:
res = await client.post(f"{base_url()}{path}", json=body) res = await client.post(f"{base_url()}{path}", json=body, timeout=timeout)
if res.status_code >= 400: if res.status_code >= 400:
raise SiteOntologyError(f"{path} HTTP {res.status_code}: {res.text[:200]}") raise SiteOntologyError(f"{path} HTTP {res.status_code}: {res.text[:200]}")
try: try:
@ -40,17 +42,18 @@ async def _post(client: httpx.AsyncClient, path: str, body: dict) -> dict:
async def match_for_merchant(merchant: dict, limit: int = DEFAULT_LIMIT) -> dict: async def match_for_merchant(merchant: dict, limit: int = DEFAULT_LIMIT) -> dict:
"""업체를 저장하고, 그 업체로 해석된 추천 결과(/v1/match 응답)를 돌려준다.""" """업체를 저장하고, 그 업체로 해석된 추천 결과(/v1/match 응답)를 돌려준다."""
body = {**merchant, "generate": False} # sync 는 처음 보는 업체에만 먹는다 — 첫 발행부터 새 키워드가 실리고, 갱신은 저쪽 큐가 한다.
body = {**merchant, "generate": True, "sync": True}
try: try:
async with httpx.AsyncClient(timeout=TIMEOUT_SEC) as client: async with httpx.AsyncClient(timeout=TIMEOUT_SEC) as client:
try: try:
await _post(client, PUBLISH_PATH, body) await _post(client, PUBLISH_PATH, body, PUBLISH_TIMEOUT_SEC)
except SiteOntologyError as ex: except SiteOntologyError as ex:
if not body.get("regionId"): if not body.get("regionId"):
raise raise
LOG.w(f"[site-ontology] regionId={body['regionId']} 로 저장하지 못했다 — 지역 없이 다시 보낸다: {ex}") LOG.w(f"[site-ontology] regionId={body['regionId']} 로 저장하지 못했다 — 지역 없이 다시 보낸다: {ex}")
body = {**body, "regionId": None} body = {**body, "regionId": None}
await _post(client, PUBLISH_PATH, body) await _post(client, PUBLISH_PATH, body, PUBLISH_TIMEOUT_SEC)
result = await _post(client, MATCH_PATH, {"query": merchant["externalId"], "limit": limit}) result = await _post(client, MATCH_PATH, {"query": merchant["externalId"], "limit": limit})
except httpx.HTTPError as ex: except httpx.HTTPError as ex:
raise SiteOntologyError(f"{type(ex).__name__}: {ex}") from ex raise SiteOntologyError(f"{type(ex).__name__}: {ex}") from ex

View File

@ -9,6 +9,9 @@ from services.external import site_ontology
from services.site_payload import _parse_address_parts from services.site_payload import _parse_address_parts
LODGING_INDUSTRY = "stay.pension" LODGING_INDUSTRY = "stay.pension"
HOTEL_INDUSTRY = "stay.hotel"
# 업종별 숙박 유형어. 자기 유형어만 자료 없이 통과한다 — 호텔에 `군산 독채펜션` 이 붙지 않게.
_TYPE_WORDS = {LODGING_INDUSTRY: "펜션", HOTEL_INDUSTRY: "호텔"}
MAX_KEYWORDS = 10 MAX_KEYWORDS = 10
MAX_NEARBY = 8 MAX_NEARBY = 8
@ -57,9 +60,9 @@ _REGION_KEYS = {
} }
# 자료에 없어도 되는 낱말 — 업종어와 "근처" 류. # 자료에 없어도 되는 낱말 — 업종어와 "근처" 류.
_GENERIC_WORDS = frozenset({"펜션", "숙소", "숙박", "스테이", "근처", "가까운", "주변", "인근", "예약", "추천"}) _GENERIC_WORDS = frozenset({"숙소", "숙박", "스테이", "근처", "가까운", "주변", "인근", "예약", "추천"})
# "독채펜션"·"감성숙소" 처럼 붙여 쓴 업종어는 떼고 앞부분만 자료와 대조한다. # "독채펜션"·"감성숙소" 처럼 붙여 쓴 업종어는 떼고 앞부분만 자료와 대조한다.
_GENERIC_SUFFIXES = ("펜션", "숙소", "스테이") _GENERIC_SUFFIXES = ("숙소", "스테이")
# 제목에는 싣지 않는 낱말. # 제목에는 싣지 않는 낱말.
_TITLE_BLOCKED_WORDS = frozenset({"예약", "추천"}) _TITLE_BLOCKED_WORDS = frozenset({"예약", "추천"})
# 한 글자 낱말("봄"·"뷰")은 소개문 어딘가에 우연히 들어 있어 대조가 무의미하다 — 통과시키지 않는다. # 한 글자 낱말("봄"·"뷰")은 소개문 어딘가에 우연히 들어 있어 대조가 무의미하다 — 통과시키지 않는다.
@ -102,6 +105,12 @@ def _locality_word(*addresses: str | None) -> str:
return word[:-1] if len(word) > 2 and word.endswith(("시", "군", "구")) else word return word[:-1] if len(word) > 2 and word.endswith(("시", "군", "구")) else word
def industry_of(place: dict) -> str:
"""외부 분류(없으면 상호)에 '호텔' 이 있으면 호텔, 아니면 펜션."""
text = f"{_text(place.get('external_category'))} {_text(place.get('name'))}"
return HOTEL_INDUSTRY if "호텔" in text else LODGING_INDUSTRY
def build_merchant(place_id: str, snapshot: dict) -> dict | None: def build_merchant(place_id: str, snapshot: dict) -> dict | None:
"""스냅샷 → SiteOntology publish 요청 본문.""" """스냅샷 → SiteOntology publish 요청 본문."""
place = (snapshot or {}).get("place") or {} place = (snapshot or {}).get("place") or {}
@ -157,7 +166,7 @@ def build_merchant(place_id: str, snapshot: dict) -> dict | None:
return { return {
"externalId": place_id, "externalId": place_id,
"name": name, "name": name,
"industryId": LODGING_INDUSTRY, "industryId": industry_of(place),
"regionId": region_key(road_address, address), "regionId": region_key(road_address, address),
"description": description, "description": description,
"profile": {key: value for key, value in profile.items() if value}, "profile": {key: value for key, value in profile.items() if value},
@ -174,13 +183,14 @@ def _evidence(merchant: dict) -> str:
return _compact(" ".join(_text(p) for p in parts if p)) return _compact(" ".join(_text(p) for p in parts if p))
def _supported(keyword: str, evidence: str) -> bool: def _supported(keyword: str, evidence: str, type_word: str = "펜션") -> bool:
"""모든 낱말이 자료에 있고, **자료로 확인한 낱말이 하나는 있어야** 한다.""" """모든 낱말이 자료에 있고, **자료로 확인한 낱말이 하나는 있어야** 한다."""
suffixes = (*_GENERIC_SUFFIXES, type_word)
specific = False specific = False
for word in keyword.split(): for word in keyword.split():
if word in _GENERIC_WORDS: if word in _GENERIC_WORDS or word == type_word:
continue continue
core = next((word[: -len(s)] for s in _GENERIC_SUFFIXES if word.endswith(s) and len(word) > len(s)), word) core = next((word[: -len(s)] for s in suffixes if word.endswith(s) and len(word) > len(s)), word)
core = _compact(core) core = _compact(core)
if len(core) < _MIN_CORE_LEN or core not in evidence: if len(core) < _MIN_CORE_LEN or core not in evidence:
return False return False
@ -193,15 +203,15 @@ def _is_question(item: dict) -> bool:
return item.get("category") == "질문형" or "?" in canonical or canonical.endswith(("요", "까")) return item.get("category") == "질문형" or "?" in canonical or canonical.endswith(("요", "까"))
def _usable(item, evidence: str) -> str | None: def _usable(item, evidence: str, type_word: str) -> str | None:
"""메타 태그에 실어도 되는 추천이면 그 표기를, 아니면 None.""" """메타 태그에 실어도 되는 추천이면 그 표기를, 아니면 None."""
if not isinstance(item, dict) or item.get("status") != "ok" or _is_question(item): if not isinstance(item, dict) or item.get("status") != "ok" or _is_question(item):
return None return None
canonical = _text(item.get("canonical")) canonical = _text(item.get("canonical"))
return canonical if canonical and _supported(canonical, evidence) else None return canonical if canonical and _supported(canonical, evidence, type_word) else None
def _title_keyword(result: dict, evidence: str, locality: str) -> str | None: def _title_keyword(result: dict, evidence: str, locality: str, type_word: str) -> str | None:
"""제목 업종어 자리에 넣을 대표 키워드 — 유형 레인에서 고른다.""" """제목 업종어 자리에 넣을 대표 키워드 — 유형 레인에서 고른다."""
lane = next( lane = next(
(l for l in result.get("byLane") or [] if isinstance(l, dict) and l.get("key") == "type"), (l for l in result.get("byLane") or [] if isinstance(l, dict) and l.get("key") == "type"),
@ -209,7 +219,7 @@ def _title_keyword(result: dict, evidence: str, locality: str) -> str | None:
) )
items = [i for i in lane.get("items") or [] if isinstance(i, dict)] items = [i for i in lane.get("items") or [] if isinstance(i, dict)]
for item in sorted(items, key=lambda i: i.get("category") != "코어"): for item in sorted(items, key=lambda i: i.get("category") != "코어"):
canonical = _usable(item, evidence) canonical = _usable(item, evidence, type_word)
if not canonical or _TITLE_BLOCKED_WORDS.intersection(canonical.split()): if not canonical or _TITLE_BLOCKED_WORDS.intersection(canonical.split()):
continue continue
if locality and locality not in canonical: if locality and locality not in canonical:
@ -221,10 +231,11 @@ def _title_keyword(result: dict, evidence: str, locality: str) -> str | None:
def select(result: dict, merchant: dict) -> dict: def select(result: dict, merchant: dict) -> dict:
"""/v1/match 응답 → {keywords, titleKeyword?}.""" """/v1/match 응답 → {keywords, titleKeyword?}."""
evidence = _evidence(merchant) evidence = _evidence(merchant)
type_word = _TYPE_WORDS.get(merchant.get("industryId"), _TYPE_WORDS[LODGING_INDUSTRY])
keywords: list[str] = [] keywords: list[str] = []
seen: set[str] = set() seen: set[str] = set()
for item in result.get("matches") or []: for item in result.get("matches") or []:
canonical = _usable(item, evidence) canonical = _usable(item, evidence, type_word)
if not canonical or _compact(canonical) in seen: if not canonical or _compact(canonical) in seen:
continue continue
seen.add(_compact(canonical)) seen.add(_compact(canonical))
@ -233,7 +244,7 @@ def select(result: dict, merchant: dict) -> dict:
break break
seo: dict = {"keywords": keywords} seo: dict = {"keywords": keywords}
title = _title_keyword(result, evidence, _locality_word((merchant.get("profile") or {}).get("address"))) title = _title_keyword(result, evidence, _locality_word((merchant.get("profile") or {}).get("address")), type_word)
if title: if title:
seo["titleKeyword"] = title seo["titleKeyword"] = title
return seo return seo

View File

@ -109,6 +109,8 @@ async def build_snapshot(place) -> dict:
"name": place.name, "name": place.name,
"category": category.value, "category": category.value,
"category_name": schema.label, "category_name": schema.label,
# 숙박 안의 유형(호텔·펜션) 판정 근거 — services/seo_keywords.industry_of
"external_category": getattr(place, "external_category", None),
"road_address": place.road_address, "road_address": place.road_address,
"address": place.address, "address": place.address,
"phone": place.phone, "phone": place.phone,

View File

@ -177,8 +177,8 @@ async def test_업체를_저장한_뒤_그_업체로_추천을_받는다(ontolog
await site_ontology.match_for_merchant(merchant) await site_ontology.match_for_merchant(merchant)
assert [url for url, _ in sent] == ["http://onto.test/v1/merchants/publish", "http://onto.test/v1/match"] assert [url for url, _ in sent] == ["http://onto.test/v1/merchants/publish", "http://onto.test/v1/match"]
# generate:false 가 빠지면 SiteOntology 가 LLM 키워드 생성을 큐에 넣는다. # 처음 보는 업체는 생성을 기다려 첫 발행부터 새 키워드가 실린다(sync 는 저쪽이 처음 보는 업체에만 먹인다).
assert sent[0][1] == {**merchant, "generate": False} assert sent[0][1] == {**merchant, "generate": True, "sync": True}
assert sent[1][1] == {"query": "place-1", "limit": site_ontology.DEFAULT_LIMIT} assert sent[1][1] == {"query": "place-1", "limit": site_ontology.DEFAULT_LIMIT}
@ -296,3 +296,32 @@ async def test_SiteOntology_가_죽어도_발행된다(auth_headers, client, db_
assert site["site"]["status"] == SiteStatus.PUBLISHED.value assert site["site"]["status"] == SiteStatus.PUBLISHED.value
payload = json.loads(open(r["payload_path"], encoding="utf-8").read()) payload = json.loads(open(r["payload_path"], encoding="utf-8").read())
assert "seo" not in payload assert "seo" not in payload
# ── 호텔 ───────────────────────────────────────────────────────────────────
def test_외부_분류에_호텔이_있으면_호텔_업종으로_보낸다():
merchant = seo_keywords.build_merchant("place-1", _snapshot(external_category="여행 > 숙박 > 호텔"))
assert merchant["industryId"] == "stay.hotel"
def test_외부_분류가_없으면_상호로_가른다():
assert seo_keywords.build_merchant("place-1", _snapshot(name="군산 베이 호텔"))["industryId"] == "stay.hotel"
assert seo_keywords.build_merchant("place-1", _snapshot())["industryId"] == "stay.pension"
def _item(canonical):
return {"canonical": canonical, "status": "ok", "category": "코어"}
def test_호텔에는_펜션_유형어를_싣지_않는다():
merchant = seo_keywords.build_merchant("place-1", _snapshot(name="군산 베이 호텔"))
result = {"matches": [_item("군산 독채펜션"), _item("군산 독채 호텔"), _item("말랭이마을 숙소")]}
assert seo_keywords.select(result, merchant)["keywords"] == ["군산 독채 호텔", "말랭이마을 숙소"]
def test_펜션에는_호텔_유형어를_싣지_않는다():
merchant = seo_keywords.build_merchant("place-1", _snapshot())
result = {"matches": [_item("군산 독채 호텔"), _item("군산 독채펜션")]}
assert seo_keywords.select(result, merchant)["keywords"] == ["군산 독채펜션"]