작업트리에 커밋되지 않은 채 쌓여 있던 것과, 오늘 찾은 문제 셋을 함께 담는다. ## 1. 콘텐츠 생성 진행 상태 (작업트리에 있던 것) COPY 잡의 실제 단계를 DB에 기록하고 응답으로 내보낸다. 폴링 횟수로 진행률을 흉내 내던 것을 걷어냈다. 새로고침·재접속해도 jobId 로 이어서 본다. - services/copy_steps.py · services/job_progress.py · common/job_errors.py (신규) - postgres-init/migrations/0013_job_progress.sql + init.sql - 프론트: useGenerationJob · generationLabels (신규), Step5Generating·pollJob 배선, orval 모델 갱신(jobProgress · jobStep · jobStepStatus · jobStepReason) - docs/GENERATION_FLOW.md (신규) ## 2. 발행된 사이트만 색인한다 실측(2026-09-15): 디스크의 발행본 33곳 중 **15곳이 draft 인데 `index, follow`** 였고 사이트맵에도 올라가 있었다. 사장님이 발행 버튼을 누른 적 없는 사이트가 짓다 만 상태로 구글에 실려 있었다는 뜻이다. head.ts 가 robots 를 하드코딩하고 payload 의 `site.status` 를 보지 않았다. "색인을 막을 이유가 없다"는 주석은 굽는 것이 곧 발행이던 시절의 말인데, 지금은 빌더 미리보기만 눌러도 draft 로 구워진다. - seo/head.ts: PUBLISHED 일 때만 index, 아니면 `noindex, follow` - 사이트맵·`/s` 목록·llms.txt 에서도 함께 빠진다 — 그쪽은 구운 HTML 의 robots 를 읽어 거른다(seo/directory.ts readBakedNoindex). 규칙을 두 자리에 두지 않으려고 한 곳에 뒀다 ## 3. [새로 크롤링하고 사이트 생성하기] 를 뒤집지 않는다ba90a19의 중복 합치기가 **일부러 다시 만들려는 경우까지** 기존 사업장으로 끌고 갔다 — 새로 만들기를 눌렀는데 기존 에디터가 열린다(사장님 보고 2026-09-15). - Req_VerifyPlaceByUrl.reuse_existing (기본 True — 다른 호출자의 동작은 그대로) - place_service.verify_place_by_url: 끄면 이어붙이지 않는다. 다만 **비어 있는 중복 행은 계속 치운다** — 원래 막으려던 누적이 그것이고 빈 행은 잃을 것이 없다 - ensureServerPlace: 위저드는 새로 만들기 경로에서만 오므로 False 로 보낸다 ## 4. 발행본 파비콘 발행본에 파비콘 링크가 아예 없어 브라우저 탭에 기본 아이콘이 떴다. 파일은 오리진 루트의 공용 자산이라 사이트마다 복사하지 않고 루트 절대경로로 가리킨다. 검증: site vitest 84건 통과 · tsc(site·frontend) · eslint 통과. 백엔드 pytest 는 로컬 DB 비밀번호가 맞지 않아 돌리지 못했다(a5b8701과 같은 자리). 발행본 반영에는 전체 재굽기가 필요하다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
355 lines
18 KiB
Python
355 lines
18 KiB
Python
"""수집 어댑터 계약 — 어디서 긁어오든 결과 모양이 같은지, 그리고 법무 게이트가 코드로 지켜지는지.
|
|
|
|
가장 중요한 것 두 개:
|
|
1. 목데이터의 fact key 가 전부 업종 스키마에 있는가
|
|
— 없는 key 는 FACT_INVALID_KEY 로 전부 거부되므로, 어긋나면 목데이터가 파이프라인을 하나도 검증하지 못한다
|
|
2. Phase 1 에 실제 크롤 어댑터가 등록돼 있지 않은가
|
|
— 법적 검토 결론 전까지 크롤이 돌면 안 된다(docs/DECISIONS.md 1-1)
|
|
|
|
DB 를 쓰지 않는 순수 단위 테스트다.
|
|
"""
|
|
import pytest
|
|
|
|
from common.category_schema import get_schema
|
|
from common.enums import LinkChannel, PlaceCategory
|
|
from services.collector import (
|
|
AdapterDisabled,
|
|
AdapterNotFound,
|
|
AdapterRegistry,
|
|
CollectedFact,
|
|
CollectedMedia,
|
|
CollectError,
|
|
MockAdapter,
|
|
RawSource,
|
|
adapter_ids,
|
|
get_adapter,
|
|
)
|
|
from services.collector.registry import REGISTRY
|
|
|
|
# 어떤 어댑터도 처리하면 안 되는 URL — 조용히 빈 결과를 주지 않고 AdapterNotFound 로 끊어야 한다.
|
|
# ★ 네이버 플레이스는 2026-08-27 부터 naver_place 어댑터가 처리하므로 여기서 뺐다.
|
|
# ★ 일반 도메인(example.co.kr)도 2026-08-28 부터 static_html 이 처리하므로 뺐다.
|
|
# 여기 남은 것은 **수집 불가 결론이 난 곳**이다(docs/DATA_SOURCE_RESEARCH.md):
|
|
# 야놀자·여기어때는 403 + Cloudflare 로 막혀 있고 민사 10억 선례가 있다.
|
|
# 카카오맵은 내부 API 406. 인스타는 Graph API(사장님 OAuth)로만 간다.
|
|
# ★ NOL 전용 어댑터(2026-09-14)가 받는 것은 `nol.yanolja.com/stay/domestic/<id>` 한 패턴뿐이다.
|
|
# 아래 야놀자 주소가 여전히 막혀야 전용 경로를 연 것과 범용 수집을 푼 것이 갈린다.
|
|
_REAL_URLS = (
|
|
"https://www.yanolja.com/pension/1000",
|
|
"https://www.goodchoice.kr/product/detail/1000",
|
|
"https://place.map.kakao.com/26338954",
|
|
"https://www.instagram.com/some_cafe",
|
|
)
|
|
|
|
|
|
# ── 목데이터 ↔ 업종 스키마 정합성 (제일 중요) ────────────────────────────
|
|
@pytest.mark.parametrize("category", list(PlaceCategory))
|
|
async def test_mock_fact_keys_exist_in_category_schema(category):
|
|
"""검증: MockAdapter 가 뱉는 fact key 를 업종 스키마와 대조한다.
|
|
기대결과: 전부 스키마에 존재한다 — 없는 key 는 fact 기록에서 전부 거부돼 목데이터가 무의미해진다."""
|
|
schema = get_schema(category)
|
|
source = await MockAdapter().fetch(MockAdapter.url_for(category))
|
|
|
|
assert source.ok is True
|
|
assert source.facts, f"{schema.name}: 목데이터 fact 가 비었다"
|
|
for fact in source.facts:
|
|
assert schema.has(fact.key), f"{schema.name} 스키마에 없는 key={fact.key}"
|
|
|
|
|
|
@pytest.mark.parametrize("category", list(PlaceCategory))
|
|
async def test_mock_fact_scope_matches_schema(category):
|
|
"""검증: fact 의 scope 가 스키마 정의와 같은지.
|
|
기대결과: place 필드는 place 로, unit 필드는 unit(+단위 이름)으로 수집된다.
|
|
scope 가 어긋나면 객실별 값이 사업장 값으로 뭉개진다."""
|
|
schema = get_schema(category)
|
|
source = await MockAdapter().fetch(MockAdapter.url_for(category))
|
|
|
|
for fact in source.facts:
|
|
assert fact.scope == schema.get(fact.key).scope, f"{schema.name}.{fact.key}: scope 불일치"
|
|
if fact.scope == "unit":
|
|
assert fact.unit_name, f"{schema.name}.{fact.key}: unit 스코프인데 단위 이름이 없다"
|
|
|
|
|
|
@pytest.mark.parametrize("category", list(PlaceCategory))
|
|
async def test_mock_does_not_collect_llm_written_fields(category):
|
|
"""검증: 수집물에 소개문 계열(allow_llm=True) 필드가 섞이는지.
|
|
기대결과: 없다 — 소개문은 수집하는 사실이 아니라 generator 가 쓰는 문장이다(절대규칙 7)."""
|
|
schema = get_schema(category)
|
|
source = await MockAdapter().fetch(MockAdapter.url_for(category))
|
|
|
|
llm_keys = set(schema.llm_writable_keys())
|
|
collected = {f.key for f in source.facts}
|
|
assert not (collected & llm_keys), f"{schema.name}: 수집물에 LLM 작성 필드가 섞였다 — {collected & llm_keys}"
|
|
|
|
|
|
@pytest.mark.parametrize("category", list(PlaceCategory))
|
|
async def test_mock_returns_three_to_five_photos(category):
|
|
"""검증: 사진 수집 결과.
|
|
기대결과: 3~5장. Gemini Vision 배치 처리를 물려볼 수 있는 최소량이다."""
|
|
source = await MockAdapter().fetch(MockAdapter.url_for(category))
|
|
assert 3 <= len(source.media) <= 5, f"사진 {len(source.media)}장"
|
|
for item in source.media:
|
|
assert item.origin_url.startswith("http")
|
|
|
|
|
|
async def test_mock_is_deterministic():
|
|
"""검증: 같은 URL 로 두 번 수집한다.
|
|
기대결과: fact·사진이 동일하다 — 목데이터가 흔들리면 파이프라인 테스트가 못 믿을 게 된다."""
|
|
adapter = MockAdapter()
|
|
url = MockAdapter.url_for(PlaceCategory.LODGING)
|
|
first, second = await adapter.fetch(url), await adapter.fetch(url)
|
|
|
|
assert [(f.key, f.value, f.unit_name) for f in first.facts] == [(f.key, f.value, f.unit_name) for f in second.facts]
|
|
assert [m.origin_url for m in first.media] == [m.origin_url for m in second.media]
|
|
|
|
|
|
async def test_mock_reads_category_and_channel_from_url():
|
|
"""검증: URL 에서 업종·채널을 읽는지.
|
|
기대결과: 업종별로 다른 목데이터가 나오고, ?channel= 로 채널이 지정된다."""
|
|
adapter = MockAdapter()
|
|
lodging = await adapter.fetch("mock://lodging/p1")
|
|
cafe = await adapter.fetch("mock://cafe/c1?channel=naver_place")
|
|
|
|
assert "check_in_time" in lodging.fact_map()
|
|
assert "break_time" in cafe.fact_map()
|
|
assert lodging.channel is LinkChannel.ETC
|
|
assert cafe.channel is LinkChannel.NAVER_PLACE
|
|
|
|
|
|
async def test_mock_unknown_category_fails_softly():
|
|
"""검증: 업종을 못 읽는 mock URL.
|
|
기대결과: 예외가 아니라 ok=False 결과 — 채널 하나가 이상해도 나머지 수집이 멈추면 안 된다."""
|
|
source = await MockAdapter().fetch("mock://unknown/x1")
|
|
assert source.ok is False
|
|
assert source.error and "업종" in source.error
|
|
assert source.facts == []
|
|
|
|
|
|
# ── can_handle / 레지스트리 ──────────────────────────────────────────────
|
|
def test_can_handle_accepts_mock_urls_only():
|
|
"""검증: MockAdapter 의 처리 범위.
|
|
기대결과: mock:// 와 mock.test 만 받고 실제 사이트 URL 은 거부한다."""
|
|
adapter = MockAdapter()
|
|
assert adapter.can_handle("mock://lodging/1") is True
|
|
assert adapter.can_handle("https://mock.test/cafe/1") is True
|
|
for url in _REAL_URLS:
|
|
assert adapter.can_handle(url) is False, f"MockAdapter 가 실제 URL 을 처리하려 한다: {url}"
|
|
assert adapter.can_handle("") is False
|
|
|
|
|
|
def test_registry_resolves_mock_url():
|
|
"""검증: 레지스트리로 어댑터를 찾는다.
|
|
기대결과: mock URL 은 MockAdapter 로 해결된다."""
|
|
assert get_adapter("mock://lodging/1").id == "mock"
|
|
|
|
|
|
@pytest.mark.parametrize("url", _REAL_URLS)
|
|
def test_registry_raises_for_unhandled_url(url):
|
|
"""검증: 등록된 어댑터가 처리 못 하는 URL 을 조회한다.
|
|
기대결과: AdapterNotFound — 조용히 빈 결과를 돌려주지 않는다."""
|
|
with pytest.raises(AdapterNotFound):
|
|
get_adapter(url)
|
|
|
|
|
|
def test_registry_rejects_duplicate_adapter_id():
|
|
"""검증: 같은 id 의 어댑터를 두 번 등록한다.
|
|
기대결과: ValueError — 어느 쪽이 쓰이는지 모르는 상태를 만들지 않는다."""
|
|
registry = AdapterRegistry()
|
|
registry.register(MockAdapter())
|
|
with pytest.raises(ValueError):
|
|
registry.register(MockAdapter())
|
|
|
|
|
|
def test_disabled_adapter_is_blocked():
|
|
"""검증: 등록은 됐지만 ENABLED 목록에 없는 어댑터를 조회한다.
|
|
기대결과: AdapterDisabled — 등록과 사용 허가를 분리해 실수로 켜지지 않게 한다."""
|
|
registry = AdapterRegistry(enabled=frozenset())
|
|
registry.register(MockAdapter())
|
|
with pytest.raises(AdapterDisabled):
|
|
registry.get_adapter("mock://lodging/1")
|
|
assert registry.can_handle("mock://lodging/1") is False
|
|
|
|
|
|
# ── 법무 게이트 (Phase 1) ────────────────────────────────────────────────
|
|
def test_registers_only_reviewed_adapters():
|
|
"""검증: 기본 레지스트리에 등록된 어댑터 목록.
|
|
|
|
기대결과: 검토를 거쳐 명시적으로 승인한 것만 있다. 이 목록이 늘어나는 것은
|
|
**의도된 결정이어야** 하므로, 코드가 몰래 늘어나면 여기서 깨진다."""
|
|
assert adapter_ids() == ["naver_place", "tour_api", "yanolja", "mock", "static_html"], (
|
|
f"예상 밖 어댑터가 등록됐다: {adapter_ids()} — 승인 없이 수집 대상을 늘리지 않는다"
|
|
)
|
|
|
|
|
|
def test_static_html_is_registered_last():
|
|
"""검증: 넓은 어댑터(static_html)가 좁은 어댑터보다 뒤에 있는가.
|
|
|
|
기대결과: 목록의 맨 끝. 앞에 있으면 http(s) 를 통째로 받는 static_html 이
|
|
네이버 플레이스 URL 까지 가로채 naver_place 가 영영 안 불린다."""
|
|
assert adapter_ids()[-1] == "static_html", (
|
|
f"static_html 은 맨 뒤여야 한다: {adapter_ids()}"
|
|
)
|
|
|
|
|
|
def test_static_html_never_touches_blocked_platforms():
|
|
"""검증: 수집 불가 결론이 난 플랫폼 URL 을 static_html 에 직접 물어본다.
|
|
|
|
기대결과: 전부 False. static_html 의 정당성은 '사장님이 확정한 자기 홈페이지만 본다' 에서
|
|
나오므로, 플랫폼 URL 이 흘러들어오면 정당성이 통째로 깨진다. 운영자가 실수로 넣어도
|
|
구조적으로 막혀야 한다(docs/DATA_SOURCE_RESEARCH.md)."""
|
|
from services.collector.static_html_adapter import StaticHtmlAdapter
|
|
|
|
adapter = StaticHtmlAdapter()
|
|
for url in (
|
|
"https://www.yanolja.com/pension/1000", # 403 + 민사 10억 선례
|
|
"https://nol.yanolja.com/hotels/1",
|
|
"https://www.goodchoice.kr/product/detail?ano=1",
|
|
"https://m.place.naver.com/restaurant/1/home", # robots Disallow: / · 전용 어댑터 있음
|
|
"https://blog.naver.com/somepension",
|
|
"https://place.map.kakao.com/26338954", # 내부 API 406
|
|
"https://www.instagram.com/some_cafe", # Graph API(OAuth)로만
|
|
"https://app.catchtable.co.kr/ct/shop/x",
|
|
):
|
|
assert adapter.can_handle(url) is False, f"건드리면 안 되는 URL 을 받았다: {url}"
|
|
|
|
|
|
def test_static_html_takes_owner_domains():
|
|
"""검증: 사장님 자체 홈페이지로 보이는 평범한 도메인.
|
|
|
|
기대결과: static_html 이 받는다. 숙박 표본의 76%가 자체 도메인을 쓰므로
|
|
여기서 놓치면 숙박 수집이 통째로 비어버린다."""
|
|
for url in ("https://gangmunstay.kr/rooms", "https://offinghouse.com/", "https://www.mulhoe.co.kr/"):
|
|
assert get_adapter(url).id == "static_html", f"static_html 이 받지 않는다: {url}"
|
|
|
|
|
|
def test_adapters_do_not_collect_llm_written_fields():
|
|
"""검증: **등록된 모든 어댑터**가 만드는 fact key 에 allow_llm=True 필드가 섞이는지.
|
|
|
|
기대결과: 없다. 소개문 계열은 수집하는 사실이 아니라 generator 가 쓰는 문장이다(절대규칙 7).
|
|
|
|
★ 실측 사고 (2026-08-31) — 이 테스트가 MockAdapter 만 보고 있어서 놓쳤다.
|
|
tour_api·naver_place 가 `intro` 를 수집했고, TourAPI overview 457자가 그대로
|
|
VERIFIED 로 들어가 사장님 사이트의 '숙소 소개' 를 차지했다. Gemini 가 쓴 162자
|
|
소개문은 PENDING_OWNER 로 뒤에 밀려 영영 안 보였다.
|
|
원문 그대로 싣는 것은 GEO 에서도 복제 콘텐츠라 손해다.
|
|
|
|
그래서 소스를 훑어 확인한다 — 어댑터가 늘어나도 같은 실수를 반복할 수 없게."""
|
|
from pathlib import Path
|
|
|
|
banned = set()
|
|
for category in PlaceCategory:
|
|
banned |= set(get_schema(category).llm_writable_keys())
|
|
assert banned, "allow_llm 필드가 하나도 없다 — 스키마 로딩이 잘못됐다"
|
|
|
|
package = Path(__file__).resolve().parents[1] / "services" / "collector"
|
|
for path in package.glob("*_adapter.py"):
|
|
body = path.read_text(encoding="utf-8")
|
|
for key in banned:
|
|
for pattern in (f'key="{key}"', f'("{key}",'):
|
|
for line in body.splitlines():
|
|
stripped = line.strip()
|
|
if pattern in stripped and not stripped.startswith("#"):
|
|
raise AssertionError(
|
|
f"{path.name}: LLM 작성 필드 '{key}' 를 수집한다 — {stripped[:80]}"
|
|
)
|
|
|
|
|
|
@pytest.mark.parametrize("url", _REAL_URLS)
|
|
def test_unregistered_sites_are_not_reachable(url):
|
|
"""검증: 어댑터가 없는 사이트 URL 로 수집을 시도한다.
|
|
기대결과: can_handle 이 False — 확정 링크 거르기 단계에서 크롤링 대상에 오르지 않는다."""
|
|
assert REGISTRY.can_handle(url) is False
|
|
|
|
|
|
def test_naver_place_hosts_are_handled():
|
|
"""검증: 사람이 실제로 공유하는 네이버 플레이스 주소들.
|
|
|
|
기대결과: 전부 naver_place 가 받는다. 하나라도 빠지면 사장님이 붙여넣은 주소가
|
|
'어댑터없음' 으로 조용히 버려진다 — 실제로 map.naver.com 이 빠져서 그랬다."""
|
|
for url in (
|
|
"https://map.naver.com/p/entry/place/1133638931",
|
|
"https://m.place.naver.com/accommodation/1133638931/home",
|
|
"https://pcmap.place.naver.com/accommodation/1133638931/home",
|
|
"https://naver.me/xxxxxxx",
|
|
):
|
|
assert REGISTRY.can_handle(url) is True, f"처리하지 못한다: {url}"
|
|
|
|
|
|
def test_no_evasion_code_in_collector_package():
|
|
"""검증: 수집 패키지에 우회 코드가 들어왔는지 소스를 훑는다.
|
|
기대결과: 캡차 우회·봇 탐지 우회·IP 회전 흔적이 없다(영구 금지, 검토 결과와 무관)."""
|
|
from pathlib import Path
|
|
|
|
banned = ("captcha", "solve_captcha", "rotate_ip", "proxy_rotate", "stealth", "undetected")
|
|
package = Path(__file__).resolve().parents[1] / "services" / "collector"
|
|
for path in package.glob("*.py"):
|
|
body = path.read_text(encoding="utf-8").lower()
|
|
for token in banned:
|
|
# 금지 사실을 적어둔 주석은 허용한다 — 실제 식별자로 쓰였는지만 본다.
|
|
for line in body.splitlines():
|
|
stripped = line.strip()
|
|
if token in stripped and not stripped.startswith("#"):
|
|
raise AssertionError(f"{path.name}: 금지된 우회 코드 흔적 '{token}' — {stripped[:80]}")
|
|
|
|
|
|
# ── 출처 강제 ────────────────────────────────────────────────────────────
|
|
async def test_every_fact_and_media_carries_source_url():
|
|
"""검증: 수집 결과의 모든 fact·사진이 출처를 들고 있는지.
|
|
기대결과: 전부 RawSource.url 과 같다 — 출처 없는 값은 fact 기록에서 거부된다."""
|
|
url = MockAdapter.url_for(PlaceCategory.LODGING, "pension-9")
|
|
source = await MockAdapter().fetch(url)
|
|
|
|
assert source.source_url == url
|
|
for item in (*source.facts, *source.media):
|
|
assert item.source_url == url, f"출처 없이 흘러가는 항목: {item}"
|
|
|
|
|
|
def test_raw_source_stamps_source_url_on_bare_items():
|
|
"""검증: source_url 을 비운 채 항목을 넣어 RawSource 를 만든다.
|
|
기대결과: 생성 시점에 출처가 찍힌다 — 어댑터가 깜빡해도 구조가 막는다."""
|
|
source = RawSource(
|
|
url="mock://lodging/x",
|
|
adapter_id="mock",
|
|
facts=[CollectedFact(key="wifi", value="true")],
|
|
media=[CollectedMedia(origin_url="https://mock.test/img/1.jpg")],
|
|
)
|
|
assert source.facts[0].source_url == "mock://lodging/x"
|
|
assert source.media[0].source_url == "mock://lodging/x"
|
|
|
|
|
|
def test_raw_source_requires_url():
|
|
"""검증: url 없이 RawSource 를 만든다.
|
|
기대결과: CollectError — 출처 없는 수집 결과는 존재할 수 없다."""
|
|
with pytest.raises(CollectError):
|
|
RawSource(url="", adapter_id="mock")
|
|
|
|
|
|
def test_collected_media_requires_origin_url():
|
|
"""검증: origin_url 없이 사진을 담는다.
|
|
기대결과: CollectError — 재게시 권리 판단(docs/DECISIONS.md 1-2)에 원본 URL 이 필요하다."""
|
|
with pytest.raises(CollectError):
|
|
CollectedMedia(origin_url=" ")
|
|
|
|
|
|
def test_unit_scoped_fact_requires_unit_name():
|
|
"""검증: 단위 이름 없이 unit 스코프 fact 를 만든다.
|
|
기대결과: CollectError — 어느 객실 값인지 모르면 저장할 수 없다."""
|
|
with pytest.raises(CollectError):
|
|
CollectedFact(key="standard_capacity", value="2", scope="unit")
|
|
|
|
|
|
async def test_unit_names_are_listed_in_order():
|
|
"""검증: 수집된 단위 이름 목록.
|
|
기대결과: 등록 순서대로 중복 없이 나온다 — place.units 시드에 그대로 쓴다."""
|
|
source = await MockAdapter().fetch(MockAdapter.url_for(PlaceCategory.LODGING))
|
|
assert source.unit_names() == ["A동 스탠다드", "B동 복층"]
|
|
|
|
|
|
async def test_failure_result_has_no_facts():
|
|
"""검증: 실패 결과(RawSource.failure)의 모양.
|
|
기대결과: ok=False, error 존재, fact·사진 없음 — 실패를 성공으로 오인할 수 없다."""
|
|
source = RawSource.failure("mock://lodging/1", "mock", "타임아웃")
|
|
assert source.ok is False and source.error == "타임아웃"
|
|
assert source.facts == [] and source.media == []
|
|
assert source.source_url == "mock://lodging/1"
|