o2o-infinith-demo/data/clinic-registry/extract_place_ids.py
Haewon Kam d5f7f24e0a feat: clinic registry DB + pipeline audit P0 fixes
## Clinic Registry
- data/clinic-registry/clinic_registry_working.csv — 91개 병원 채널 마스터 DB
- data/clinic-registry/INFINITH_Outbound_List.csv — BD팀 아웃바운드 리스트 (17컬럼)
- data/clinic-registry/update_csv.py — 안전 CSV 업데이트 스크립트 (빈 필드만 채움)
- data/clinic-registry/extract_place_ids.py — 네이버 플레이스 ID 추출기
- scripts/import-registry.ts — CSV → Supabase clinic_registry 테이블 임포트
- supabase/migrations/20260406_clinic_registry.sql — clinic_registry 테이블 스키마

## Pipeline P0 Bug Fixes (전수 감사 후)
- fix(collect-channel-data): 강남언니 rating 0-10 스케일 오변환 제거
  - 기존: rating ≤ 5이면 ×2 → 4.8/10을 9.6/10으로 잘못 변환
  - 수정: Firecrawl 프롬프트가 이미 0-10 지시 → rawValue 직접 신뢰
- fix(generate-report): Perplexity 단일 fetch → fetchWithRetry 교체
  - maxRetries:2, backoffMs:[5000,15000], timeoutMs:90s
  - 기존: 타임아웃/429 시 리포트 생성 전체 실패
  - 수정: 자동 재시도로 일시적 API 오류 극복

## Docs
- docs/PIPELINE_IMPROVEMENT_PLAN.md — Sprint 0/1/2 완료 표시 + 전수 감사 결과 추가
- docs/REGISTRY_FUNCTIONAL_SPECS.md, DB_SCHEMA_V3.md 외 기획 문서 다수 추가

## New Components & Features
- supabase/functions/generate-content-plan, adjust-strategy — 콘텐츠 플랜/전략 조정
- src/components/plan/EditEntryModal, StrategyAdjustmentSection — 플랜 편집 UI
- supabase/functions/_shared/dataQuality, foundingYearExtractor, urlClassifier — 데이터 품질 유틸

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-07 09:33:25 +09:00

38 lines
1.2 KiB
Python

#!/usr/bin/env python3
"""Extract Naver Place IDs from links arrays"""
import re
import json
import sys
def extract_place_id(links):
"""Extract first valid Naver Place ID from list of URLs"""
place_ids = set()
for link in links:
# Pattern: place/DIGITS in map.naver.com URLs
# But NOT in search URLs or directions URLs with coordinates
if 'map.naver.com' in link:
matches = re.findall(r'place/(\d{7,12})', link)
for m in matches:
# Filter out coordinate-like numbers (14140xxx pattern)
if not m.startswith('1414'):
place_ids.add(m)
if place_ids:
# Return the most common ID (first one found in entry/place URLs)
for link in links:
if 'entry/place/' in link:
match = re.search(r'entry/place/(\d{7,12})', link)
if match and not match.group(1).startswith('1414'):
return match.group(1)
# Fallback: return smallest ID (usually the main one)
return min(place_ids, key=len)
return None
if __name__ == '__main__':
data = json.load(sys.stdin)
pid = extract_place_id(data)
if pid:
print(pid)
else:
print("NOT_FOUND")