백엔드

네이버 플레이스 크롤러 패턴

백엔드

네이버 플레이스 크롤러 패턴 — 실전 적용 구조와 코드 예시

언제 쓰나 · 1. **CDP + 로그인 세션** — Playwright의 `launch()` 대신 `connect_over_cdp()`로 평소 쓰던(로그인된) Chrome에 붙기. 이것만으로 차단 대부분 회피. · 2. **헤드리스 금지** — 반드시 GUI Chrome에 사람이 직접 로그인한 프로필 · 3. **데이터센터 IP 금지** — AWS/OCI/GCP IP는 즉시 블랙. 로컬 또는 레지덴셜 프록시 · 4. **블록 감지 → 즉시 중단** — "서비스 이용이 제한", "자동화된 접근" 등 시그널 만나면 6~12h 대기

#backend#네이버#플레이스#크롤러
다운로드
---
title: 네이버 플레이스 크롤러
tags: [backend, crawler, playwright, anti-bot, naver, python]
created: 2026-04-20
---

# 네이버 플레이스 크롤러 패턴

네이버 플레이스에서 장소별 정보(층수/호수/이미지 등)를 수집하는 재사용 가능 패턴.
봇탐지 회피 전략 9종 기본 + 누적 볼륨 제어 3종 고급 (총 12종) 적용.

## 핵심 포인트 — 가장 중요한 것부터

1. **CDP + 로그인 세션** — Playwright의 `launch()` 대신 `connect_over_cdp()`로 평소 쓰던(로그인된) Chrome에 붙기. 이것만으로 차단 대부분 회피.
2. **헤드리스 금지** — 반드시 GUI Chrome에 사람이 직접 로그인한 프로필
3. **데이터센터 IP 금지** — AWS/OCI/GCP IP는 즉시 블랙. 로컬 또는 레지덴셜 프록시
4. **블록 감지 → 즉시 중단** — "서비스 이용이 제한", "자동화된 접근" 등 시그널 만나면 6~12h 대기

## 적용된 봇 회피 전략

### 기본 9종

| 전략 | 구현 |
|---|---|
| CDP 연결 | `p.chromium.connect_over_cdp(f"http://localhost:{CDP_PORT}")` |
| 세션 웜업 | map.naver.com 먼저 방문 + 3~10초 대기 |
| 지수분포 지연 | `random.expovariate(1/mean)` + floor/cap |
| Long Break | 30~50건마다 20~60초 |
| 인간 행동 시뮬 | `human_move`(워밍업) + **`human_browse`(레코드간 50%, 끊어 스크롤 2~4스텝)** |
| 데코이 검색 | 5% 확률로 무관한 검색어 |
| 블록 감지 | `detect_block(page)` — 시그널 리스트 스캔 |
| 활동시간 제한 | 9~24시만 작동 (`--force`로 무시 가능) |
| 순서 셔플 | `random.shuffle(records)` — ID 순차 접근 패턴 회피 |

### 누적 볼륨 제어 3종 (장기·대량 크롤 전용 — 추가)

| 전략 | 구현 | 발동 기준 |
|---|---|---|
| **Long Break 상향** | `long_break(60, 180)` — 30~50건마다 1~3분 | 하루 2,000건+ 크롤 |
| **하루 상한 소프트캡** | `_daily_count` 카운터 → 400건 도달 시 자동 종료 | 장기 크롤 프로젝트 |
| **프로필 로테이션** | 프로필 A/B/C + 포트 9222/9223/9224 | 10,000건+ 정기 크롤 |

> 기본 9종은 모두 적용해도 **같은 IP·프로필로 하루 485건 누적 시 차단됐음** (2026-04-20 경험칙).
> 장기 크롤에는 위 3종을 반드시 추가 적용.

## 층/호 파서 지원 패턴

| 입력 | 결과 |
|---|---|
| `5층` | `floors:["5층"]` |
| `5층 501호` | `floors:["5층"], unit:"501호"` |
| `1~3층` / `1층~3층` | `floors:["1층","2층","3층"]` |
| `1, 2, 3층` / `1·2층` | `floors:["1층","2층","3층"]` (나열) |
| `지하1층 101호` / `B1` | `floors:["지1층"], unit:"101호"` |
| `전층` | `floors:["전층"]` (건축물대장 조회 시 전체 지상층 자동 합산) |

## 실전 속도 (참고)

| 모드 | 분당 | 시간당 | 하루 상한 |
|---|---|---|---|
| 기본 (9종) | 15~20건 | 400~600건 | ~1,000건 |
| **방어 모드** (9종 + 볼륨 제어 3종) | **3~5건** | **200~300건** | **400건 (소프트캡)** |

> **경험칙 (2026-04-20)**: 같은 IP·프로필로 하루 485건 누적 처리 시 차단.
> 장기·대량 크롤에서는 방어 모드 권장.

## 재사용 시 변경 포인트

- `parse_floor_info()` : 층수 대신 다른 정보 추출 시 변경
- `extract_*_from_page()` : DOM 추출 로직 교체
- `INPUT_FILE` / `OUTPUT_FILE` : 입출력 파일
- 크롤링 URL : `m.place.naver.com` (모바일) vs `map.naver.com/p/entry/place` (데스크톱)

## 환경 요구사항

- Python 3.9+
- playwright (`pip install playwright`)
- macOS Chrome (또는 경로 조정)

## ⚠️ 주의

- 헤드리스 서버(OCI/AWS 등)에서는 효과 없음 — 로컬 PC + 로그인 Chrome 조합 필수
- IP 차단되면 데이터센터 IP는 오래 유지되므로 복구 어려움
- **API 키 등 민감 정보는 .env 로 분리하고 코드에 하드코딩 금지**

## 파생 프로젝트

- SmartPicks / 260225_MugJJang — 대형카페 1,580건 층수·면적 수집 (2026-04)
- SmartPicks / space_rental — 공간대여 2,417건 detail 수집 + 외부/내부 사진 수집 (2026-04-27)
  - false positive "잠시 후 다시" 트리거 발견 (record_id 1450061243 = 아임키즈룸 안내문)
  - 카테고리 탭 사진 추출에 JS-click + set-diff 패턴 적용
  - `human_browse(page)` 레코드 간 스크롤 모션 추가

---

## 상세 가이드

> Playwright CDP 방식 + 봇 탐지 회피 전략 (실전 적용 기준)

### 1. 전제 조건 및 환경 설정

#### 필수 패키지

```bash
pip install playwright
playwright install chromium
```

#### Chrome 프로필 설정 (가장 중요)

네이버는 로그인된 Chrome 세션을 신뢰합니다. **전용 프로필을 만들고, 네이버에 로그인한 상태를 유지**해야 합니다.

```
# 전용 프로필 경로 (예시, PC마다 다르게 설정)
~/chrome-naver-auto
```

#### Chrome 디버그 모드로 실행

```bash
# Windows (CMD/PowerShell)
"C:/Program Files/Google/Chrome/Application/chrome.exe" ^
  --user-data-dir="C:/Users/{유저명}/chrome-naver-auto" ^
  --remote-debugging-port=9222

# 주의: --user-data-dir 없으면 Chrome이 CDP 포트를 거부함
# 기본 프로필(AppData/Local/Google/Chrome/User Data)은 사용 불가
```

#### 최초 1회 설정

1. 위 명령으로 Chrome 실행
2. 네이버(naver.com)에서 로그인
3. 네이버 지도(map.naver.com) 방문하여 정상 작동 확인
4. Chrome을 그대로 두고 크롤러 실행

### 2. Chrome CDP 연결 방식

#### 절대 금칙

```python
# ❌ 절대 금지 — 새 브라우저 인스턴스 = 로그인 없음 = 탐지됨
browser = p.chromium.launch()
```

#### 올바른 방식 (CDP 연결)

```python
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    # CDP로 기존 Chrome에 붙기
    browser = p.chromium.connect_over_cdp("http://localhost:9222")

    # 기존 컨텍스트(로그인 세션) 사용
    ctx  = browser.contexts[0]
    page = ctx.new_page()

    # 팝업/다이얼로그 자동 닫기
    page.on("dialog", lambda d: d.dismiss())

    # 작업...

    page.close()
```

### 3. 데이터 추출 전략

네이버 플레이스에서 데이터를 추출하는 방법은 **2가지 URL 패턴**.

#### 방법 A: place_id를 알고 있을 때 (빠름, 추천)

```
URL: https://map.naver.com/p/entry/place/{place_id}
```

```python
def get_data_by_place_id(page, place_id: str) -> dict:
    url = f"https://map.naver.com/p/entry/place/{place_id}"
    page.goto(url, wait_until="domcontentloaded", timeout=30000)
    time.sleep(random.uniform(1.5, 3.0))

    result = {}

    # 1순위: og:image (서버사이드 렌더링, 가장 안정적)
    og_image = page.evaluate(
        "document.querySelector('meta[property=\"og:image\"]')?.getAttribute('content')"
    )
    if og_image and "pstatic" in og_image:
        result["image_url"] = og_image

    # 2순위: Apollo State (SPA 데이터 저장소)
    try:
        apollo = page.evaluate("window.__APOLLO_STATE__ || {}") or {}
        for val in apollo.values():
            if isinstance(val, dict):
                img = val.get("imageUrl") or val.get("mainPhotoUrl")
                name = val.get("name")
                if img and "naver" in img:
                    result["image_url"] = img
                if name:
                    result["name"] = name
    except Exception:
        pass

    # 3순위: DOM에서 직접 추출
    if "image_url" not in result:
        img = page.evaluate("""
            (() => {
                const imgs = [...document.querySelectorAll('img[src]')];
                const cdn = imgs.find(i =>
                    i.src.includes('pstatic.net') || i.src.includes('ldb-phinf')
                );
                return cdn ? cdn.src : null;
            })()
        """)
        if img:
            result["image_url"] = img

    return result
```

#### 방법 B: 이름+주소로 검색 (place_id 모를 때)

```
URL: https://pcmap.place.naver.com/place/list?query={검색어}
```

```python
def search_place(page, name: str, addr: str) -> dict:
    import re
    m = re.search(r"(\S+(?:시|도))\s+(\S+(?:시|군|구))", addr or "")
    region_hint = f"{m.group(1)} {m.group(2)}" if m else ""

    query = f"{name} {region_hint}".strip()
    url = f"https://pcmap.place.naver.com/place/list?query={urllib.parse.quote(query)}"

    page.goto(url, wait_until="load", timeout=25000)
    time.sleep(random.uniform(1.0, 2.5))

    try:
        apollo = page.evaluate("window.__APOLLO_STATE__") or {}
    except Exception:
        return {}

    for val in apollo.values():
        if isinstance(val, dict):
            cname = val.get("name", "")
            cid   = str(val.get("id", ""))
            if cid.isdigit() and len(cid) >= 7:
                if normalize(cname) == normalize(name):
                    return {
                        "place_id":  cid,
                        "name":      cname,
                        "image_url": val.get("imageUrl"),
                        "category":  val.get("category"),
                        "map_url":   f"https://map.naver.com/p/entry/place/{cid}",
                    }
    return {}

def normalize(s: str) -> str:
    """공백/괄호 제거 후 비교"""
    import re
    return re.sub(r"[\s\(\)(\)]", "", s or "")
```

### 4. 봇 탐지 회피 전략 상세

#### 4-1. 세션 웜업 (시작 전 필수)

```python
page.goto("https://map.naver.com/", wait_until="load", timeout=20000)
time.sleep(random.uniform(3.0, 6.0))
human_move(page)
time.sleep(random.uniform(1.5, 3.0))
```

#### 4-2. 지수분포 자연 지연

```python
def natural_wait(mean: float = 3.0, cap: float = 8.0, floor: float = 1.0):
    """
    지수분포로 대기시간 생성 — 실제 사람처럼 불규칙한 간격
    - mean: 평균 대기시간(초) — 권장 3.0
    - cap : 최대 대기시간 — 권장 8.0 (15초 outlier 제거)
    - floor: 최소 대기시간 — 권장 1.0 (0.8초는 봇 의심)
    """
    t = random.expovariate(1.0 / mean)
    t = max(floor, min(t, cap))
    time.sleep(t)
```

> **권장값 변경 이력 (2026-04-27)**: 기존 `(mean=2.5, floor=0.8, cap=15)` →
> `(mean=3.0, floor=1.0, cap=8.0)` 로 상향. 이유:
> - **floor 1초**: 0.8초는 인간 평균 페이지 전환보다 빠름. 봇 의심 신호.
> - **cap 8초**: 15초 outlier 는 long-tail 인간 행동 시뮬엔 좋지만 실제 사용성에서 너무 큼.
> - **mean 3초**: 평균 +0.5초로 rate-limit 마진 ↑. 분당 처리량은 ~16% 감소.
> - **uniform vs exponential**: uniform(1~8s) 도 안전하지만 stddev 작음(~2초) → 패턴 균일.
>   exponential 분포는 stddev=mean(3초) 으로 더 자연스러움. **exp 권장**.

#### 4-3. 인간 행동 시뮬레이션

```python
def human_move(page):
    """마우스 이동 + 스크롤로 사람처럼 행동 (단발성, 워밍업용)"""
    try:
        page.mouse.move(random.randint(200, 900), random.randint(200, 600))
        time.sleep(random.uniform(0.15, 0.5))
        page.evaluate(f"window.scrollBy(0, {random.randint(80, 400)})")
        time.sleep(random.uniform(0.2, 0.7))
        if random.random() < 0.3:
            page.evaluate(f"window.scrollBy(0, {-random.randint(40, 200)})")
            time.sleep(random.uniform(0.1, 0.4))
    except Exception:
        pass


def human_browse(page):
    """레코드 처리 후 페이지를 훑어보는 모션 — 끊어서 스크롤 다운, 가끔 위로.

    **레코드 사이마다** 50% 확률로 호출. bot-detection 회피용.
    `human_move` 와 차이점: smooth scroll + 2~4 multi-step (한 번에 X).
    평균 1~2초 소요.
    """
    if random.random() >= 0.5:
        return
    try:
        # 마우스 살짝 이동
        page.mouse.move(random.randint(150, 1000), random.randint(150, 700))
        time.sleep(random.uniform(0.1, 0.3))
        # 2~4번 끊어서 스크롤 다운 (한 번에 X)
        steps = random.randint(2, 4)
        for _ in range(steps):
            delta = random.randint(120, 400)
            page.evaluate(
                f"window.scrollBy({{top: {delta}, left: 0, behavior: 'smooth'}})"
            )
            time.sleep(random.uniform(0.25, 0.7))
        # 30% 확률로 살짝 위로 (다시 보는 척)
        if random.random() < 0.3:
            page.evaluate(
                f"window.scrollBy({{top: -{random.randint(60, 220)}, left: 0, behavior: 'smooth'}})"
            )
            time.sleep(random.uniform(0.15, 0.4))
    except Exception:
        pass


# 메인 루프에서 사용
for i, record in enumerate(records, 1):
    # ... 처리 ...
    natural_wait()
    human_browse(page)   # 50% 확률로 스크롤 모션
    if i % long_break_every == 0:
        long_break(60, 180)
```

> **2026-04-27 추가**: `human_move` 는 워밍업 1회용. 레코드 사이 자연스러운
> 스크롤은 `human_browse` 로 분리. smooth + 끊어진 multi-step 으로 단발 scrollBy 보다
> 더 인간답게 보임.

#### 4-4. 긴 휴식 (Long Break) — 가장 중요

```python
def long_break(min_s: float = 30, max_s: float = 120):
    """30~120초 긴 휴식 — 30~50건마다 한 번"""
    t = random.uniform(min_s, max_s)
    print(f"  [☕ long break {t:.0f}s]")
    time.sleep(t)

# 사용 예시
long_break_every = random.randint(30, 50)
for i, record in enumerate(records, 1):
    natural_wait()
    if i % long_break_every == 0:
        long_break(20, 60)
        long_break_every = random.randint(30, 50)
```

#### 4-5. 데코이 검색 (5% 확률)

```python
DECOY_QUERIES = [
    "강남역 카페", "홍대 맛집", "여의도 공원 카페",
    "명동 쇼핑", "이태원 브런치", "성수동 카페",
    "잠실 맛집", "건대입구 술집",
]

if random.random() < 0.05 and i > 1:
    dq = random.choice(DECOY_QUERIES)
    decoy_url = f"https://pcmap.place.naver.com/place/list?query={urllib.parse.quote(dq)}"
    page.goto(decoy_url, wait_until="load", timeout=20000)
    natural_wait(mean=2.0, cap=6.0)
```

#### 4-6. 블록 감지 → 즉시 중단

```python
BLOCK_SIGNALS = [
    "비정상적인 접근", "비정상 접근",
    "자동화된 접근", "자동화 접근",
    "보안문자", "보안 문자",
    "captcha", "CAPTCHA", "reCAPTCHA",
    "일시적으로 차단", "접근이 차단",
    "잠시 후 다시",
]

def detect_block(page) -> str:
    try:
        text = page.inner_text("body", timeout=3000)[:3000]
    except Exception:
        return ""
    for sig in BLOCK_SIGNALS:
        if sig.lower() in text.lower():
            return sig
    return ""

# 매 페이지 방문 후 반드시 체크
blk = detect_block(page)
if blk:
    print(f"🚨 BLOCK 감지: {blk} — 즉시 중단, 6~12시간 후 재시도")
    break
```

##### ⚠️ False Positive 사례 — "잠시 후 다시" (2026-04-27)

`BLOCK_SIGNALS` 의 **"잠시 후 다시"** 는 가장 false positive 가 잦은 시그널.
한국어 안내문에 흔하게 등장하는 표현이라 **사장님이 작성한 정상 본문**이
차단 시그널로 오인되는 케이스가 발생.

**실전 케이스**: detail crawler 1647/2417 시점에 record_id `1450061243`
("아임키즈룸" 경산 키즈룸/장소대여) 페이지에서 차단 판정 → 전체 run 중단.
조사 결과 사장님 예약 안내문에 다음 문장이 있었음:

> "뒤로가기 시 발생하는 일시적 예약 오류는 5~10분 뒤 자동 해제됩니다.
> **잠시 후 다시 시도해 주시면** 감사하겠습니다."

`page.inner_text("body")[:3000]` 안에 매칭되어 차단 판정. 실제로는 정상 페이지(HTTP 200, 콘텐츠 풀로드).

**대응 옵션 (셋 중 택 1)**:

```python
# 옵션 A: BLOCK_SIGNALS 에서 "잠시 후 다시" 제거 (가장 단순)
BLOCK_SIGNALS = [
    "비정상적인 접근", "비정상 접근",
    "자동화된 접근", "자동화 접근",
    "보안문자", "보안 문자",
    "captcha", "CAPTCHA", "reCAPTCHA",
    "일시적으로 차단", "접근이 차단",
    # "잠시 후 다시",  # ← false positive 다발 — 제거
]

# 옵션 B: 더 구체적인 phrase 로 교체
BLOCK_SIGNALS = [
    ...,
    "잠시 후 다시 접속",          # 차단 페이지의 정확한 문구만
    "잠시 후 다시 시도해 주시기 바랍니다",  # 네이버 차단 안내문
]

# 옵션 C: 본문 전체가 아닌 차단 페이지 전용 셀렉터 검사
def detect_block(page) -> str:
    # 차단 페이지는 보통 .error_page 또는 별도 title
    title = page.title() or ""
    if any(s in title for s in ["접근 제한", "비정상", "차단"]):
        return f"title: {title}"
    err = page.locator(".error_page, #error_message").first
    if err.count() > 0:
        try:
            return err.inner_text(timeout=1000)[:50]
        except Exception:
            pass
    return ""

# 옵션 D: 그 ID 만 skip 리스트에 추가하고 계속 (당장 대응)
SKIP_IDS = {"1450061243"}  # known false-positive triggers
records = [r for r in records if str(r["id"]) not in SKIP_IDS]
```

**권장**: 신규 프로젝트는 **B** (정확한 문구 매칭). 운영 중인 크롤러는 **D** (즉시 우회) → 다음 maintenance 때 **B** 로 마이그레이션.

**검증 방법**: blocked 으로 멈추면 그 record_id 페이지를 직접 (다른 Chrome) 열어보고:
1. HTTP 200 인지
2. 콘텐츠가 풀로드 됐는지
3. 차단 시그널 텍스트가 *고객 안내문* 인지 *네이버 차단 안내* 인지
세 가지 확인. 1~3 모두 해당하면 false positive.

#### 4-7. 활동 시간대 제한

```python
def in_active_hours(start: int = 9, end: int = 23) -> bool:
    """새벽 크롤링은 봇 패턴 — 낮 시간대만 운영"""
    return start <= datetime.now().hour < end

if not in_active_hours():
    print("⛔ 활동시간대 아님 — 중단")
    sys.exit(0)
```

#### 4-8. 레코드 순서 셔플

```python
# 순차적 접근은 패턴 탐지됨 → 반드시 셔플
random.shuffle(records)
```

#### 4-9. 누적 볼륨 제어 (고빈도 재차단 방지)

**실전 사례** (2026-04-20): 같은 IP·같은 Chrome 프로필로 하루 2세션
(오전 320건 + 오후 165건 = 485건)을 돌렸더니 두 번째 세션 165건째에 블록 감지.
구현은 완벽(4-1 ~ 4-8 전부 적용)했지만 **누적 페이스가 네이버 기준
"동일 세션 집중 접근"으로 판정됨**.

##### (1) Long Break 상향 (방어 모드)
```python
# 기본: long_break(20, 60)     — 30~50건마다 20~60초
# 방어 모드: long_break(60, 180) — 30~50건마다 60~180초
long_break(min_s=60, max_s=180)
```
건당 평균 +1분, 시속 200~300건으로 감속. 하루 2,000건+ 장기 크롤에서 권장.

##### (2) 하루 상한 소프트캡 (자동 종료)
```python
from datetime import date

DAILY_CAP = 400  # 하루 처리 한도 (성공+실패 합산)

today = date.today().isoformat()
daily = progress.get("_daily_count", {}).get(today, 0)

for i, record in enumerate(records, 1):
    if daily >= DAILY_CAP:
        print(f"⛔ 하루 상한 {DAILY_CAP}건 도달 — 내일 재개")
        save_progress(progress)
        break
    # ... 처리 ...
    daily += 1
    progress.setdefault("_daily_count", {})[today] = daily
```
progress 파일에 `_daily_count: {"2026-04-20": 400}` 식으로 유지.
재실행 시 오늘 한도 도달 여부 확인 → 자동 종료.

##### (3) 프로필 2~3개 로테이션 (고급)
```
~/chrome-naver-crawl-A  (네이버 계정 A, 포트 9222)
~/chrome-naver-crawl-B  (네이버 계정 B, 포트 9223)
~/chrome-naver-crawl-C  (네이버 계정 C, 포트 9224)
```
- 세션마다 프로필 선택 (round-robin 또는 차단 감지 시 자동 교체)
- **주의**: 같은 IP에서 여러 계정 돌리면 오히려 의심 살 수 있음
  → IP도 분리(VPN·라우터 재연결)가 이상적
- **비용 대비**: 단일 프로젝트엔 과투자 — 10,000건+ 정기 크롤일 때만 의미

> 블록 회피 철학: "무한 방어"보다 **"걸릴 때 우아하게 중단 + 다음날 재개"** 가 현실적.
> 위 3가지는 누적 페이스를 여유있게 낮추는 안전 마진.

### 봇 회피 강도 요약

| 전략 | 효과 | 비용 |
|------|------|------|
| CDP 연결 (로그인 세션) | ★★★★★ | Chrome 한 번 실행 |
| 세션 웜업 | ★★★★☆ | +5~10초 |
| 지수분포 지연 | ★★★★☆ | 속도 50% 감소 |
| Long Break | ★★★★★ | 속도 30% 감소 |
| 인간 행동 시뮬 | ★★★☆☆ | 거의 없음 |
| 데코이 검색 | ★★★☆☆ | 5% 오버헤드 |
| 블록 감지 | ★★★★★ | 없음 (필수 안전장치) |
| 시간대 제한 | ★★★☆☆ | 없음 |
| 순서 셔플 | ★★★☆☆ | 없음 |
| **Long Break 상향 (방어 모드)** | ★★★★☆ | 추가 속도 50% 감소 |
| **하루 상한 소프트캡** | ★★★★★ | 하루 처리량 상한 (자동 종료) |
| **프로필 로테이션** | ★★★☆☆ | 계정 2~3개 + IP 분리 |

### 5. 백그라운드 실행 및 모니터링

```bash
# 백그라운드 실행 + 로그 저장
nohup python -u crawler.py >> crawler.log 2>&1 &
echo "PID: $!"
# -u 플래그: stdout 버퍼링 비활성화 → 로그 실시간 출력

# 진행 상황 모니터링
grep -c "저장완료" crawler.log
tail -10 crawler.log
ps aux | grep python | grep -v grep
```

크래시 자동 재시작 (bash):

```bash
#!/bin/bash
# auto_restart.sh
while true; do
    python -u crawler.py --force-hour >> crawler.log 2>&1
    echo "[$(date)] 크래시 감지 — 30초 후 재시작" >> crawler.log
    sleep 30
done
```

### 6. 자주 겪는 문제와 해결법

| 문제 | 원인 / 해결 |
|---|---|
| `TimeoutError: The read operation timed out` | 네이버 서버 응답 지연. try/except로 감싸고 다음 레코드로 continue |
| 이미지 URL 항상 None | og:image 없는 장소거나 로딩 불완전. `wait_until="domcontentloaded"` + 추가 대기 |
| `__APOLLO_STATE__ = null` | JS 실행 전 추출 시도. 페이지 이동 후 1.5초 이상 대기 |
| Chrome CDP 포트 연결 실패 | `--user-data-dir` 없이 실행했을 것. `http://localhost:9222/json` 직접 접속 테스트 |
| 검색 결과 엉뚱한 장소 매칭 | 이름이 너무 일반적. 주소 지역 힌트 쿼리에 추가 (`f"{name} {시도} {시군구}"`) |
| Supabase 1000행 제한 | offset 페이지네이션으로 배치 순회 |

### 카테고리 탭 사진 추출 (외부/내부) — JS-click + set-diff

`pcmap.place.naver.com/{type}/{id}/photo` 의 카테고리 탭(외부/내부/음식·음료/메뉴판/...)
별 사진을 추출하는 패턴. **URL 파라미터(`?subFilter=EXTERIOR`)는 동작 불안정**
(errorCode 408/AbortError) — JavaScript 클릭 + 차집합으로 분리.

#### URL 패턴

```
https://pcmap.place.naver.com/{place_type}/{place_id}/photo
```

`place_type` 우선순위 fallback: `restaurant` → `cafe` → `place`. 카테고리 탭은 `<li class="Zt2Kl">{label}</li>` (label = "외부", "내부", "음식·음료" 등).

#### 함정 1: URL 파라미터로는 안 됨

```python
# ❌ 동작 안함 — errorCode 408 또는 AbortError
url = f"https://pcmap.place.naver.com/restaurant/{id}/photo?filterType=AI%20View&subFilter=EXTERIOR"
```

#### 함정 2: 작은 viewport 에서 click 안 먹음

Headless Chrome 의 viewport 가 작으면(800x600 등) 탭이 화면 밖에 있어 click 이 트리거 안 됨. **`scrollIntoView` 후 `force=True` click 필수**:

```python
loc = page.locator(".Zt2Kl").filter(has_text="외부").first
loc.scroll_into_view_if_needed(timeout=3000)
loc.click(force=True, timeout=5000)
# URL 변화 확인
page.wait_for_url("**subFilter=EXTERIOR**", timeout=5000)
```

또는 evaluate 안에서 보강:
```js
const t = [...document.querySelectorAll('.Zt2Kl')].find(t => t.textContent.trim() === '외부');
t.scrollIntoView({behavior: 'instant', block: 'center'});
const a = t.querySelector('a') || t;
a.click();
// React onClick 대비 dispatch 시퀀스
for (const type of ['pointerdown','mousedown','pointerup','mouseup','click']) {
    a.dispatchEvent(new MouseEvent(type, {bubbles: true, cancelable: true, view: window}));
}
```

#### 함정 3: 단순 추출 → 외부/내부에 같은 사진 17장씩 섞임

각 탭 활성화 후 `document.querySelectorAll('img')` 로 추출하면, **헤더/sticky preview 의 카페 대표 이미지(~17장)** 가 두 탭 모두에 보여서 외부/내부 사진이 교차 오염됨.

**해결 — set-diff 차집합 방식**:

```python
def extract_per_tab(page) -> dict[str, str]:
    """현재 활성 탭의 모든 pstatic 이미지 → {origkey: best_url}"""
    return page.evaluate("""
    () => {
      function origKey(src) {
        if (!src) return '';
        if (src.includes('search.pstatic.net')) {
          const m = src.match(/[&?]src=([^&]+)/);
          if (m) return decodeURIComponent(m[1]);
        }
        return src.split('#')[0];
      }
      function sizeRank(src) {
        const m = src.match(/type=([wf]\\d+|f\\d+_\\d+)/);
        return m ? parseInt((m[1].match(/\\d+/) || [0])[0]) : 0;
      }
      const imgs = [...document.querySelectorAll('img')]
        .filter(i => /pstatic\\.net|naver\\.net/.test(i.src||''))
        .filter(i => !i.naturalWidth || i.naturalWidth >= 100);
      const byOrig = {};
      for (const i of imgs) {
        const key = origKey(i.src);
        if (!key) continue;
        if (!byOrig[key] || sizeRank(i.src) > sizeRank(byOrig[key])) byOrig[key] = i.src;
      }
      return byOrig;
    }
    """) or {}

# 외부 탭 클릭 → lazy-load 스크롤 → 추출
click_tab(page, "외부"); ext_raw = extract_per_tab(page)
# 내부 탭 클릭 → lazy-load 스크롤 → 추출
click_tab(page, "내부"); int_raw = extract_per_tab(page)

# 차집합으로 진짜 카테고리만 추리기
ext_keys = set(ext_raw.keys())
int_keys = set(int_raw.keys())
common = ext_keys & int_keys           # 헤더/sticky — 카페 대표 이미지 (~17장)
ext_only_keys = ext_keys - common      # 진짜 외부
int_only_keys = int_keys - common      # 진짜 내부
exterior_urls = [ext_raw[k] for k in ext_only_keys][:MAX_PER_TAB]
interior_urls = [int_raw[k] for k in int_only_keys][:MAX_PER_TAB]
```

**검증된 수치 (라운지티 테스트 카페)**:
- EXT 탭 활성화 시 보이는 unique URL: 48 (= 헤더 17 + 진짜 외부 31)
- INT 탭 활성화 시 보이는 unique URL: 55 (= 헤더 17 + 진짜 내부 38)
- 공통(common) = 17장 → **제거**, 외부 only 31장 + 내부 only 38장만 사용

#### `origkey()` — search.pstatic.net wrapper 분해

네이버 사진 URL 은 보통 wrapper 형식:
```
https://search.pstatic.net/common/?autoRotate=true&type=w560_sharpen&src=https%3A%2F%2Fldb-phinf.pstatic.net%2F.../IMG_9528.jpg
```

같은 원본 사진이 다양한 wrapper 사이즈(`w560`, `w278`, `f320_320` ...) 로 페이지에 여러 번 등장. dedup 시 wrapper URL 그대로 비교하면 같은 사진을 중복으로 셈. **`src=` 파라미터의 원본 URL 추출**해서 그것을 key 로 dedup:

```python
import re
from urllib.parse import unquote

def origkey(src: str) -> str:
    if not src: return ""
    if "search.pstatic.net" in src:
        m = re.search(r"[&?]src=([^&]+)", src)
        if m:
            return unquote(m.group(1))
    return src.split("#")[0]
```

### 빠른 참고 — URL 패턴

| 용도 | URL |
|------|-----|
| place_id로 직접 방문 | `https://map.naver.com/p/entry/place/{id}` |
| 이름으로 검색 | `https://pcmap.place.naver.com/place/list?query={검색어}` |
| 지도 홈 (웜업용) | `https://map.naver.com/` |
| 모바일 장소 페이지 | `https://m.place.naver.com/place/{id}/home` |
| 사진 페이지 (PC) | `https://pcmap.place.naver.com/{type}/{id}/photo` (`type` = restaurant/cafe/place) |
| 상세 페이지 (PC) | `https://pcmap.place.naver.com/place/{id}/home` |

### 권장 크롤링 속도

| 상황 | 권장 간격 | 긴 휴식 | 하루 상한 |
|------|----------|---------|---------|
| place_id 직접 방문 | 1.5~4초/건 | 30~50건마다 20~60초 | ~1,000건 |
| 검색 방식 | 2~5초/건 | 10~20건마다 30~120초 | ~500건 |
| **방어 모드** (장기 크롤) | 3~6초/건 | 30~50건마다 **60~180초** | **400건 (소프트캡)** |
| 블록 감지 후 | 6~12시간 대기 | - | - |

> **결론**: 시간당 약 400~600건 처리 가능 (place_id 직접 방문 기준) · 24시간 기준 최대 약 10,000건  
> **방어 모드**: 시속 200~300건 / 하루 400건 소프트캡 (경험칙 2026-04-20 — 같은 IP·프로필 누적 485건에서 블록)

---

## 실전 파이썬 템플릿 (층수/호수 크롤러)

`crawl_naver_floors_local.py` — 봇탐지 회피 전략 전부 적용, 층/호 파서 포함:

```python
# -*- coding: utf-8 -*-
"""
네이버 플레이스 층수/호수 크롤러 — 로컬 CDP 방식
- CDP 연결 (로그인된 Chrome 세션 재사용)
- 세션 웜업 + 지수분포 자연 지연 + Long Break
- 인간 행동 시뮬레이션 + 데코이 검색
- 블록 감지 → 즉시 중단
- 활동 시간대 제한 + 순서 셔플

추출 대상:
- floors : ["1층", "2층"] — 여러 층 점유 시 전부 리스트로
- unit   : "301호" — 호수가 있으면 추출 (optional)

사용법:
  1. (최초 1회) 별도 터미널에서: ./start_chrome.sh
  2. 열린 Chrome에서 naver.com 로그인
  3. python3 crawl_naver_floors_local.py
     (심야 강행: python3 crawl_naver_floors_local.py --force)
"""

import sys, re, json, time, random, subprocess, urllib.parse
from datetime import datetime
from pathlib import Path
from playwright.sync_api import sync_playwright

sys.stdout.reconfigure(encoding="utf-8", errors="replace")

# ── 경로 ─────────────────────────────────────────────────────────
BASE_DIR = Path(__file__).parent
INPUT_FILE = BASE_DIR / "large_cafes_final.json"
OUTPUT_FILE = BASE_DIR / "large_cafes_final.json"
PROGRESS_FILE = BASE_DIR / "progress_naver.json"

# ── 설정 ─────────────────────────────────────────────────────────
CDP_PORT = 9222
# ⚠️ 민감한 키는 환경변수/env 파일로 옮기세요. 코드에 하드코딩 금지.
API_KEY = "<YOUR_API_KEY>"  # 건축물대장 API — .env에서 로드 권장
API_BASE = "https://apis.data.go.kr/1613000/BldRgstHubService"

BLOCK_SIGNALS = [
    "비정상적인 접근", "비정상 접근", "자동화된 접근", "자동화 접근",
    "보안문자", "보안 문자", "captcha", "CAPTCHA", "reCAPTCHA",
    "일시적으로 차단", "접근이 차단", "잠시 후 다시", "서비스 이용이 제한",
]
DECOY_QUERIES = [
    "강남역 카페", "홍대 맛집", "여의도 공원 카페", "명동 쇼핑",
    "이태원 브런치", "성수동 카페", "잠실 맛집", "건대입구 술집", "한남동 카페",
]

# ── 봇 방지 유틸 ─────────────────────────────────────────────────
def natural_wait(mean=2.5, cap=15.0, floor_=0.8):
    t = random.expovariate(1.0 / mean)
    time.sleep(max(floor_, min(t, cap)))

def long_break(min_s=20, max_s=60):
    t = random.uniform(min_s, max_s)
    print(f"  [☕ long break {t:.0f}s]")
    time.sleep(t)

def human_move(page):
    try:
        page.mouse.move(random.randint(200, 900), random.randint(200, 600))
        time.sleep(random.uniform(0.2, 0.5))
        page.evaluate(f"window.scrollBy(0, {random.randint(80, 400)})")
        time.sleep(random.uniform(0.2, 0.7))
        if random.random() < 0.3:
            page.evaluate(f"window.scrollBy(0, -{random.randint(40, 200)})")
            time.sleep(0.2)
    except Exception:
        pass

def detect_block(page):
    try:
        text = page.inner_text("body", timeout=3000)[:3000]
        for sig in BLOCK_SIGNALS:
            if sig.lower() in text.lower():
                return sig
    except Exception:
        pass
    return ""

def in_active_hours(start=9, end=24):
    return start <= datetime.now().hour < end

# ── 층/호 파서 ───────────────────────────────────────────────────
def parse_floor_info(text):
    """
    텍스트에서 층 목록과 호수를 추출.
    반환:
      {"floors": ["1층", "2층"], "unit": "301호"}  # 호수 없으면 None
      None  (층 정보를 하나도 못 찾았을 때)
    """
    if not text:
        return None

    floors = []
    def add(f):
        if f and f not in floors:
            floors.append(f)

    consumed = text

    # 1) 전층 / 전관
    if re.search(r'전\s*(층|관)', consumed):
        add("전층")
        consumed = re.sub(r'전\s*(층|관)', lambda m: " " * len(m.group(0)), consumed)

    # 2) 지하 — "지하 N층" 또는 "BN" / "BN층"
    def _basement(m):
        n = m.group(1) or m.group(2)
        add(f"지{n}층")
        return " " * len(m.group(0))
    consumed = re.sub(r'지하\s*(\d+)\s*층', _basement, consumed)
    consumed = re.sub(r'\b[Bb](\d+)(?:\s*층)?\b', _basement, consumed)

    # 3) 지상 N층 (명시적)
    def _ground(m):
        add(f"{m.group(1)}층")
        return " " * len(m.group(0))
    consumed = re.sub(r'지상\s*(\d+)\s*층', _ground, consumed)

    # 4) 범위 "1~3층", "1층~3층", "1-3층"
    def _range(m):
        a, b = int(m.group(1)), int(m.group(2))
        if 1 <= a <= b and b - a <= 15:
            for n in range(a, b + 1):
                add(f"{n}층")
        return " " * len(m.group(0))
    consumed = re.sub(r'(\d+)\s*층?\s*[~\-–]\s*(\d+)\s*층', _range, consumed)

    # 5) 나열 "1,2층" / "1, 2, 3층" / "1·2·3층"
    def _multi(m):
        for num in re.split(r'[,·\s]+', m.group(1)):
            if num.strip().isdigit():
                add(f"{num.strip()}층")
        return " " * len(m.group(0))
    consumed = re.sub(r'((?:\d+\s*[,·]\s*)+\d+)\s*층', _multi, consumed)

    # 6) 단일 "N층"
    for m in re.finditer(r'(\d+)\s*층', consumed):
        add(f"{m.group(1)}층")
        break

    if not floors:
        return None

    # 호수
    unit = None
    m = re.search(r'(\d+)\s*호', text)
    if m:
        unit = f"{m.group(1)}호"

    return {"floors": floors, "unit": unit}

def extract_floor_info_from_page(page):
    """페이지 본문에서 층·호 정보 추출. 주소로 보이는 라인 우선."""
    try:
        text = page.inner_text("body", timeout=5000)
    except Exception:
        return None

    for line in text.split("\n"):
        line = line.strip()
        if not line or len(line) > 150:
            continue
        if "층" not in line and "호" not in line:
            continue
        if not any(kw in line for kw in ["층", "호", "번길", "로 ", "동 ", "구 ", "시 ", "군 "]):
            continue
        info = parse_floor_info(line)
        if info and info.get("floors"):
            info["raw"] = line[:120]
            return info

    info = parse_floor_info(text[:3000])
    if info and info.get("floors"):
        info["raw"] = "(fallback from body)"
        return info
    return None

# ── 메인 ─────────────────────────────────────────────────────────
def main():
    force = "--force" in sys.argv

    if not force and not in_active_hours():
        print("⛔ 활동시간대(9~24시) 아님. 심야에도 돌리려면 --force 플래그")
        return

    with open(INPUT_FILE) as f:
        cafes = json.load(f)

    progress = {}
    if PROGRESS_FILE.exists():
        with open(PROGRESS_FILE) as f:
            progress = json.load(f)
        # blocked는 재시도 대상으로 간주
        progress = {k: v for k, v in progress.items() if v.get("error") != "blocked"}

    # 타겟: naver_place_id 있고, cafeArea 아직 없고, progress에 없는 것
    targets = [
        c for c in cafes
        if str(c["id"]) not in progress
        and not (c.get("building") or {}).get("cafeArea")
        and c.get("naver_place_id")
    ]

    print(f"대상: {len(targets)}건 (이전 진행: {len(progress)}건)")
    if not targets:
        print("처리할 항목 없음.")
        return

    random.shuffle(targets)

    with sync_playwright() as p:
        try:
            browser = p.chromium.connect_over_cdp(f"http://localhost:{CDP_PORT}")
        except Exception as e:
            print(f"❌ Chrome CDP 연결 실패 (localhost:{CDP_PORT}): {e}")
            print(f"   먼저 ./start_chrome.sh 를 실행하고 네이버 로그인하세요.")
            return

        if not browser.contexts:
            print("❌ 기존 컨텍스트 없음. Chrome을 --user-data-dir로 실행했는지 확인")
            return

        ctx = browser.contexts[0]
        page = ctx.new_page()
        page.on("dialog", lambda d: d.dismiss())

        try:
            # 세션 웜업
            print("[warmup] map.naver.com 방문...")
            page.goto("https://map.naver.com/", wait_until="load", timeout=20000)
            natural_wait(mean=4.0, cap=10.0, floor_=2.0)
            human_move(page)
            natural_wait(mean=2.0, cap=5.0, floor_=1.0)

            success = fail = 0
            long_break_every = random.randint(30, 50)
            start_time = time.time()

            for i, c in enumerate(targets, 1):
                if not force and not in_active_hours():
                    print(f"⛔ 활동시간대 벗어남 — 중단 ({i-1}/{len(targets)})")
                    break

                cid = str(c["id"])
                pid = c["naver_place_id"]
                name = (c.get("name") or "")[:20]

                # 데코이 (5%)
                if random.random() < 0.05 and i > 1:
                    dq = random.choice(DECOY_QUERIES)
                    print(f"  [decoy: {dq}]")
                    try:
                        page.goto(
                            f"https://pcmap.place.naver.com/place/list?query={urllib.parse.quote(dq)}",
                            wait_until="load", timeout=20000
                        )
                        natural_wait(mean=2.0, cap=5.0)
                    except Exception:
                        pass

                # 추출
                try:
                    url = f"https://m.place.naver.com/place/{pid}/home"
                    page.goto(url, wait_until="domcontentloaded", timeout=25000)
                    natural_wait(mean=2.5, cap=8.0, floor_=1.5)
                    human_move(page)

                    blk = detect_block(page)
                    if blk:
                        progress[cid] = {"error": "blocked", "sig": blk}
                        fail += 1
                        print(f"🚨 BLOCK 감지 ({blk}) — 즉시 중단")
                        print(f"   6~12시간 후 재시도 권장")
                        break

                    info = extract_floor_info_from_page(page)
                    if info and info.get("floors"):
                        entry = {"floors": info["floors"]}
                        if info.get("unit"):
                            entry["unit"] = info["unit"]
                        progress[cid] = entry
                        success += 1
                        floors_str = "·".join(info["floors"])
                        unit_str = f" {info['unit']}" if info.get("unit") else ""
                        print(f"[{i}/{len(targets)}] {name}: {floors_str}{unit_str} ✓")
                    else:
                        progress[cid] = {"floors": None, "error": "not_found"}
                        fail += 1

                except Exception as e:
                    progress[cid] = {"floors": None, "error": str(e)[:60]}
                    fail += 1

                if i % 10 == 0:
                    with open(PROGRESS_FILE, "w", encoding="utf-8") as f:
                        json.dump(progress, f, ensure_ascii=False, indent=2)

                natural_wait(mean=2.5, cap=10.0, floor_=1.0)

                if i % long_break_every == 0:
                    long_break(20, 60)
                    long_break_every = random.randint(30, 50)

            with open(PROGRESS_FILE, "w", encoding="utf-8") as f:
                json.dump(progress, f, ensure_ascii=False, indent=2)
            print(f"\n크롤링 완료: 성공 {success} / 실패 {fail}")

        finally:
            try:
                page.close()
            except Exception:
                pass

if __name__ == "__main__":
    main()
```

---

## Chrome 실행 스크립트 (macOS)

`start_chrome.sh`:

```bash
#!/bin/bash
# 네이버 크롤링용 전용 Chrome 프로필 실행
# 이 프로필은 평소 쓰는 Chrome과 완전히 분리됨 (기존 Chrome 세션에 영향 없음)

PROFILE_DIR="$HOME/chrome-naver-crawl"

echo "🌐 크롤링 전용 Chrome 실행 중..."
echo "   프로필: $PROFILE_DIR"
echo "   CDP 포트: 9222"
echo ""
echo "📌 최초 1번만: 열린 Chrome에서 네이버(naver.com) 로그인하세요."
echo "📌 이후엔 이 창만 켜두면 자동 로그인 유지됩니다."
echo ""

"/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" \
  --user-data-dir="$PROFILE_DIR" \
  --remote-debugging-port=9222 \
  --no-first-run \
  --no-default-browser-check \
  "https://nid.naver.com/nidlogin.login"
```

실행 권한: `chmod +x start_chrome.sh`

---

## 구성 요약 (이 패턴의 가치)

이 패턴은 **4가지 전략이 맞물려야** 효과를 냅니다. 하나라도 빠지면 차단률이 급등:

1. **로그인 세션 재사용** (CDP + `--user-data-dir`)
2. **인간화된 행동 패턴** (지연 분포, 스크롤, 데코이)
3. **블록 조기 감지** (매 페이지 후 시그널 스캔)
4. **속도 자발적 제한** (Long Break, 시간대)

재사용 시 `parse_floor_info()` / `extract_*_from_page()` 만 바꾸면 다른 데이터(카테고리, 별점, 영업시간 등)도 같은 방식으로 수집 가능.
jt · v1 · CC0-1.0 · 복사 0