백업: DB수집 전체 스냅샷 (공공기관2 정리 전)

공공기관2 작업 중. _temp 몽타주(재생성가능)는 제외.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
hehihoho3 2026-06-18 18:15:40 +09:00
commit df16c98366
2133 changed files with 530967 additions and 0 deletions

View File

@ -0,0 +1 @@
{"sessionId":"3c6000fc-2299-4a1f-993b-223dcca20883","pid":20156,"acquiredAt":1779975159075}

12
.gitignore vendored Normal file
View File

@ -0,0 +1,12 @@
# 재생성 가능 임시/대용량 (push 제외)
/_temp/
작업파일/_temp/
**/__pycache__/
**/.playwright-mcp/
**/_nshots/
**/_naudit/
**/nshot_*/
~$*.xlsx
*.pyc
*.tmp
.DS_Store

35441
.owncloudsync.log Normal file

File diff suppressed because it is too large Load Diff

141793
.owncloudsync.log.1 Normal file

File diff suppressed because it is too large Load Diff

BIN
.sync_journal.db Normal file

Binary file not shown.

BIN
.sync_journal.db-shm Normal file

Binary file not shown.

BIN
.sync_journal.db-wal Normal file

Binary file not shown.

38
2단계 안내사항.txt Normal file
View File

@ -0,0 +1,38 @@
※ 각자 시트에 게시판명/카테고리(소-세부사항2) 열 추가해주세요 ※
--------------------------------------------------------------------------------------
[분배 완료 - 17개 지방자치단체]
강원특별자치도/경기도/경상남도/경상북도/광주광역시/대구광역시/대전광역시/부산광역시/서울특별시/세종특별자치시/울산광역시/인천광역시/전라남도/제주특별자치도/충청북도/전북특별자치도/충청남도
▶ 김재민: 경상남도, 울산광역시, 충청남도
(강릉시, 속초시)
▶ 구소영: 경상북도,서울특별시
(고성군,강원도경찰청)
▶ 심예진: 강원특별자치도, 대전광역시, 세종특별자치시, 제주특별자치도
(동해시, 강원특별자치도 교육청)
▶ 방유정: 경기도, 부산광역시, 전북특별자치도
(삼척시, 강원영동병무지청, 강원지방기상청, 강원지방병무청, 강원특별자치도 양구군)
▶ 황의정: 광주광역시
▶ 정대원: 대구광역시,충청북도
▶ 프리랜서1: 인천광역시, 전라남도, 고용노동부, 과학기술정보통신부, 교육부
--------------------------------------------------------------------------------------
✓ 대분류:지방자치단체 - 17개 우선적으로 작성
✓ 중분류:부/처/청 - 상단기관만 작성(단, 중분류에 '_기타' 들어가는건 제외)
① '게시물형태' == '사이트' : '수량'~'마크 하이퍼링크' 공백
② '저작물 유형' & '공공누리 연계' 작성 X
③ '공공누리 부착' 유형 많을 경우에는 모두 작성
ex) 1유형, 4유형
④ 게시판에 게시물이 0개인 경우/자물쇠가 걸려있는 경우 '공공누리 부착' == '미부착'
⑤ 게시판 및 게시물 동시에 공공누리 마크가 부착되어 있다면 '마크 부착위치' == '게시판'
⑥ 페이지에 공공누리 마크가 부착되어 있다면 '마크 부착위치' == '게시물'
⑦ 로그인/본인인증 해야 보이는 페이지들은 모두 삭제

5
Desktop.ini Normal file
View File

@ -0,0 +1,5 @@
[.ShellClassInfo]
IconResource=D:\\ownCloud\\owncloud.exe
[ownCloud]
UpdateIcon=true

119
README.md Normal file
View File

@ -0,0 +1,119 @@
# 공공저작물 실태조사 DB수집
신유형개방지원사업 대상기관(1,160개)의 홈페이지를 사이트맵 단위로 점검하여
공공누리 부착 여부·게시물 형태·저작물 유형을 엑셀로 정리하는 데이터 수집 프로젝트.
루트 디렉터리: `D:\01.프로젝트\DB수집`
---
## 1. 프로젝트 구성
```
DB수집/
├─ ★[아이티앤] 공공저작물 자료조사 매뉴얼자료Ver1.2_260525.hwp # 조사 매뉴얼(v1.2)
├─ ★공공저작물 실태조사 조사원 자료 취합본(관리자)_2605212137.xlsx # 관리자 취합본
├─ 붙임1_...개방 대상기관(1160개) 실태조사 목록...양식_260518.xlsx # 1,160개 기관 마스터 양식
├─ 자료_취합_예시.xlsx # 작성 예시
├─ 2단계 안내사항.txt # 2단계 분배·작성 규칙
├─ 하이퍼링크활성화.py # K열 URL을 엑셀 하이퍼링크로 변환
├─ 1주차(5.19~5.25)/ # 지자체 2개 + 부처 3개 완료
│ ├─ 인천광역시(완료)/ # 분야별 메뉴 .txt + 작업본 xlsx + 크롤러
│ ├─ 전라남도(완료)/
│ ├─ 고용노동부(완료)/
│ ├─ 과학기술정보통신부(완료)/
│ ├─ 교육부(완료)/
│ └─ 수정/ # 샘플 .xlsm + 심예진 수정본
└─ 2주차/ # 부처 3개 진행
├─ 보건복지부/ # test.py = mohw.go.kr 크롤러
├─ 성평등가족부/ # test.py = 저작물 8유형 자동 분류기
└─ 외교부/ # test.py = /list.do 게시판 건수 수집기
```
각 부처/지자체 폴더의 전형적 구성:
- `<기관명>_홈페이지_사이트_*.xlsx` — 단계별 작업본(r1, r2, ... 결과, 완료)
- `<기관명>_메뉴.txt` / `메뉴구조.txt` — 사이트맵 메뉴 텍스트 덤프
- `*.py` — 해당 기관 사이트 구조에 맞춘 크롤링 스크립트
---
## 2. 엑셀 데이터 스키마 (작업 컬럼)
크롤러 코드 기준 핵심 열:
| 열 | 번호 | 의미 | 채워지는 방법 |
|----|----|------|-------------|
| K | 11 | URL 주소 | 수기 입력 (1차 사이트맵 수집) |
| L | 12 | 게시판/페이지 구분 | 크롤러가 URL 패턴·DOM으로 자동 판별 |
| M | 13 | 게시물 수량 | 게시판일 때 `total`·`totalCount` 등에서 추출 |
| N | 14 | 저작물 유형 | 본문 분석 → 어문/이미지/영상/오디오/글꼴/3D/기타/없음 |
| O | 15 | 공공누리 부착 유형 | `kogl.or.kr/info/licenseType` 링크에서 1~4유형 추출 |
| P | 16 | 마크 부착위치 | 게시판/게시물 |
| Q | 17 | 부착 여부 보조 | (스크립트별 상이) |
데이터는 보통 3행 또는 4행부터 시작(헤더 2행).
---
## 3. 작성 규칙 (2단계 안내사항 발췌)
- 대분류는 **지방자치단체 17개 우선**, 중분류는 **부/처/청 상단기관**만 (단, `_기타` 포함 제외).
- `게시물형태 == 사이트`: `수량`~`마크 하이퍼링크` 공백 처리.
- `저작물 유형``공공누리 연계`는 작성 X.
- `공공누리 부착` 유형이 여러 개이면 모두 기재(예: `1유형, 4유형`).
- 게시물 0건 또는 자물쇠 페이지 → `공공누리 부착 = 미부착`.
- 게시판·게시물 동시 부착 → `마크 부착위치 = 게시판`.
- 페이지에만 부착 → `마크 부착위치 = 게시물`.
- 로그인·본인인증 페이지는 전부 삭제.
분배 현황(2단계 17개 지자체):
김재민(경남/울산/충남) · 구소영(경북/서울) · 심예진(강원/대전/세종/제주) ·
방유정(경기/부산/전북) · 황의정(광주) · 정대원(대구/충북) ·
프리랜서1(인천/전남/고용노동부/과기정통부/교육부).
---
## 4. 스크립트 카탈로그
루트 공용:
- `하이퍼링크활성화.py` — K열의 http(s) URL을 엑셀 하이퍼링크로 변환 + 파란색 밑줄 스타일 적용. 입력/출력 경로는 스크립트 내부 상수.
기관별 크롤러(공통 패턴: `requests` + `BeautifulSoup` + `openpyxl`로 기존 서식을 보존하며 셀만 갱신):
| 위치 | 용도 |
|------|------|
| `2주차/외교부/test.py` | URL이 `/list.do`로 끝나면 L열에 "게시판" 표시, `<div class="total"><span>` 건수를 M열 기입 |
| `2주차/성평등가족부/test.py` | 게시판 목록에서 `fn_selectView(n)` 패턴으로 상세 URL 최대 5개 수집 → 본문(`table.brdView01` 또는 `div#contents`) 분석 → 저작물 8유형 판별 |
| `2주차/보건복지부/test.py` | `mohw.go.kr` 페이지 판별, `#totalCount` 또는 `span.total b`로 건수 수집, 상세글 `article.board_view` 본문 영역 분석 |
| `1주차/인천광역시(완료)/페이지_마크찾기.py` | `<main class="content">``kogl.or.kr/info/licenseType{n}` 링크에서 공공누리 유형 추출 → O열 기입 |
| `1주차/인천광역시(완료)/저작물유형구분.py` | 저작물 유형 자동 분류 |
| `1주차/전라남도(완료)/전라남도_페이지_마크찾기_게시판.py` | 전남 게시판형 마크 탐지 |
| `1주차/전라남도(완료)/페이지_마크찾기_페이지.py` | 전남 페이지형 마크 탐지 |
| `1주차/고용노동부(완료)/페이지_마크찾기_allinone.py` | 고용부 전 영역 통합 마크 탐지 |
공통 동작 원칙:
- `User-Agent`를 Chrome으로 위장, 0.5~1초 슬리프로 차단 회피.
- `openpyxl.load_workbook` → 셀만 수정 → 새 파일명(`_r2`, `_결과`, `_완료` 등)으로 저장.
- 본문 영역 셀렉터는 사이트별로 다름(상세 코드 주석 참조).
---
## 5. 산출물 네이밍 컨벤션
같은 데이터를 단계적으로 가공하면서 접미사로 버전 표시:
`{기관명}_홈페이지_사이트_{r1|r2|r3|결과|완료}.xlsx`
예) `외교부.xlsx → 외교부_업데이트.xlsx → ..._활성화_r4.xlsx → ..._활성화_r5.xlsx → 외교부_완료.xlsx`
`_완료` 또는 `_완료_ai` 접미사가 붙은 파일이 최종본.
---
## 6. 실행 환경
- Python 3 + `openpyxl`, `requests`, `beautifulsoup4`, `urllib3`.
- Windows 경로(`D:\01.프로젝트\DB수집`)와 `ownCloud\알바\...` 경로가 스크립트에 혼재 — 실행 전 입출력 경로 확인 필요.
- 일부 사이트(성평등가족부 등) SSL 검증 우회(`verify=False`) 사용.

143
Superpowers_사용법.md Normal file
View File

@ -0,0 +1,143 @@
# Superpowers 사용법
> Claude Code용 플러그인. 버전 **5.1.0** (claude-plugins-official)
> 제작: Jesse Vincent · https://github.com/obra/superpowers
---
## 1. Superpowers가 뭔가
Superpowers는 **코딩 에이전트에게 "개발 방법론"을 입히는 스킬 모음집**이다.
설치하면 Claude가 코드를 짤 때 곧장 코드부터 쓰지 않고,
1. **무엇을 만들려는지 먼저 캐묻고(brainstorming)**
2. **설계 문서를 보여주고 승인을 받고**
3. **실행 계획서(plan)를 만들고**
4. **TDD(빨강→초록) 기반으로 한 단계씩 구현하고**
5. **검증하고 코드 리뷰까지** 하는
체계적인 흐름을 자동으로 따른다.
핵심은 **"스킬이 자동으로 발동된다"**는 것. 사용자가 특별히 명령하지 않아도, Claude가 상황을 보고 알아서 맞는 스킬을 꺼내 쓴다.
---
## 2. 가장 중요한 사실 — 따로 외울 명령어가 없다
Superpowers는 슬래시 명령(`/superpowers ...` 같은 것)이 **없다.**
대신 **14개의 스킬**로 구성되어 있고, Claude가 대화 맥락에 맞춰 **알아서 호출**한다.
| 사용자가 하는 일 | Claude가 자동으로 하는 일 |
|---|---|
| "○○ 기능 만들어줘" | `brainstorming` 발동 → 질문으로 요구사항 정리 |
| 설계 승인함 | `writing-plans` 발동 → 실행 계획서 작성 |
| "이제 진행해" | `test-driven-development` + `executing-plans` 발동 |
| 버그/에러 발생 | `systematic-debugging` 발동 |
| "다 됐어/끝났어" | `verification-before-completion``requesting-code-review` |
**평소처럼 한국어로 시키기만 하면 된다.** Claude가 "Using [스킬명] to [목적]" 이라고 알리고 해당 절차를 따른다.
---
## 3. 14개 스킬 한눈에 보기
### 🧭 시작 & 설계
| 스킬 | 언제 발동 |
|---|---|
| **using-superpowers** | 모든 대화 시작 시 — 어떤 스킬을 쓸지 판단하는 진입점 |
| **brainstorming** | 새 기능·컴포넌트·동작 변경 등 **창작 작업 전 필수**. 의도·요구사항·설계를 먼저 탐색 |
| **writing-plans** | 스펙이 정해진 다단계 작업의 **구현 계획서** 작성 |
### 🔨 구현
| 스킬 | 언제 발동 |
|---|---|
| **test-driven-development** | 모든 기능/버그픽스 구현 시 — 구현 코드 전에 테스트 먼저 |
| **executing-plans** | 작성된 계획서를 **별도 세션에서** 체크포인트 단위로 실행 |
| **subagent-driven-development** | 계획서의 독립 작업들을 **현재 세션에서** 서브에이전트로 실행 |
| **dispatching-parallel-agents** | 의존성 없는 2개 이상 작업을 **병렬** 처리 |
| **using-git-worktrees** | 현재 작업공간과 격리된 별도 워크스페이스가 필요할 때 |
### 🐞 디버깅 & 검증
| 스킬 | 언제 발동 |
|---|---|
| **systematic-debugging** | 버그·테스트 실패·예상 밖 동작 — **고치기 전에** 원인부터 체계적으로 |
| **verification-before-completion** | "완료/수정됨/통과" 라고 말하기 전 — 실제로 명령 돌려 **증거** 확인 |
### 👀 코드 리뷰 & 마무리
| 스킬 | 언제 발동 |
|---|---|
| **requesting-code-review** | 작업 완료·주요 기능 구현·머지 전 검증 |
| **receiving-code-review** | 리뷰 피드백 받았을 때 — 맹목적 수용 말고 기술적 검증 |
| **finishing-a-development-branch** | 구현 끝 + 테스트 통과 후 머지/PR/정리 결정 |
### 🛠 메타
| 스킬 | 언제 발동 |
|---|---|
| **writing-skills** | 새 스킬을 만들거나 기존 스킬을 수정할 때 |
---
## 4. 전형적인 작업 흐름 (예시)
```
나: "사이트맵 수집 결과를 검증하는 작은 스크립트 만들어줘"
[brainstorming] Claude가 질문을 하나씩 던짐
- 입력 형식? 어떤 검증? 실패 시 동작?
- 2~3가지 접근법 + 추천안 제시
- 설계를 짧게 정리해 보여주고 승인 요청
- docs/superpowers/specs/YYYY-MM-DD-검증스크립트-design.md 로 저장
▼ (내가 "좋아" 승인)
[writing-plans] 구현 계획서 작성 → docs/superpowers/plans/ 에 저장
▼ (내가 "진행해")
[test-driven-development] 테스트 먼저 작성(빨강) → 구현(초록) → 리팩터
[subagent-driven-development] 작업 단위로 서브에이전트가 진행/검토
[verification-before-completion] 실제로 돌려서 통과 증거 확보
[requesting-code-review] 변경분 리뷰
[finishing-a-development-branch] 머지/PR/정리 옵션 제시
```
산출물 저장 위치:
- 설계 문서 → `docs/superpowers/specs/`
- 계획서 → `docs/superpowers/plans/`
---
## 5. 실전 팁
- **그냥 평소대로 시키면 된다.** "○○ 만들어줘"라고 하면 brainstorming부터 시작한다.
- **간단한 작업이라도 설계 단계를 거친다.** "이건 너무 간단한데" 싶어도 Claude가 짧게라도 설계를 보여주고 승인을 받는다 — 이게 의도된 동작이다.
- **특정 스킬을 직접 부르고 싶으면** 이름을 말하면 된다. 예: "systematic-debugging 스킬 써서 이 에러 봐줘"
- **건너뛰고 싶으면 말하면 된다.** 사용자 지시가 스킬보다 우선이다. 예: "설계 단계 생략하고 바로 짜줘", "TDD 말고 그냥 구현해줘" → Claude가 따른다.
- **CLAUDE.md 규칙이 항상 최우선.** 우선순위: ① 사용자 지시(CLAUDE.md·직접 요청) → ② Superpowers 스킬 → ③ 기본 동작.
---
## 6. 설치 / 관리 명령 (Claude Code 터미널)
```bash
# 설치 (이미 설치됨)
/plugin install superpowers@claude-plugins-official
# 적용 (설치 직후 1회)
/reload-plugins
# 플러그인 관리 UI
/plugin
```
설치 경로:
`~/.claude/plugins/cache/claude-plugins-official/superpowers/5.1.0/`
---
## 7. 한 줄 요약
> **"평소처럼 한국어로 작업을 시켜라. Superpowers가 알아서 설계 → 계획 → TDD 구현 → 검증 → 리뷰 흐름을 태운다. 건너뛰고 싶으면 그렇게 말하면 된다."**

10
_avg_result.txt Normal file
View File

@ -0,0 +1,10 @@
01_김재민 4463 5 892.6
02_심예진 4433 5 886.6
03_방유정 4499 7 642.7
04_구소영 3234 9 359.3
05_정대원 2639 8 329.9
06_황의정 3512 6 585.3
07_프리랜서 2224 3 741.3
08_이희재 1372 5 274.4
09_신성범 3677 4 919.2
10_신규 0 0 0.0

BIN
_test_buyeo.xlsx Normal file

Binary file not shown.

55
_김제_E셸삭제.py Normal file
View File

@ -0,0 +1,55 @@
# -*- coding: utf-8 -*-
"""김제시: E가 병합(같은 D,E 연속 ≥2행)인데 그 블록 '첫 행'의 F가 비어있으면 그 행 삭제(중분류 랜딩 셸).
F 카테고리 라벨은 자식행이 병합 승계. 사용: python -X utf8 _김제_E셸삭제.py plan | run
"""
import os, sys, re, io, shutil, importlib.util, warnings
warnings.filterwarnings('ignore')
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8', line_buffering=True)
import openpyxl
HERE = os.path.dirname(os.path.abspath(__file__))
def _imp(n, p):
s = importlib.util.spec_from_file_location(n, p); m = importlib.util.module_from_spec(s); s.loader.exec_module(m); return m
dd = _imp('dd', os.path.join(HERE, '_스크립트', '_dedup_all.py'))
XP = os.path.join(HERE, '작업파일', '광역_사이트맵', '전북특별자치도', '3.김제시', '전북특별자치도_김제시.xlsx')
def lab(v):
return ' > '.join(str(v.get(c)) for c in range(4, 11) if v.get(c) not in (None, ''))
def main():
mode = sys.argv[1] if len(sys.argv) > 1 else 'plan'
wb = openpyxl.load_workbook(XP); ws = wb.active
rows = dd.load_flat(ws)
# E-block = 같은 (D,E) 연속 run, E 비어있지 않음, run≥2
delete = set()
n = len(rows); i = 0
while i < n:
v = rows[i]['vals']
E = v.get(5)
if E in (None, ''):
i += 1; continue
j = i
while j + 1 < n and rows[j+1]['vals'].get(5) == E and rows[j+1]['vals'].get(4) == v.get(4):
j += 1
runlen = j - i + 1
if runlen >= 2 and rows[i]['vals'].get(6) in (None, ''):
delete.add(id(rows[i]))
i = j + 1
del_rows = [r for r in rows if id(r) in delete]
keep = [r for r in rows if id(r) not in delete]
print('=== 김제 E셸삭제 plan ===')
print(f'기존행 {len(rows)}{len(keep)} (삭제 {len(del_rows)})')
for r in del_rows:
v = r['vals']
print(f' r{r["src"]}: {lab(v)} | L={v.get(12)} M={v.get(13)} | K={str(v.get(11))[-38:]}')
if mode != 'run':
return
bak = XP.replace('.xlsx', '_backup_E셸전.xlsx')
shutil.copy(XP, bak)
dd.write_back(ws, keep)
try:
wb.save(XP); print(f'\n저장완료 {len(rows)}{len(keep)}행. 백업 {os.path.basename(bak)}')
except PermissionError:
wb.save(XP.replace('.xlsx', '_LP.xlsx')); print('\n!! 잠김 → _LP')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,116 @@
# -*- coding: utf-8 -*-
"""김제시 인페이지 탭(#anchor) 수량 → M 기입.
페이지 본문의 div[class*=basic_tab] 링크가 전부 '#anchor' 그룹( basic_tab2>ul.col4,
#tab1~#tabN)을 인페이지 스크롤탭으로 보고 M=탭수 기입(공주/금산 전례, 매뉴얼 1-5b: #anchor는 행 분리 안함).
menuCd 형제 nav(basic_tab depth4 , href=실제 .gimje) 제외. L은 페이지 유지.
대상: 17~(요청). 3~16행은 참고 보고만.
사용: python -X utf8 _김제_인페이지탭수.py dry | run
"""
import os, sys, re, io, time, shutil
from concurrent.futures import ThreadPoolExecutor, as_completed
import warnings; warnings.filterwarnings('ignore')
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8', line_buffering=True)
import requests
from bs4 import BeautifulSoup
import openpyxl
HERE = os.path.dirname(os.path.abspath(__file__))
XP = os.path.join(HERE, '작업파일', '광역_사이트맵', '전북특별자치도', '3.김제시', '전북특별자치도_김제시.xlsx')
UA = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/120 Safari/537.36'}
def fetch(session, url):
r = session.get(url, headers=UA, verify=False, timeout=10, allow_redirects=True)
meta = re.search(rb'charset=["\']?\s*([\w-]+)', r.content[:3000], re.I)
r.encoding = meta.group(1).decode('ascii', 'ignore') if meta else r.apparent_encoding
return BeautifulSoup(r.text, 'html.parser')
BASIC_TAB = re.compile('basic_tab')
def inpage_tabs(soup):
"""가장 큰 '전부 #anchor' basic_tab 그룹의 (탭수, 라벨)."""
best, labels = 0, None
for d in soup.find_all('div', class_=BASIC_TAB):
ul = d.find('ul')
if not ul:
continue
links = ul.select('li > a')
if len(links) < 2:
continue
hrefs = [(a.get('href') or '').strip() for a in links]
if all(h.startswith('#') for h in hrefs):
if len(links) > best:
best = len(links)
labels = [a.get_text(strip=True) for a in links]
return best, labels
def main():
mode = sys.argv[1] if len(sys.argv) > 1 else 'dry'
wb = openpyxl.load_workbook(XP); ws = wb.active
last = max(r for r in range(3, ws.max_row + 1) if ws.cell(r, 2).value not in (None, ''))
def lab(r):
return ' > '.join(str(ws.cell(r, c).value) for c in range(4, 11) if ws.cell(r, c).value not in (None, ''))
# 대상 수집: L=페이지 & http URL
def targets(lo, hi):
out = []
for r in range(lo, hi + 1):
L = ws.cell(r, 12).value
url = ws.cell(r, 11).value
if L == '페이지' and isinstance(url, str) and url.startswith('http'):
out.append((r, url))
return out
main_t = targets(17, last)
pre_t = targets(3, 16)
session = requests.Session()
def scan(rows_urls):
res = {}
def w(ru):
r, u = ru
try:
n, ls = inpage_tabs(fetch(session, u))
return r, n, ls
except Exception:
return r, 0, None
with ThreadPoolExecutor(max_workers=10) as ex:
for f in as_completed([ex.submit(w, ru) for ru in rows_urls]):
r, n, ls = f.result()
if n >= 2:
res[r] = (n, ls)
return res
print(f'스캔: 17~{last} ({len(main_t)}개 페이지), 참고 3~16 ({len(pre_t)}개)')
t0 = time.time()
main_hits = scan(main_t)
pre_hits = scan(pre_t)
print(f'스캔완료 {time.time()-t0:.0f}s')
print(f'\n[대상 17~끝] 인페이지탭 발견 {len(main_hits)}행:')
for r in sorted(main_hits):
n, ls = main_hits[r]
old = ws.cell(r, 13).value
print(f' r{r}: M {old}{n} [{lab(r)}] 탭={ls}')
if pre_hits:
print(f'\n[참고 3~16] 인페이지탭 {len(pre_hits)}행 (요청범위 밖, 미적용):')
for r in sorted(pre_hits):
n, ls = pre_hits[r]
print(f' r{r}: M {ws.cell(r,13).value}{n}? [{lab(r)}] 탭={ls}')
if mode != 'run':
return
bak = XP.replace('.xlsx', '_backup_인페이지탭전.xlsx')
shutil.copy(XP, bak)
for r, (n, ls) in main_hits.items():
ws.cell(r, 13).value = n
try:
wb.save(XP); print(f'\n저장완료. {len(main_hits)}행 M 갱신. 백업 {os.path.basename(bak)}')
except PermissionError:
wb.save(XP.replace('.xlsx', '_LP.xlsx')); print('\n!! 잠김 → _LP')
if __name__ == '__main__':
main()

170
_김제_재분류.py Normal file
View File

@ -0,0 +1,170 @@
# -*- coding: utf-8 -*-
"""김제시 재검토: 게시판↔페이지 재분류 + 페이지 인페이지탭(basic_tab) 부모 병합.
- 판별: 본문에 게시판 리스트 클래스(bbs_list/news_list/photo_list/video_list/magazine_list/gallery_list/board_list) 게시판, 없으면 페이지.
- 외부 리다이렉트/ERR (새창열림 외부시스템) 손대지 않음.
- basic_tab 탭그룹: 페이지 탭만 부모(페이지) 접고 M=페이지탭수. 게시판 탭은 별도 유지.
사용: python -X utf8 _김제_재분류.py dry | run
"""
import os, sys, re, io, hashlib, shutil, importlib.util, warnings, pickle
warnings.filterwarnings('ignore')
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8')
import openpyxl
HERE=os.path.dirname(os.path.abspath(__file__))
spec=importlib.util.spec_from_file_location('dd', os.path.join(HERE,'_스크립트','_dedup_all.py'))
dd=importlib.util.module_from_spec(spec); spec.loader.exec_module(dd)
XP=r"작업파일\광역_사이트맵\전북특별자치도\3.김제시\전북특별자치도_김제시.xlsx"
CACHE=r"_cache_gimje"
LIST=re.compile(r'class="[^"]*(bbs_list|news_list|photo_list|video_list|magazine_list|gallery_list|board_list)[^"]*"')
TAB_DIV=re.compile(r'<div class="basic_tab[^"]*">(.*?)</div>', re.S)
EXTERNAL={69,249,440,441,148,195,277,318,390,400,134,368} # 외부 리다이렉트/ERR/외부새창 (현행 유지)
TYPE_ORDER=['어문','이미지','영상','음악','소프트웨어','데이터','3D','기타']
def union_types(vals_list):
seen=[]
for s in vals_list:
if not s: continue
for t in str(s).split(','):
t=t.strip()
if t and t not in seen: seen.append(t)
pri=[t for t in TYPE_ORDER if t in seen]+[t for t in seen if t not in TYPE_ORDER]
return ','.join(pri) if pri else None
def mc_of(K):
m=re.search(r'menuCd=(DOM_\w+)', str(K)); return m.group(1) if m else None
def cpath(url):
return os.path.join(CACHE, hashlib.md5(str(url).encode()).hexdigest()+".html")
def html_of(K):
p=cpath(K); return open(p,encoding='utf-8').read() if os.path.exists(p) else None
def tail(mc):
m=re.search(r'(\d{18})$', mc); return m.group(1) if m else None
def parent_mc(mc):
t=tail(mc)
if not t: return None
g=[t[i:i+3] for i in range(0,18,3)]
for i in range(5,-1,-1):
if g[i]!='000':
g2=g[:]; g2[i]='000'; return mc[:mc.rindex(t)]+''.join(g2)
return None
def main():
mode=sys.argv[1] if len(sys.argv)>1 else 'dry'
wb=openpyxl.load_workbook(XP); ws=wb.active
rows=dd.load_flat(ws) # src 기준
by_src={r['src']:r for r in rows}
# menuCd 매핑 (src 기준)
src_mc={}; mc_src={}
for r in rows:
mc=mc_of(r['vals'].get(11))
if mc: src_mc[r['src']]=mc; mc_src.setdefault(mc,r['src'])
# 판별
det={}
for r in rows:
s=r['src']; K=r['vals'].get(11)
if s in EXTERNAL or s not in src_mc:
det[s]=None; continue
h=html_of(K)
det[s]='게시판' if (h and LIST.search(h)) else '페이지'
# 탭그룹 발굴 → parent별 distinct 탭 menuCd
plan={}
for r in rows:
mc=src_mc.get(r['src']);
if not mc: continue
h=html_of(r['vals'].get(11))
if not h: continue
m=TAB_DIV.search(h)
if not m: continue
tabs=re.findall(r'menuCd=(DOM_\w+)', m.group(1))
seen=set(); tabs=[t for t in tabs if not (t in seen or seen.add(t))]
if len(tabs)<2: continue
par=parent_mc(tabs[0])
plan.setdefault(par,set()).update(tabs)
# 병합 액션 산출
fold_src=set() # 흡수(삭제)될 src
survivor_M={} # src -> M(페이지탭수)
survivor_N={} # src -> 저작물유형 합집합
survivor_clear_leaf=set() # 라벨 leaf 컬럼 비울 survivor
actions=[]
warn_data=[]; warn_new=[]
for par, tabset in plan.items():
tab_srcs=[mc_src[t] for t in tabset if t in mc_src]
if len(tab_srcs)<2: continue
# 폴드 가능한 페이지 탭 = det 페이지 & 외부아님
page_tabs=[s for s in tab_srcs if det.get(s)=='페이지' and s not in EXTERNAL]
board_tabs=[s for s in tab_srcs if not (det.get(s)=='페이지' and s not in EXTERNAL)]
if len(page_tabs)<1:
continue
par_src=mc_src.get(par)
if par_src and det.get(par_src)=='페이지' and par_src not in tab_srcs:
survivor=par_src; members=page_tabs
else:
survivor=min(page_tabs); members=[s for s in page_tabs if s!=survivor]
survivor_clear_leaf.add(survivor)
if len(members)<1:
continue
for s in members: fold_src.add(s)
survivor_M[survivor]=len(page_tabs)
survivor_N[survivor]=union_types([by_src[survivor]['vals'].get(14)]+[by_src[s]['vals'].get(14) for s in members])
actions.append((survivor, members, board_tabs, len(page_tabs)))
# 경고: 흡수행에 KOGL/저작물 데이터(N=14,O=15,..S=19) 있나
for s in members:
v=by_src[s]['vals']
if any(v.get(c) not in (None,'') for c in (15,16,17,18)): # O~R 공공누리
warn_data.append((s, v.get(7) or v.get(6), {c:v.get(c) for c in (14,15,19)}))
lab=' > '.join(str(v.get(c)) for c in range(4,11) if v.get(c))
if '새창열림' in lab:
warn_new.append((s,lab))
# 새 행 구성
out=[]
reclass={'게시판→페이지':0,'페이지→게시판(미적용/외부)':0,'유지':0}
for r in rows:
s=r['src']
if s in fold_src:
continue
v=r['vals']
oldL=v.get(12)
d=det.get(s)
# 재분류 (외부/비menuCd 제외)
if d in ('게시판','페이지'):
if oldL!=d and not (oldL=='페이지' and d=='게시판'):
# 페이지→게시판은 외부 제외했으므로 여기 오면 진짜 게시판
pass
newL=d
if oldL=='페이지' and d=='게시판' and s in EXTERNAL:
newL=oldL
# 통계
if oldL=='게시판' and newL=='페이지': reclass['게시판→페이지']+=1
v[12]=newL
# M 설정
if newL=='페이지':
v[13]=survivor_M.get(s, 1)
# 게시판이면 기존 M 유지
# survivor: 저작물유형 합집합 갱신
if s in survivor_N and survivor_N[s]:
v[14]=survivor_N[s]
# survivor 라벨 정리(탭이 survivor가 된 경우 leaf 비움)
if s in survivor_clear_leaf:
lc=dd.leaf_depth(v)
v[lc]=None
out.append(r)
print(f"=== 김제시 재분류 dry ===")
print(f"현재행 {len(rows)} → 병합후 {len(out)} (흡수 {len(fold_src)}행)")
print(f"게시판→페이지 재분류: {reclass['게시판→페이지']}")
print(f"병합 그룹 수: {len(actions)}")
print(f"\n[경고] 흡수행 중 공공누리(O~R) 데이터 보유: {len(warn_data)}")
for s,nm,d2 in warn_data: print(f" r{s} {nm}: {d2}")
print(f"\n[확인] 흡수 대상 '새창열림' 페이지: {len(warn_new)}")
for s,lab in warn_new[:20]: print(f" r{s}: {lab}")
# 병합 그룹 요약
print("\n=== 병합 그룹 (survivor M=페이지탭수, 게시판탭은 유지) ===")
def lab(s):
v=by_src[s]['vals']; return ' > '.join(str(v.get(c)) for c in range(4,11) if v.get(c))
for survivor,members,boards,m in sorted(actions):
print(f" survivor r{survivor} M={m} [{lab(survivor)}] 흡수 {len(members)}" + (f", 게시판유지 {len(boards)}" if boards else ""))
if mode=='run':
bak=XP.replace('.xlsx','_backup_재분류전.xlsx')
shutil.copy(XP,bak)
dd.write_back(ws,out)
try: wb.save(XP); print(f"\n저장완료. 백업:{os.path.basename(bak)}")
except PermissionError: wb.save(XP.replace('.xlsx','_LP.xlsx')); print("\n!! 잠김→_LP")
if __name__=='__main__':
main()

93
_김제_정리.py Normal file
View File

@ -0,0 +1,93 @@
# -*- coding: utf-8 -*-
"""김제시 후처리 2규칙:
R1) F가 병합(같은 D,E,F 연속 2)인데 블록 '첫 행' G가 비어있으면 삭제(부모 랜딩 ).
R2) F 또는 G 텍스트 끝이 '새창열림'이면 L=사이트, M·N·O·P·Q 삭제, 텍스트에서 '새창열림' 제거.
사용: python -X utf8 _김제_정리.py plan | run
"""
import os, sys, re, io, shutil, importlib.util, warnings
warnings.filterwarnings('ignore')
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8', line_buffering=True)
import openpyxl
HERE = os.path.dirname(os.path.abspath(__file__))
def _imp(n, p):
s = importlib.util.spec_from_file_location(n, p); m = importlib.util.module_from_spec(s); s.loader.exec_module(m); return m
dd = _imp('dd', os.path.join(HERE, '_스크립트', '_dedup_all.py'))
XP = os.path.join(HERE, '작업파일', '광역_사이트맵', '전북특별자치도', '3.김제시', '전북특별자치도_김제시.xlsx')
NEWWIN = re.compile(r'\s*새\s*창\s*열림\s*$')
def lab(v):
return ' > '.join(str(v.get(c)) for c in range(4, 11) if v.get(c) not in (None, ''))
def main():
mode = sys.argv[1] if len(sys.argv) > 1 else 'plan'
wb = openpyxl.load_workbook(XP); ws = wb.active
rows = dd.load_flat(ws)
# --- R2: 새창열림 — 행의 leaf(가장 깊은 카테고리) 라벨 끝이 새창열림이면 사이트화 ---
r2 = []
for r in rows:
v = r['vals']
lc = dd.leaf_depth(v) # 가장 깊은 카테고리 컬럼(E/F/G…)
t = v.get(lc)
if isinstance(t, str) and NEWWIN.search(t):
r2.append(r)
# 적용: 사이트화 + M/N/O/P/Q 삭제. 텍스트 strip 은 전체 행·전 카테고리 컬럼에서.
for r in r2:
v = r['vals']
r['_had_pq'] = any(v.get(c) not in (None, '') for c in (16, 17))
v[12] = '사이트' # L
for c in (13, 14, 15, 16, 17): # M N O P Q
v[c] = None
# 새창열림 텍스트 제거(모든 행, D~J)
for r in rows:
v = r['vals']
for c in range(4, 11):
t = v.get(c)
if isinstance(t, str) and NEWWIN.search(t):
v[c] = NEWWIN.sub('', t).strip()
# --- R1: F-block 첫행 G 비면 삭제 ---
# F-block = 같은 (D,E,F) 연속 run, F 비어있지 않음, run길이>=2
delete = set() # id(row)
n = len(rows)
i = 0
while i < n:
v = rows[i]['vals']
F = v.get(6)
if F in (None, ''):
i += 1; continue
j = i
while j + 1 < n and rows[j+1]['vals'].get(6) == F \
and rows[j+1]['vals'].get(4) == v.get(4) \
and rows[j+1]['vals'].get(5) == v.get(5):
j += 1
runlen = j - i + 1
if runlen >= 2 and rows[i]['vals'].get(7) in (None, ''):
delete.add(id(rows[i]))
i = j + 1
del_rows = [r for r in rows if id(r) in delete]
keep = [r for r in rows if id(r) not in delete]
print('=== 김제 정리 plan ===')
print(f'기존행 {len(rows)}')
print(f'\n[R2] 새창열림 → 사이트: {len(r2)}건 (그중 P/Q 보유 {sum(1 for r in r2 if r.get("_had_pq"))}건)')
for r in r2[:40]:
print(f' r{r["src"]}: {lab(r["vals"])}')
print(f'\n[R1] F병합 첫행 G빈칸 → 삭제: {len(del_rows)}')
for r in del_rows[:60]:
print(f' r{r["src"]}: {lab(r["vals"])} (K={str(r["vals"].get(11))[-40:]})')
print(f'\n결과행: {len(rows)}{len(keep)}')
if mode != 'run':
return
bak = XP.replace('.xlsx', '_backup_정리전.xlsx')
shutil.copy(XP, bak)
dd.write_back(ws, keep)
try:
wb.save(XP); print(f'\n저장완료 {len(rows)}{len(keep)}행. 백업 {os.path.basename(bak)}')
except PermissionError:
alt = XP.replace('.xlsx', '_LP.xlsx'); wb.save(alt); print(f'\n!! 잠김 → {os.path.basename(alt)}')
if __name__ == '__main__':
main()

299
_김제_확장.py Normal file
View File

@ -0,0 +1,299 @@
# -*- coding: utf-8 -*-
"""김제시 누락 하위메뉴 확장 — 공식 사이트맵(전체메뉴) 기준.
basic_tab 부모병합 등으로 빠진 말단 하위메뉴(각자 menuCd 보유) 사이트맵 트리에서
복원해 부모 아래 올바른 컬럼(D=4+tree_depth)·위치로 삽입한다. 부모행(병합 survivor,
M=탭수) 자식이 분리되므로 자기 페이지로 Phase2~4 재수집(M=1 ). 기존 행은 보존.
L 판별은 김제 방식(본문 리스트클래스=게시판) + 미디어/KOGL은 전북 phase234 로직 재사용.
사용: python -X utf8 _김제_확장.py plan | run
"""
import os, sys, re, io, time, importlib.util, warnings
from urllib.parse import urljoin
from copy import copy
from concurrent.futures import ThreadPoolExecutor, as_completed
warnings.filterwarnings('ignore')
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8', line_buffering=True)
import requests
from bs4 import BeautifulSoup
import openpyxl
HERE = os.path.dirname(os.path.abspath(__file__))
def _imp(name, path):
spec = importlib.util.spec_from_file_location(name, path)
m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m); return m
dd = _imp('dd', os.path.join(HERE, '_스크립트', '_dedup_all.py'))
ph = _imp('ph', os.path.join(HERE, '_스크립트', '_jeonbuk_phase234_all.py'))
XP = os.path.join(HERE, '작업파일', '광역_사이트맵', '전북특별자치도', '3.김제시', '전북특별자치도_김제시.xlsx')
BASE = 'https://www.gimje.go.kr'
ALLMENU = BASE + '/index.gimje?menuCd=DOM_000000107002000000'
LIST = re.compile(r'class="[^"]*(bbs_list|news_list|photo_list|video_list|magazine_list|gallery_list|board_list)[^"]*"')
NEWWIN = re.compile(r'\s*새\s*창\s*열림\s*$')
UA = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/120 Safari/537.36'}
def mc_of(s):
m = re.search(r'menuCd=(DOM_\w+)', str(s) or ''); return m.group(1) if m else None
def parse_tree():
"""반환: nodes(dict mc->{label,url,depth,parent,children[]}), order(list of mc in DFS)."""
r = requests.get(ALLMENU, headers=UA, verify=False, timeout=20)
soup = BeautifulSoup(r.content, 'html.parser')
smap = soup.select_one('div.sitemap')
nodes = {}; order = []
def add(mc, label, url, depth, parent):
label = NEWWIN.sub('', label).strip()
if mc in nodes:
return
nodes[mc] = {'label': label, 'url': url, 'depth': depth,
'parent': parent, 'children': []}
order.append(mc)
if parent and parent in nodes:
nodes[parent]['children'].append(mc)
def walk(ul, depth, parent):
for li in ul.find_all('li', recursive=False):
a = li.find('a', recursive=False) or li.find('a')
if not a:
continue
mc = mc_of(a.get('href'))
if not mc:
continue
label = a.get_text(' ', strip=True)
url = urljoin(BASE, a.get('href'))
add(mc, label, url, depth, parent)
sub = li.find('ul', recursive=False)
if sub:
walk(sub, depth + 1, mc)
for mdiv in smap.find_all('div', recursive=False):
head = mdiv.find(['h2', 'h3', 'strong', 'a'])
if not head:
continue
# 대분류 head 의 menuCd (a면) 아니면 라벨만 (컨테이너) — 트리 노드로 등록(depth0)
hmc = mc_of(head.get('href')) if head.name == 'a' else None
htxt = head.get_text(' ', strip=True)
if not hmc:
# 컨테이너 라벨용 가짜 mc
hmc = 'CAT_' + str(len(nodes))
add(hmc, htxt, urljoin(BASE, head.get('href')) if head.name == 'a' and head.get('href') else '', 0, None)
topul = mdiv.find('ul')
if topul:
walk(topul, 1, hmc)
return nodes, order
def subtree(nodes, mc):
out = set()
stack = [mc]
while stack:
x = stack.pop(); out.add(x)
stack.extend(nodes[x]['children'])
return out
def path_labels(nodes, mc):
"""mc 의 조상→자기 라벨 리스트 (depth 순)."""
chain = []
cur = mc
while cur is not None:
chain.append(cur)
cur = nodes[cur]['parent']
chain.reverse()
return chain # list of mc by depth
# ---- Phase 2~4 (김제 list-class L + 전북 미디어/KOGL) ----
def collect(session, url):
out = {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
try:
r = session.get(url, headers=UA, verify=False, timeout=8, allow_redirects=True)
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
r.encoding = meta.group(1).decode('ascii', 'ignore') if meta else r.apparent_encoding
if r.status_code != 200:
out['note'] = '접근 실패'; return out
html = r.text
except Exception:
out['note'] = '접근 실패'; return out
soup = BeautifulSoup(html, 'html.parser')
body = ph.get_body(soup, ph.BODY_SEL)
is_board = bool(LIST.search(html))
if is_board:
out['L'] = '게시판'
_, cnt = ph.detect_form(body)
out['M'] = cnt
else:
out['L'] = '페이지'; out['M'] = 1
has_img, has_vid, has_txt = ph.detect_media(body)
types, q = ph.detect_kogl(body)
if is_board:
for du in ph.extract_detail_urls(body, url, limit=2):
try:
dr = session.get(du, headers=UA, verify=False, timeout=8)
ds = BeautifulSoup(dr.content, 'html.parser')
db = ph.get_body(ds, ph.BODY_SEL)
di, dv, dt = ph.detect_media(db)
has_img |= di; has_vid |= dv; has_txt |= dt
dt_types, dt_q = ph.detect_kogl(db)
if dt_types and not types:
out['P'] = '게시물'
types |= dt_types
if dt_q == 'Y':
q = 'Y'
except Exception:
pass
out['N'] = ph.n_string(has_txt, has_img, has_vid)
if types:
out['O'] = ','.join(f'{n}유형' for n in sorted(types))
out['P'] = out['P'] or '게시판'
out['Q'] = q or 'N'
else:
out['O'] = '미부착'
return out
def main():
mode = sys.argv[1] if len(sys.argv) > 1 else 'plan'
nodes, order = parse_tree()
real = {mc: n for mc, n in nodes.items() if not mc.startswith('CAT_')}
leaves = [mc for mc in order if not mc.startswith('CAT_') and not nodes[mc]['children']]
wb = openpyxl.load_workbook(XP); ws = wb.active
rows = dd.load_flat(ws)
pos = {}; row_by_mc = {}
for i, r in enumerate(rows):
mc = mc_of(r['vals'].get(11))
if mc:
pos[mc] = i; row_by_mc[mc] = r
have = set(row_by_mc)
missing = [mc for mc in leaves if mc not in have]
# parents that will gain children
gain_parents = {}
for mc in missing:
p = nodes[mc]['parent']
gain_parents.setdefault(p, []).append(mc)
# anchor index for each parent group: after last existing row in subtree of parent
inserts = {} # anchor_idx -> [mc,...] in tree order
no_anchor = []
for p, kids in gain_parents.items():
# nearest existing ancestor (incl parent) to anchor after its subtree
anc = p
anchor_idx = None
while anc is not None:
st = subtree(nodes, anc)
present = [pos[m] for m in st if m in pos]
if present:
anchor_idx = max(present); break
anc = nodes[anc]['parent']
if anchor_idx is None:
no_anchor.append((p, kids)); continue
# order kids by their order in tree (children list of p)
ordered = [m for m in nodes[p]['children'] if m in kids]
inserts.setdefault(anchor_idx, []).extend(ordered)
survivors = [p for p in gain_parents if p in have]
print('=== 김제 확장 plan ===')
print(f'트리 노드(실): {len(real)} 말단leaf: {len(leaves)} 기존행: {len(rows)}')
print(f'누락 말단메뉴(추가대상): {len(missing)}')
print(f'자식 얻는 부모: {len(gain_parents)} (그중 기존행=재수집대상 survivor: {len(survivors)})')
print(f'앵커 못찾음: {len(no_anchor)}')
# sample
def lab(mc):
return ' > '.join(nodes[m]['label'] for m in path_labels(nodes, mc) if not m.startswith('CAT_') or nodes[m]['label'])
print('\n[샘플 추가 행 20]')
for mc in missing[:20]:
d = nodes[mc]['depth']; col = chr(ord('A') + 3 + d)
print(f' +{col}({d}) {lab(mc)} ({mc})')
print('\n[survivor 부모 재수집 대상]')
for p in survivors:
r = row_by_mc[p]; v = r['vals']
print(f' r{r["src"]} M={v.get(13)} L={v.get(12)} [{nodes[p]["label"]}] 자식 {len(gain_parents[p])}')
if no_anchor:
print('\n[!] 앵커 못찾은 그룹:')
for p, kids in no_anchor:
print(f' parent {p} kids {len(kids)}')
if mode != 'run':
return
# ---- build new rows + fetch ----
session = requests.Session()
template = row_by_mc.get(list(survivors)[0]) if survivors else rows[0]
tmpl_styles = template['styles']
def make_row(mc):
v = {c: None for c in range(1, dd.MAXCOL + 1)}
v[3] = '김제시'
chain = path_labels(nodes, mc)
for m in chain:
d = nodes[m]['depth']
col = 4 + d
if col <= 10 and nodes[m]['label']:
v[col] = nodes[m]['label']
v[11] = nodes[mc]['url']
return {'src': None, 'vals': v,
'styles': {c: tuple(copy(x) if hasattr(x, 'copy') or True else x for x in tmpl_styles[c]) for c in tmpl_styles},
'hyperlink': nodes[mc]['url']}
# fetch all missing + survivor parents
fetch_targets = list(missing) + survivors
print(f'\nPhase2~4 수집 {len(fetch_targets)}건 ...')
res = {}
t0 = time.time()
def work(mc):
return mc, collect(session, nodes[mc]['url'])
with ThreadPoolExecutor(max_workers=10) as ex:
futs = [ex.submit(work, mc) for mc in fetch_targets]
done = 0
for f in as_completed(futs):
mc, o = f.result(); res[mc] = o; done += 1
if done % 30 == 0 or done == len(fetch_targets):
print(f' {done}/{len(fetch_targets)} ({time.time()-t0:.0f}s)')
# apply to survivors (overwrite own page data; M back to own)
for p in survivors:
v = row_by_mc[p]['vals']; o = res.get(p, {})
for col, key in [(12, 'L'), (13, 'M'), (14, 'N'), (15, 'O'), (16, 'P'), (17, 'Q')]:
if o.get(key) != '':
v[col] = o[key]
# build new row objects with phase data
new_obj = {}
for mc in missing:
ro = make_row(mc); o = res.get(mc, {})
v = ro['vals']
for col, key in [(12, 'L'), (13, 'M'), (14, 'N'), (15, 'O'), (16, 'P'), (17, 'Q')]:
if o.get(key) != '':
v[col] = o[key]
if o.get('note'):
v[19] = o['note']
new_obj[mc] = ro
# assemble final ordered list
out = []
for i, r in enumerate(rows):
out.append(r)
if i in inserts:
for mc in inserts[i]:
out.append(new_obj[mc])
# backup + write
import shutil
bak = XP.replace('.xlsx', '_backup_확장전.xlsx')
shutil.copy(XP, bak)
dd.write_back(ws, out)
try:
wb.save(XP)
print(f'\n저장완료 {len(rows)}{len(out)}행. 백업 {os.path.basename(bak)}')
except PermissionError:
alt = XP.replace('.xlsx', '_LP.xlsx'); wb.save(alt)
print(f'\n!! 원본 잠김 → {os.path.basename(alt)} 로 저장')
if __name__ == '__main__':
main()

53
_스크립트/_br_scan.py Normal file
View File

@ -0,0 +1,53 @@
# -*- coding: utf-8 -*-
"""보령시 210행~끝: E병합 그룹별 F자식 형태/수량 스캔. 합치기 대상 판정."""
import openpyxl
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\6.보령시\충청남도_보령시.xlsx'
ws = openpyxl.load_workbook(XLSX).active
last = 3
for r in range(3, ws.max_row + 1):
if ws.cell(r, 11).value:
last = r
print('마지막 데이터행', last)
# E 병합 범위(>=210)
emerges = {}
for mr in ws.merged_cells.ranges:
if mr.min_col == 5 and mr.max_row >= 210:
emerges[mr.min_row] = (mr.min_row, mr.max_row, ws.cell(mr.min_row, 5).value)
# 210부터 행 훑기
r = 210
elig = noelig = 0
while r <= last:
e = ws.cell(r, 5).value
# E 병합?
rng = None
for mr in ws.merged_cells.ranges:
if mr.min_col == 5 and mr.min_row <= r <= mr.max_row:
rng = (mr.min_row, mr.max_row); break
if rng and rng[1] > rng[0]:
r1, r2 = rng
elabel = ws.cell(r1, 5).value
dlabel = ws.cell(r1, 4).value
ls = [ws.cell(x, 12).value for x in range(r1, r2 + 1)]
ms = [ws.cell(x, 13).value for x in range(r1, r2 + 1)]
fs = [ws.cell(x, 6).value for x in range(r1, r2 + 1)]
allpage = all((l in ('페이지', '사이트')) for l in ls)
kinds = set(ls)
tag = 'O합치기' if allpage else 'X(게시판섞임)'
if allpage: elig += 1
else: noelig += 1
try:
ssum = sum(int(m) for m in ms if m is not None)
except Exception:
ssum = '?'
print('E[%d~%d] %s > %s (%d행) L=%s M합=%s %s' % (r1, r2, dlabel, elabel, r2 - r1 + 1, kinds, ssum, tag))
if not allpage:
for x in range(r1, r2 + 1):
print(' - %s L=%s M=%s' % (ws.cell(x, 6).value, ws.cell(x, 12).value, ws.cell(x, 13).value))
r = r2 + 1
else:
# 단일행(E 비병합) 또는 E없음
r += 1
print('\n합치기대상 E그룹:', elig, '/ 제외(게시판섞임):', noelig)

View File

@ -0,0 +1,140 @@
# -*- coding: utf-8 -*-
"""부여군 N(저작물 유형) 재판정 — 장식/공통UI 이미지 false '이미지' 제거.
- phase234 detect_media KOGL 마크만 제외 file_icon.gif·see_btn.gif·btn_page_*.gif
게시판 스킨/공통버튼 아이콘을 '이미지' 오탐.
- [[feedback_N_image_rule]] 적용: 장식/공통UI·페이징버튼·첨부아이콘 제외, 실제 사진/업로드·지도/PDF임베드만 인정.
- 대상: 현재 N '이미지' 포함된 같은도메인 행만(false 양성 제거 전용 이미지 추가 ).
- 게시판은 상세글 상위5 추적. body=#txt|#contents|main.
사용: python -X utf8 _buyeo_rejudge_N.py [--write]
"""
import sys, warnings, re
from urllib.parse import urljoin
import requests, openpyxl
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
S = requests.Session(); S.headers.update({'User-Agent': 'Mozilla/5.0'})
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\7.부여군\충청남도_부여군.xlsx'
DOMAIN = 'buyeo.go.kr'
KOGL = re.compile(r'(?:new_)?img_open(?:type|code)\d', re.I)
# 장식/공통UI: 기존 논산 DECO + 부여 게시판 스킨/공통버튼/첨부아이콘 패턴 보강
DECO = re.compile(
r'/site/common/img/|move\.png|no[-_]?img|blank\.|spacer\.|'
r'/img/(?:icon|ico|bul|bullet|arrow|btn|bg|tit|h\d)|'
r'ico_|btn_|bul_|bg_|_bg\.|icon_|'
r'file_icon|_icon\.|icon\.gif|_btn\.|see_btn|' # 첨부아이콘·바로보기버튼
r'/skin/|/images/(?:kr/)?common/|/common/img/|' # 게시판 스킨·공통 UI 디렉터리
r'btn_page|page_(?:next|prev|first|last)', re.I) # 페이징 버튼
MAPPDF = re.compile(r'pdf|viewer\.html|map|kakao|daum|naver.*map|google.*map|/map', re.I)
VIDEO = re.compile(r'youtube\.com|youtu\.be|/embed/|\.mp4|\.webm|vimeo', re.I)
DETAIL = re.compile(r'mode=V|view\.do|/view|seq=|idx=|mng_no=|bbsSeq|nttId|no=', re.I)
def get(url):
r = S.get(url, verify=False, timeout=20)
return BeautifulSoup(r.content, 'html.parser')
def body(soup):
return soup.select_one('#txt') or soup.select_one('#contents') or soup.select_one('main') or soup
def real_imgs(node):
n = 0
for im in node.find_all('img'):
src = im.get('src') or ''
if not src or KOGL.search(src) or DECO.search(src):
continue
n += 1
return n
def media_flags(node):
img = real_imgs(node) > 0
vid = False
for ifr in node.find_all('iframe'):
s = ifr.get('src') or ''
if VIDEO.search(s):
vid = True
elif MAPPDF.search(s):
img = True
for a in node.find_all('a'):
if VIDEO.search(a.get('href') or ''):
vid = True
if node.find('video'):
vid = True
return img, vid
def detail_links(soup, base):
out = []
for a in body(soup).find_all('a'):
h = a.get('href') or ''
if h.startswith('#'):
continue
if DETAIL.search(h):
out.append(urljoin(base, h))
seen = set(); res = []
for u in out:
if u not in seen:
seen.add(u); res.append(u)
return res[:5]
def judge(k, L, M, N):
if not isinstance(k, str) or DOMAIN not in k:
return N, 'skip(외부)'
try:
soup = get(k)
except Exception as e:
return N, f'fetch실패:{str(e)[:25]}'
bd = body(soup)
txtlen = len(bd.get_text(strip=True))
if L == '게시판' and (M in (0, '0', None)):
return '없음', 'board0'
has_img, has_vid = media_flags(bd)
detail = False
if L == '게시판' and not has_img:
for du in detail_links(soup, k):
try:
di, dv = media_flags(body(get(du)))
if di: has_img = True
if dv: has_vid = True
if di or dv: detail = True
if has_img and has_vid: break
except Exception:
pass
parts = []
if txtlen >= 1: parts.append('어문')
if has_img: parts.append('이미지')
if has_vid: parts.append('영상')
return (','.join(parts) if parts else '없음'), ('detail' if detail else '')
def main():
write = '--write' in sys.argv
wb = openpyxl.load_workbook(XLSX); ws = wb.active
changes = []
for r in range(3, ws.max_row + 1):
N = ws.cell(r, 14).value
if not (isinstance(N, str) and '이미지' in N):
continue
k = ws.cell(r, 11).value; L = ws.cell(r, 12).value; M = ws.cell(r, 13).value
new, tag = judge(k, L, M, N)
if new != N:
changes.append((r, N, new, tag, ws.cell(r, 7).value or ws.cell(r, 6).value, k))
print(f'현재 N에 이미지 포함 행 재판정 → 변경 {len(changes)}')
for r, old, new, tag, label, k in changes:
print(f'{r} [{old}{new}] {tag} | {label} | {str(k)[-45:]}')
if write and changes:
for r, old, new, tag, label, k in changes:
ws.cell(r, 14).value = new
wb.save(XLSX)
print(f'\n저장: {XLSX} ({len(changes)}건 수정)')
elif not write:
print('\n(계획만. --write 로 반영)')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,266 @@
# -*- coding: utf-8 -*-
"""부여군 N(저작물 유형) + O/P/Q(공공누리) 전수 재판정.
사용자 검수(3~37) 정답으로 검증 38~ 적용.
핵심(부여 buyeo 템플릿 특성):
- KOGL 마크는 게시판 '상세글' `div.gnuri_layer > .codeView01 > img[src=/_module/gnuri/images/img_opentypeNN.png]`
+ kogl.or.kr/info/licenseTypeN.do 링크. **본문(#txt) 바깥**이라 기존 phase234가 놓침.
- 상세 URL = `?mode=V&no=...`.
- 게시판은 글마다 유형 다를 있음 O = ** 마크 글의 유형**(사용자 컨벤션: 01030502 글1=1유형 O=1유형).
- N 이미지: 장식/공통UI(file_icon·see_btn·btn_page·/skin/·/common/·ico/btn/bg) + KOGL 마크 제외.
지도/PDF iframe = 이미지. 게시판 M=0 = 없음. 실제 업로드 사진만 이미지.
사용: python -X utf8 _buyeo_rejudge_NO.py --validate # 3~37 사용자값과 비교
python -X utf8 _buyeo_rejudge_NO.py --write # 38행~ 기입(백업)
"""
import sys, warnings, re, time
from urllib.parse import urljoin
from concurrent.futures import ThreadPoolExecutor, as_completed
import requests, openpyxl
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
S = requests.Session(); S.headers.update({'User-Agent': 'Mozilla/5.0'})
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\7.부여군\충청남도_부여군.xlsx'
DOMAIN = 'buyeo.go.kr'
REF_MAX = 37 # 3~37 = 사용자 검수 정답
KOGL_IMG = re.compile(r'(?:new_)?img_open(?:type|code)(\d{1,2})\.(?:png|jpe?g|gif)', re.I)
KOGL_LINK = re.compile(r'licenseType(\d)', re.I)
DECO = re.compile(
r'/site/common/img/|move\.png|no[-_]?img|blank\.|spacer\.|'
r'/img/(?:icon|ico|bul|bullet|arrow|btn|bg|tit|h\d)|'
r'ico_|btn_|bul_|bg_|_bg\.|icon_|'
r'file_icon|_icon\.|icon\.gif|_btn\.|see_btn|flag\.|webaccess|'
r'/skin/|/images/(?:kr/)?common(?:_new)?/|/common/img/|/_module/gnuri/|'
r'btn_page|page_(?:next|prev|first|last)', re.I)
MAPPDF = re.compile(r'pdf|viewer\.html|/map|kakao|daum|=map|naver.*map|google.*map', re.I)
VIDEO = re.compile(r'youtube\.com|youtu\.be|/embed/|\.mp4|\.webm|vimeo', re.I)
DETAIL = re.compile(r'mode=V&no=', re.I)
def get(url):
last = None
for _ in range(3):
try:
r = S.get(url, verify=False, timeout=20)
return BeautifulSoup(r.content, 'html.parser')
except Exception as e:
last = e
time.sleep(1.2)
raise last
def body(soup):
return soup.select_one('#txt') or soup.select_one('#contents') or soup.select_one('main') or soup
def real_img(node):
for im in node.find_all('img'):
src = im.get('src') or ''
if not src or KOGL_IMG.search(src) or DECO.search(src):
continue
return True
return False
def media_flags(node):
img = real_img(node)
vid = bool(node.find('video'))
for ifr in node.find_all('iframe'):
s = ifr.get('src') or ''
if VIDEO.search(s):
vid = True
elif MAPPDF.search(s):
img = True
for a in node.find_all('a'):
if VIDEO.search(a.get('href') or ''):
vid = True
return img, vid
def kogl_from(html):
"""전체 HTML에서 KOGL 이미지유형·링크유형 추출."""
imgs = [int(x) for x in KOGL_IMG.findall(html) if 1 <= int(x) <= 4]
links = [int(x) for x in KOGL_LINK.findall(html) if 1 <= int(x) <= 4]
return imgs, links
def _menu_cd(u):
m = re.search(r'menu_dvs_cd=([^&]+)', u or '')
return m.group(1) if m else None
def detail_links(soup, base):
"""같은 게시판(_prog/_board + 동일 menu_dvs_cd) 상세글만. 타 게시판 stray 링크 배제."""
if '/_prog/_board/' not in base:
return [] # 게시판 템플릿이 아니면 상세 없음(scate 등)
base_menu = _menu_cd(base)
out = []
for a in body(soup).find_all('a', href=True):
h = a['href']
if not DETAIL.search(h):
continue
full = urljoin(base, h)
if base_menu and _menu_cd(full) and _menu_cd(full) != base_menu:
continue # 다른 게시판 글 → 제외
out.append(full)
seen = set(); res = []
for u in out:
if u not in seen:
seen.add(u); res.append(u)
return res[:5]
def decide_O(img_types, link_types):
"""KOGL 규칙: 이미지유형 우선. 링크만이면 licenseType, 1·2·3·4 전부=범례→미부착."""
if img_types:
t = img_types[0]
mism = bool(link_types) and (link_types[0] != t)
return f'{t}유형', mism
if link_types:
if {1, 2, 3, 4}.issubset(set(link_types)):
return '미부착', False
return f'{link_types[0]}유형', False
return '미부착', False
def judge(k, L, M):
"""반환 dict: N, O, P, Q, S(mismatch)."""
try:
soup = get(k)
except Exception as e:
return {'err': f'fetch:{str(e)[:25]}'}
bd = body(soup); full = str(soup)
txt = len(bd.get_text(strip=True)) >= 1
is_board_url = '/_prog/_board/' in k # 실제 게시판 URL만(.html·서브사이트는 게시판 아님)
has_img, has_vid = media_flags(bd)
# O: 페이지/리스트 자체 마크 우선
img_t, link_t = kogl_from(full)
o_loc = '게시판' if L == '게시판' else '페이지'
o_imgs, o_links = list(img_t), list(link_t)
# '빈 게시판→없음'은 진짜 게시판(_prog/_board)에 글이 0개일 때만.
# L=게시판 M=0 이라도 .html 페이지·서브사이트는 본문내용으로 판정(원본 L/M 오기 보정).
dls = detail_links(soup, k) if (L == '게시판') else []
board0 = (L == '게시판' and is_board_url and (M in (0, '0', None)) and not dls)
if L == '게시판' and not board0:
# N: 상위 5글 본문에서 실제 이미지/영상 누적(이미지 콘텐츠 회수율↑)
for idx, du in enumerate(dls):
try:
ds = get(du); db = body(ds); df = str(ds)
except Exception:
continue
di, dv = media_flags(db)
if di: has_img = True
if dv: has_vid = True
# O: 리스트 마크 없으면 '첫(대표) 글'만으로 판정
# (사용자 컨벤션: 01030502 글1=1유형→1유형 / 업무추진비 글1무·글5만3유형→미부착)
if idx == 0 and not o_imgs and not o_links:
dimg, dlink = kogl_from(df)
if dimg or dlink:
o_imgs, o_links, o_loc = dimg, dlink, '게시물'
if has_img and has_vid:
break
# N
if board0:
N = '없음'
else:
parts = []
if txt: parts.append('어문')
if has_img: parts.append('이미지')
if has_vid: parts.append('영상')
N = ','.join(parts) if parts else '없음'
O, mism = decide_O(o_imgs, o_links)
P = (o_loc if O != '미부착' else None)
Q = ('Y' if (O != '미부착' and o_links) else None)
return {'N': N, 'O': O, 'P': P, 'Q': Q, 'mismatch': mism}
def main():
mode = '--validate' if '--validate' in sys.argv else ('--write' if '--write' in sys.argv else '--dry')
wb = openpyxl.load_workbook(XLSX); ws = wb.active
rows = []
for r in range(3, ws.max_row + 1):
k = ws.cell(r, 11).value; L = ws.cell(r, 12).value; M = ws.cell(r, 13).value
if not (isinstance(k, str) and DOMAIN in k):
continue
if L == '사이트':
continue
if mode == '--validate' and r > REF_MAX:
continue
if mode == '--write' and r <= REF_MAX:
continue
rows.append((r, k, L, M))
print(f'mode={mode} | 대상 {len(rows)}')
res = {}
t0 = time.time()
with ThreadPoolExecutor(max_workers=3) as ex:
futs = {ex.submit(judge, k, L, M): (r, k, L, M) for r, k, L, M in rows}
done = 0
for fut in as_completed(futs):
r, k, L, M = futs[fut]
res[r] = fut.result()
done += 1
if done % 40 == 0:
print(f' {done}/{len(rows)} ({time.time()-t0:.0f}s)')
print(f'크롤링 완료 ({time.time()-t0:.0f}s)')
if mode == '--validate':
miss = 0
for r, k, L, M in rows:
d = res[r]
if d.get('err'):
print(f'{r} ERR {d["err"]}'); continue
curN = ws.cell(r, 14).value; curO = ws.cell(r, 15).value
curP = ws.cell(r, 16).value; curQ = ws.cell(r, 17).value
dn = (d['N'] or '') != (curN or '')
do = (d['O'] or '') != (curO or '')
dp = (d['P'] or '') != (curP or '')
dq = (d['Q'] or '') != (curQ or '')
if dn or do or dp or dq:
miss += 1
g = ws.cell(r, 7).value or ws.cell(r, 6).value
print(f' ✗행{r} {str(g)[:16]:16}| N:{curN}{d["N"]} {"" if dn else "="} | '
f'O:{curO}{d["O"]}{"" if do else "="} P:{curP}{d["P"]}{"" if dp else ""} Q:{curQ}{d["Q"]}{"" if dq else ""}')
print(f'\n불일치 {miss}/{len(rows)}행 (0이면 로직이 사용자 검수와 100% 일치)')
return
# write (38행~)
if mode == '--write':
import shutil
bk = XLSX.replace('.xlsx', '_backup_NO재판정전.xlsx'); shutil.copy(XLSX, bk)
print('백업:', bk)
chg = 0
for r, k, L, M in rows:
d = res[r]
if d.get('err'):
print(f'{r} ERR {d["err"]}'); continue
old = (ws.cell(r,14).value, ws.cell(r,15).value, ws.cell(r,16).value, ws.cell(r,17).value)
ws.cell(r, 14).value = d['N']
ws.cell(r, 15).value = d['O']
ws.cell(r, 16).value = d['P']
ws.cell(r, 17).value = d['Q']
if d['mismatch']:
cur = (ws.cell(r,19).value or '').strip()
ws.cell(r,19).value = '링크주소 오기' if not cur else cur+' / 링크주소 오기'
new = (d['N'], d['O'], d['P'], d['Q'])
if old != new:
chg += 1
g = ws.cell(r,7).value or ws.cell(r,6).value
print(f'{r} {str(g)[:16]:16}| N:{old[0]}{new[0]} O:{old[1]}{new[1]} P:{old[2]}{new[2]} Q:{old[3]}{new[3]}')
wb.save(XLSX)
print(f'\n저장: {XLSX} ({chg}행 변경)')
return
print('(--validate 또는 --write 지정)')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,166 @@
# -*- coding: utf-8 -*-
"""부여군 사전정보공개(정보목록공개) fiexd_tab1 scate=1~9 분야탭 전개.
- _tab_expand 일반 필터탭 제외 규칙(같은 경로·쿼리만 다름) 걸려 누락된 케이스.
사용자 지시로 9 분야탭을 각각 별도 게시판 행으로 전개(보령시 사전정보공표 컨벤션).
- 분야 = 게시판 / M=목록 항목수(테이블 데이터행) / N=어문 / O=미부착.
- 평탄화치환재구성(D/E/F 재병합·K하이퍼링크·B순번) _tab_expand 재사용.
"""
import sys, shutil, warnings, re
from copy import copy
import requests, openpyxl
from bs4 import BeautifulSoup
from openpyxl.styles import Alignment, Font
from openpyxl.utils import get_column_letter
sys.path.insert(0, r'D:\01.프로젝트\DB수집\_스크립트')
from _tab_expand import load_flat, MAXCOL, CAT_COLS, HEADER_MERGES
warnings.filterwarnings('ignore')
H = {'User-Agent': 'Mozilla/5.0'}
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\7.부여군\충청남도_부여군.xlsx'
BASE = 'https://www.buyeo.go.kr/_prog/service/index.php?scate='
def collect():
"""9개 scate 라벨+항목수 수집."""
r = requests.get(BASE + '1', headers=H, timeout=15, verify=False)
soup = BeautifulSoup(r.content, 'html.parser')
labels = {}
for li in soup.select('.fiexd_tab1 li'):
a = li.find('a'); m = re.search(r'scate=(\d+)', a.get('href') or '')
if m:
labels[int(m.group(1))] = a.get_text(strip=True)
out = []
for sc in sorted(labels):
rr = requests.get(BASE + str(sc), headers=H, timeout=15, verify=False)
s = BeautifulSoup(rr.content, 'html.parser')
t = s.find('table')
cnt = sum(1 for tr in t.find_all('tr') if tr.find_all('td'))
out.append((sc, labels[sc], cnt))
return out
def main():
write = '--write' in sys.argv
scates = collect()
print('=== scate 분야탭 ===')
for sc, lab, cnt in scates:
print(f' scate={sc} | 항목={cnt} | {lab}')
print(f'합계 항목 {sum(c for _,_,c in scates)} / 9행 전개\n')
wb = openpyxl.load_workbook(XLSX)
ws = wb.active
rows = load_flat(ws)
# 대상 행 식별: K가 scate=1
tgt = None
for row in rows:
k = row['vals'].get(11)
if isinstance(k, str) and 'service/index.php?scate=1' in k:
tgt = row; break
if tgt is None:
print('!! 대상(scate=1) 행 없음'); return
lc = 6 # leaf=F (사전정보공개)
childc = lc + 1 # G
print(f'대상 sheet행 {tgt["src"]} (F={tgt["vals"].get(6)}) → 자식 {get_column_letter(childc)} 에 9분야 전개')
if not write:
print('\n(계획만. --write 로 기입)')
return
backup = XLSX.replace('.xlsx', '_backup_scate전.xlsx')
shutil.copy(XLSX, backup)
print('백업:', backup)
out_rows = []
for row in rows:
if row is tgt:
for sc, lab, cnt in scates:
nv = dict(tgt['vals'])
for c in CAT_COLS:
if c > lc:
nv[c] = None
nv[childc] = lab
nv[11] = BASE + str(sc)
nv[12] = '게시판'
nv[13] = cnt
nv[14] = '어문'
nv[15] = '미부착'
for c in (16, 17, 18): # P,Q,R 비움
nv[c] = None
# 신규 행(scate>1)은 비고~ 비움, scate=1은 기존 보존
if sc != 1:
for c in range(19, MAXCOL + 1):
nv[c] = None
out_rows.append({'vals': nv, 'style': tgt, 'url': nv[11]})
else:
out_rows.append({'vals': dict(row['vals']), 'style': row, 'url': row['vals'].get(11)})
# 데이터 클리어
for r in range(3, ws.max_row + 1):
for c in range(1, MAXCOL + 1):
ws.cell(r, c).value = None
ws.cell(r, c).hyperlink = None
START = 3
for i, orow in enumerate(out_rows):
r = START + i
sty = orow['style']['styles']
for c in range(1, MAXCOL + 1):
cell = ws.cell(r, c)
f, fl, bd, al, nf, pr = sty[c]
cell.font = copy(f); cell.fill = copy(fl); cell.border = copy(bd)
cell.alignment = copy(al); cell.number_format = nf; cell.protection = copy(pr)
v = orow['vals']
for c in range(1, MAXCOL + 1):
ws.cell(r, c).value = v.get(c)
ws.cell(r, 2).value = i + 1
END = START + len(out_rows) - 1
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
def merge_runs(col_letter, col_idx, group_cols=()):
cur_val = ws.cell(START, col_idx).value
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
run_start = START; runs = []
for r in range(START + 1, END + 1):
val = ws.cell(r, col_idx).value
grp = tuple(ws.cell(r, gg).value for gg in group_cols)
if val == cur_val and grp == cur_grp:
continue
if cur_val not in (None, '') and r - 1 > run_start:
runs.append((run_start, r - 1))
cur_val, cur_grp, run_start = val, grp, r
if cur_val not in (None, '') and END > run_start:
runs.append((run_start, END))
for s, e in runs:
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
ws.cell(s, col_idx).alignment = center
merge_runs('F', 6, group_cols=(4, 5))
merge_runs('E', 5, group_cols=(4,))
merge_runs('D', 4)
for r in range(START, END + 1):
for c in (4, 5, 6, 7):
if ws.cell(r, c).value is not None:
ws.cell(r, c).alignment = center
for r in range(1, END + 1):
ws.row_dimensions[r].height = 15
for r in range(START, END + 1):
cell = ws.cell(r, 11)
u = cell.value
if isinstance(u, str) and u.startswith(('http://', 'https://')):
cell.hyperlink = u
old = cell.font
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
bold=old.bold, italic=old.italic, color='0000FF', underline='single')
cell.alignment = left
wb.save(XLSX)
print(f'저장 완료: {XLSX} (총 {len(out_rows)}행)')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,660 @@
"""충청북도 11개 시·군 Phase 1 일괄 처리.
매뉴얼: D:\\01.프로젝트\\DB수집\\사이트맵_수집_매뉴얼.md
출력: 폴더의 {기관명}.xlsx
"""
import os
import re
import shutil
import ssl
import sys
import warnings
from copy import copy
from urllib.parse import urljoin
import openpyxl
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.ssl_ import create_urllib3_context
from bs4 import BeautifulSoup
from openpyxl.styles import Alignment, Font
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
TEMPLATE = r'D:\01.프로젝트\DB수집\자료_취합_예시.xlsx'
class WeakSSLAdapter(HTTPAdapter):
"""레거시 SSL 핸드셰이크 허용 (영동군 등)."""
def init_poolmanager(self, *args, **kwargs):
ctx = create_urllib3_context()
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
ctx.options |= 0x4 # OP_LEGACY_SERVER_CONNECT
ctx.check_hostname = False
ctx.verify_mode = ssl.CERT_NONE
kwargs['ssl_context'] = ctx
return super().init_poolmanager(*args, **kwargs)
def make_session(weak_ssl=False):
s = requests.Session()
s.headers.update(H)
if weak_ssl:
s.mount('https://', WeakSSLAdapter())
return s
def fetch_html(url, session=None, timeout=20, force_encoding=None):
s = session or requests.Session()
if not session:
s.headers.update(H)
r = s.get(url, timeout=timeout, verify=False, allow_redirects=True)
if force_encoding:
r.encoding = force_encoding
else:
# 메타 태그에서 charset 시도 → 없으면 apparent_encoding
meta_charset = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
if meta_charset:
r.encoding = meta_charset.group(1).decode('ascii', errors='ignore')
else:
r.encoding = r.apparent_encoding
return r.text
def clean_text(s):
return re.sub(r'\s+', ' ', s or '').strip().replace('\xa0', '')
def extract_href(a):
if a is None:
return ''
href = (a.get('href') or '').strip()
if not href or href.startswith('#') or href.lower().startswith('javascript:'):
return ''
return href
ENCODE_URI_PAT = re.compile(r"""(?:location\.href|window\.location(?:\.href)?)\s*=\s*encodeURI\(\s*['"]([^'"]+)['"]\s*\)""")
ONCLICK_HREF_PAT = re.compile(r"""(?:location\.href|window\.open)\s*\(?\s*['"]([^'"]+)['"]""")
def extract_href_with_onclick(a):
if a is None:
return ''
href = (a.get('href') or '').strip()
if href and not href.startswith('#') and not href.lower().startswith('javascript:'):
return href
onclick = a.get('onclick', '')
if onclick:
m = ENCODE_URI_PAT.search(onclick)
if m:
return m.group(1)
m = ONCLICK_HREF_PAT.search(onclick)
if m:
return m.group(1)
return ''
# ================================================================
# 파서들
# ================================================================
def parse_depth1_chungbuk(soup, base):
"""충북 e-Gov depth1 형: div.depth1 > ul.depth1_list > li.depth1_item > a.depth1_text + div.depth2 > [div.depth2_content|depth2_wrap] > ul.depth2_list."""
rows = []
container = soup.select_one('div.depth.depth1, div.depth1')
if not container:
return rows
top_ul = container.find('ul', class_=re.compile(r'depth1?_list'), recursive=False)
if not top_ul:
return rows
def walk_depth(ul, level, base_path, out):
for li in ul.find_all('li', recursive=False):
a = li.find('a', class_=re.compile(rf'depth{level}_text'), recursive=False) or li.find('a', recursive=False)
if not a:
continue
text = clean_text(a.get_text())
href = extract_href(a)
path = base_path + [(text, href)]
# Find next depth container — 다양한 wrapper 클래스 지원: _content / _wrap / 없음
next_div = li.find('div', class_=re.compile(rf'depth\s+depth{level+1}\b'), recursive=False) or \
li.find('div', class_=re.compile(rf'\bdepth{level+1}\b'), recursive=False)
if next_div:
# wrapper 가 있을 수도, 없을 수도
next_ul = next_div.find('ul', class_=re.compile(rf'depth{level+1}_list'), recursive=True)
if next_ul:
# 같은 깊이의 첫번째 ul.depth_list (자손 검색이지만 보통 1번 wrapper 안)
out.append({'path': list(path), 'href': href})
walk_depth(next_ul, level + 1, path, out)
continue
out.append({'path': list(path), 'href': href})
tmp = []
walk_depth(top_ul, 1, [], tmp)
cols = 'DEFGHIJ'
for item in tmp:
p = item['path']
row = {c: '' for c in 'DEFGHIJ'}
row['href'] = item['href']
for i, (t, _) in enumerate(p):
col = cols[i] if i < len(cols) else 'J'
row[col] = t
rows.append(row)
return rows
def parse_danyang_ld(soup, base):
"""단양군: ul#menu_sitemap.ld1 > li.cd1 > a.l1 + div.lb1 > ul.ld2 > li.cd2 > a.l2 + div.lb2 > ul.ld3 > ...
'menutype_empty' (메인 ) 스킵.
"""
rows = []
container = soup.select_one('ul#menu_sitemap')
if not container:
return rows
def walk(ul, level, base_path, out):
for li in ul.find_all('li', recursive=False):
a = li.find('a', class_=re.compile(rf'\bl{level}\b'), recursive=False) or li.find('a', recursive=False)
if not a:
continue
a_cls = ' '.join(a.get('class', []))
if 'menutype_empty' in a_cls:
continue # 메인 같은 빈 항목
text = clean_text(a.get_text())
href = extract_href(a)
path = base_path + [(text, href)]
next_div = li.find('div', class_=re.compile(rf'\blb{level}\b'), recursive=False)
if next_div:
next_ul = next_div.find('ul', class_=re.compile(rf'\bld{level+1}\b'), recursive=False)
if next_ul:
out.append({'path': list(path), 'href': href})
walk(next_ul, level + 1, path, out)
continue
out.append({'path': list(path), 'href': href})
tmp = []
walk(container, 1, [], tmp)
cols = 'DEFGHIJ'
for item in tmp:
p = item['path']
row = {c: '' for c in cols}
row['href'] = item['href']
for i, (t, _) in enumerate(p):
col = cols[i] if i < len(cols) else 'J'
row[col] = t
rows.append(row)
return rows
def parse_dl_dt_dd(soup, base, use_onclick=False):
"""계룡시·홍성군·영동군형: div.sitemap > dl > dt + dd > b > a + ul > li > a (+ ul > li > a)."""
rows = []
sm = soup.select_one('div.sitemap[class*=type2]') or soup.select_one('div.sitemap[class*=type1]') or soup.select_one('div.sitemap')
if not sm:
return rows
href_fn = extract_href_with_onclick if use_onclick else extract_href
def walk(ul, depth, base_path, out):
for li in ul.find_all('li', recursive=False):
a = li.find('a', recursive=False)
if not a:
continue
text = clean_text(a.get_text())
href = href_fn(a)
path = base_path[:depth] + [(text, href)]
nested = li.find('ul', recursive=False)
if nested:
out.append({'path': list(path), 'href': href})
walk(nested, depth + 1, path, out)
else:
out.append({'path': list(path), 'href': href})
for dl in sm.find_all('dl', recursive=False):
dt = dl.find('dt')
dt_a = dt.find('a') if dt else None
D = clean_text(dt_a.get_text() if dt_a else (dt.get_text() if dt else ''))
for dd in dl.find_all('dd', recursive=False):
b = dd.find('b')
b_a = b.find('a') if b else None
E = clean_text(b_a.get_text()) if b_a else ''
E_href = href_fn(b_a) if b_a else ''
nested = dd.find('ul', recursive=False)
if nested:
tmp = []
walk(nested, 0, [], tmp)
for item in tmp:
p = item['path']
row = {'D': D, 'E': E, 'href': item['href'],
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''}
for di, (t, _) in enumerate(p):
col = 'FGHIJ'[di] if di < 5 else 'J'
row[col] = t
rows.append(row)
else:
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
return rows
def parse_yesan_depth(soup, base):
"""예산·청양·증평형: ul.depth1_ul > li > a (D, .th_1st/.th1_lnk) + [div.item >]? ul.depth2_ul > li > a (E) + ul.depth3_ul > li > a (F).
div.item wrapper 있을수도 없을수도 지원.
a 클래스: .th_1st (예산), .th1_lnk (증평) 번째 a 가져옴.
"""
rows = []
container = soup.select_one('ul.depth1_ul')
if not container:
return rows
for top_li in container.find_all('li', recursive=False):
a1 = top_li.find('a', class_=re.compile(r'th[_]?1?(?:_1st|_lnk)?'), recursive=False) or \
top_li.find('a', recursive=False)
if not a1:
continue
D = clean_text(a1.get_text())
# div.item wrapper 있을 수도 없을 수도
item = top_li.find('div', class_='item', recursive=False)
d2_ul = (item.find('ul', class_='depth2_ul') if item else None) or \
top_li.find('ul', class_='depth2_ul', recursive=False)
if not d2_ul:
continue
for d2_li in d2_ul.find_all('li', recursive=False):
d2_a = d2_li.find('a', recursive=False)
if not d2_a:
continue
E = clean_text(d2_a.get_text())
E_href = extract_href(d2_a)
d3_ul = d2_li.find('ul', class_='depth3_ul', recursive=False)
if not d3_ul:
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
for d3_li in d3_ul.find_all('li', recursive=False):
d3_a = d3_li.find('a', recursive=False)
if not d3_a:
continue
F = clean_text(d3_a.get_text())
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(d3_a),
'G': '', 'H': '', 'I': '', 'J': ''})
return rows
def parse_jincheon_recursive(soup, base):
"""진천군: div.sitemap > ul > li > a + div > ul > li > a + div > ul > ... 재귀."""
rows = []
container = soup.select_one('div.sitemap')
if not container:
return rows
top_ul = container.find('ul', recursive=False)
if not top_ul:
return rows
def walk(ul, base_path, out):
for li in ul.find_all('li', recursive=False):
a = li.find('a', recursive=False)
if not a:
continue
text = clean_text(a.get_text())
if not text or text == '메뉴명이 없습니다.':
continue
href = extract_href(a)
path = base_path + [(text, href)]
next_div = li.find('div', recursive=False)
if next_div:
next_ul = next_div.find('ul', recursive=False)
if next_ul:
out.append({'path': list(path), 'href': href})
walk(next_ul, path, out)
continue
out.append({'path': list(path), 'href': href})
tmp = []
walk(top_ul, [], tmp)
cols = 'DEFGHIJ'
for item in tmp:
p = item['path']
row = {c: '' for c in cols}
row['href'] = item['href']
for i, (t, _) in enumerate(p):
col = cols[i] if i < len(cols) else 'J'
row[col] = t
rows.append(row)
return rows
def parse_cheongju_sitemap(soup, base):
"""청주시: div#sitemap > div.site_map_col > div.sitemap_box > h3 > a (D) + ul.sm2depth > li > a (E) + ul.sm3depth > li > a (F).
별도의 '인트로' sitemap_box D만 있고 ul.sm2depth 단순한 외부 링크 묶음 그대로 .
"""
rows = []
container = soup.select_one('div#sitemap')
if not container:
return rows
for box in container.select('div.sitemap_box'):
h3 = box.find('h3')
h3_a = h3.find('a') if h3 else None
D = clean_text(h3_a.get_text() if h3_a else (h3.get_text() if h3 else ''))
D_href = extract_href(h3_a) if h3_a else ''
sm2 = box.find('ul', class_='sm2depth')
if not sm2:
rows.append({'D': D, 'href': D_href, 'E': '', 'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
for li2 in sm2.find_all('li', recursive=False):
a2 = li2.find('a', recursive=False)
if not a2:
continue
E = clean_text(a2.get_text())
E_href = extract_href(a2)
sm3 = li2.find('ul', class_='sm3depth', recursive=False)
if not sm3:
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
for li3 in sm3.find_all('li', recursive=False):
a3 = li3.find('a', recursive=False)
if not a3:
continue
F = clean_text(a3.get_text())
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(a3),
'G': '', 'H': '', 'I': '', 'J': ''})
return rows
def parse_chungju_sitemap(soup, base):
"""충주시: div#sitemap > div.site_map_col > div.sitemap_box > h3.h0 > a (D) + ul > li > a.h4 (E) + ul.bu > li > a (F)."""
rows = []
container = soup.select_one('div#sitemap')
if not container:
return rows
for box in container.select('div.sitemap_box'):
h3 = box.find('h3')
h3_a = h3.find('a') if h3 else None
D = clean_text(h3_a.get_text() if h3_a else (h3.get_text() if h3 else ''))
D_href = extract_href(h3_a) if h3_a else ''
e_ul = box.find('ul', recursive=False)
if not e_ul:
rows.append({'D': D, 'href': D_href, 'E': '', 'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
for e_li in e_ul.find_all('li', recursive=False):
a_E = e_li.find('a', class_='h4', recursive=False) or e_li.find('a', recursive=False)
if not a_E:
continue
E = clean_text(a_E.get_text())
E_href = extract_href(a_E)
f_ul = e_li.find('ul', class_='bu', recursive=False)
if not f_ul:
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
for f_li in f_ul.find_all('li', recursive=False):
f_a = f_li.find('a', recursive=False)
if not f_a:
continue
F = clean_text(f_a.get_text())
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(f_a),
'G': '', 'H': '', 'I': '', 'J': ''})
return rows
# ================================================================
# 사이트 설정
# ================================================================
SITES = [
{
'idx': 1, 'name': '괴산군', 'base': 'https://www.goesan.go.kr',
'sitemap': 'https://www.goesan.go.kr/www/sitemap.do?key=28',
'sheet': '01_괴산군', 'parser': parse_depth1_chungbuk,
'domain': 'goesan.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\1.괴산군',
'weak_ssl': False,
},
{
'idx': 2, 'name': '단양군', 'base': 'https://www.danyang.go.kr',
'sitemap': 'https://www.danyang.go.kr/dy21/98',
'sheet': '02_단양군', 'parser': parse_danyang_ld,
'domain': 'danyang.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\2.단양군',
'weak_ssl': False,
},
{
# 보은군 — www.boeun.go.kr DNS 차단, apex boeun.go.kr 는 정상(2026-05-30 재확인)
'idx': 3, 'name': '보은군', 'base': 'https://boeun.go.kr',
'sitemap': 'https://boeun.go.kr/www/sitemap.do?key=1323',
'sheet': '03_보은군', 'parser': parse_depth1_chungbuk,
'domain': 'boeun.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\3.보은군',
'weak_ssl': False,
},
{
'idx': 4, 'name': '영동군', 'base': 'https://www.yd21.go.kr',
'sitemap': 'https://www.yd21.go.kr/kr/html/guide/0701.html',
'sheet': '04_영동군', 'parser': parse_dl_dt_dd,
'domain': 'yd21.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\4.영동군',
'weak_ssl': True,
},
{
'idx': 5, 'name': '옥천군', 'base': 'https://www.oc.go.kr',
'sitemap': 'https://www.oc.go.kr/www/sub.do?key=121',
'sheet': '05_옥천군', 'parser': parse_depth1_chungbuk,
'domain': 'oc.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\5.옥천군',
'weak_ssl': False,
},
{
'idx': 6, 'name': '음성군', 'base': 'https://www.eumseong.go.kr',
'sitemap': 'https://www.eumseong.go.kr/www/sub.do?key=722',
'sheet': '06_음성군', 'parser': parse_depth1_chungbuk,
'domain': 'eumseong.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\6.음성군',
'weak_ssl': False,
},
{
'idx': 7, 'name': '제천시', 'base': 'https://www.jecheon.go.kr',
'sitemap': 'https://www.jecheon.go.kr/www/sitemap.do?key=553',
'sheet': '07_제천시', 'parser': parse_depth1_chungbuk,
'domain': 'jecheon.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\7.제천시',
'weak_ssl': False,
},
{
'idx': 8, 'name': '증평군', 'base': 'https://www.jp.go.kr',
'sitemap': 'https://www.jp.go.kr/kor/sitemap_11.do',
'sheet': '08_증평군', 'parser': parse_yesan_depth,
'domain': 'jp.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\8.증평군',
'weak_ssl': False,
},
{
'idx': 9, 'name': '진천군', 'base': 'https://www.jincheon.go.kr',
'sitemap': 'https://www.jincheon.go.kr/home/sub.do?menukey=445',
'sheet': '09_진천군', 'parser': parse_jincheon_recursive,
'domain': 'jincheon.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\9.진천군',
'weak_ssl': False,
},
{
'idx': 10, 'name': '청주시', 'base': 'https://www.cheongju.go.kr',
'sitemap': 'https://www.cheongju.go.kr/www/sitemap.do?key=589',
'sheet': '10_청주시', 'parser': parse_cheongju_sitemap,
'domain': 'cheongju.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\10.청주시',
'weak_ssl': False,
},
{
'idx': 11, 'name': '충주시', 'base': 'https://www.chungju.go.kr',
'sitemap': 'https://www.chungju.go.kr/www/sub.do?key=692',
'sheet': '11_충주시', 'parser': parse_chungju_sitemap,
'domain': 'chungju.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\11.충주시',
'weak_ssl': False,
},
]
# ================================================================
# 엑셀 생성 (공통)
# ================================================================
def write_excel(site, raw_rows):
name = site['name']
base = site['base']
domain = site['domain']
output = f"{site['folder']}\\{os.path.basename(os.path.dirname(site['folder']))}_{name}.xlsx"
def abs_url(href):
if not href:
return ''
href = href.strip()
if href.startswith(('javascript:', '#')):
return ''
if href.startswith(('http://', 'https://')):
return href
return urljoin(base + '/', href)
def is_external(url):
return url.startswith(('http://', 'https://')) and domain not in url
# 부모-자식 URL 중복 제거
final_rows = []
i = 0
removed = 0
while i < len(raw_rows):
row = raw_rows[i]
if (i + 1 < len(raw_rows)
and row.get('G', '') == ''
and raw_rows[i + 1].get('D') == row.get('D')
and raw_rows[i + 1].get('E') == row.get('E')
and raw_rows[i + 1].get('F') == row.get('F')
and raw_rows[i + 1].get('G', '') != ''
and raw_rows[i + 1].get('href') == row.get('href')):
removed += 1
i += 1
continue
final_rows.append(row)
i += 1
print(f' [{name}] 원본 {len(raw_rows)} → 중복 {removed} → 최종 {len(final_rows)}')
if not final_rows:
print(f' [{name}] !! 행 0개 — 파서 점검 필요. 엑셀 생성 스킵.')
return False
shutil.copy(TEMPLATE, output)
wb = openpyxl.load_workbook(output)
ws = wb.active
ws.title = site['sheet']
HEADER = {'B1:R1', 'S1:W1', 'Y1:AA1'}
for rng in [str(m) for m in ws.merged_cells.ranges if str(m) not in HEADER]:
ws.unmerge_cells(rng)
for row in ws.iter_rows(min_row=3, max_row=ws.max_row, min_col=1, max_col=ws.max_column):
for cell in row:
cell.value = None
START = 3
template_r = 3
cur_max = ws.max_row
for idx, item in enumerate(final_rows, start=START):
if idx > cur_max:
for c in range(1, ws.max_column + 1):
srcc = ws.cell(template_r, c)
tgt = ws.cell(idx, c)
if srcc.has_style:
tgt.font = copy(srcc.font)
tgt.fill = copy(srcc.fill)
tgt.border = copy(srcc.border)
tgt.alignment = copy(srcc.alignment)
tgt.number_format = srcc.number_format
tgt.protection = copy(srcc.protection)
url = abs_url(item.get('href', ''))
ws.cell(idx, 2).value = idx - 2
ws.cell(idx, 3).value = name
ws.cell(idx, 4).value = item.get('D', '')
ws.cell(idx, 5).value = item.get('E', '')
ws.cell(idx, 6).value = item.get('F', '')
ws.cell(idx, 7).value = item.get('G', '')
ws.cell(idx, 8).value = item.get('H', '')
ws.cell(idx, 9).value = item.get('I', '')
ws.cell(idx, 10).value = item.get('J', '')
ws.cell(idx, 11).value = url
if is_external(url):
ws.cell(idx, 19).value = '외부링크'
END = START + len(final_rows) - 1
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
def merge_runs(col_letter, col_idx, group_cols=()):
runs = []
cur_val = ws.cell(START, col_idx).value
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
run_start = START
for r in range(START + 1, END + 1):
v = ws.cell(r, col_idx).value
g = tuple(ws.cell(r, gg).value for gg in group_cols)
if v == cur_val and g == cur_grp:
continue
if cur_val not in (None, '') and r - 1 > run_start:
runs.append((run_start, r - 1))
cur_val, cur_grp, run_start = v, g, r
if cur_val not in (None, '') and END > run_start:
runs.append((run_start, END))
for s, e in runs:
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
ws.cell(s, col_idx).alignment = center
return len(runs)
# 병합 순서 F→E→D (D를 먼저 병합하면 2행부터 D=None이 되어 E 그룹키가 깨짐)
n_f = merge_runs('F', 6, group_cols=(4, 5))
n_e = merge_runs('E', 5, group_cols=(4,))
n_d = merge_runs('D', 4)
for r in range(START, END + 1):
for c in (4, 5, 6, 7):
if ws.cell(r, c).value is not None:
ws.cell(r, c).alignment = center
for r in range(1, END + 1):
ws.row_dimensions[r].height = 15
link_n = 0
for r in range(START, END + 1):
cell = ws.cell(r, 11)
u = cell.value
if u and isinstance(u, str) and u.startswith(('http://', 'https://')):
cell.hyperlink = u
old = cell.font
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
bold=old.bold, italic=old.italic,
color='0000FF', underline='single')
cell.alignment = left
link_n += 1
wb.save(output)
ext_n = sum(1 for r in range(START, END + 1) if ws.cell(r, 19).value == '외부링크')
print(f' [{name}] 병합 D:{n_d} E:{n_e} F:{n_f} | 외부링크 {ext_n} | K링크 {link_n}{output}')
return True
def main():
targets = sys.argv[1:] if len(sys.argv) > 1 else [s['name'] for s in SITES]
for site in SITES:
if site['name'] not in targets:
continue
print(f"\n{'='*60}\n[{site['idx']}.{site['name']}] {site['sitemap']}\n{'='*60}")
try:
sess = make_session(weak_ssl=site.get('weak_ssl', False))
html = fetch_html(site['sitemap'], session=sess)
soup = BeautifulSoup(html, 'html.parser')
raw_rows = site['parser'](soup, site['base'])
write_excel(site, raw_rows)
except Exception as e:
print(f' [{site["name"]}] !! 실패: {e}')
import traceback
traceback.print_exc()
if __name__ == '__main__':
main()

View File

@ -0,0 +1,357 @@
"""충청북도 10개 시·군 Phase 2~4 일괄 처리 (보은군 제외 — DNS 차단).
매뉴얼: D:\\01.프로젝트\\DB수집\\사이트맵_수집_매뉴얼.md
"""
import re
import ssl
import sys
import time
import warnings
from urllib.parse import urljoin
from concurrent.futures import ThreadPoolExecutor, as_completed
import openpyxl
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.ssl_ import create_urllib3_context
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
class WeakSSLAdapter(HTTPAdapter):
def init_poolmanager(self, *args, **kwargs):
ctx = create_urllib3_context()
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
ctx.options |= 0x4
ctx.check_hostname = False
ctx.verify_mode = ssl.CERT_NONE
kwargs['ssl_context'] = ctx
return super().init_poolmanager(*args, **kwargs)
# ================================================================
# 정규식
# ================================================================
TOTAL_PAT = re.compile(r'\s*(?:게시물\s*)?(\d[\d,]*)\s*(?:건|개|page|페이지)', re.I)
TOTAL_PAT_LOOSE = re.compile(r'(?:전체|총)\s*[:\-]?\s*(\d[\d,]*)\s*건', re.I)
KOGL_IMG_PAT = re.compile(r'(?:new_)?img_opentype(\d{2})\.png', re.I)
KOGL_LINK_PAT = re.compile(r'kogl\.or\.kr/info/licenseType(\d)', re.I)
YOUTUBE_PAT = re.compile(r'(?:youtube\.com|youtu\.be)', re.I)
VIDEO_EXT = re.compile(r'\.(mp4|webm|mov|avi)(?:\?|$)', re.I)
DETAIL_PAT = re.compile(r'(mode=V|view\.do|bbtSn=|seqRepeat=|nttId=|articleNo=|boardSeq=|menukey=)', re.I)
# 사이트별: xlsx, body selectors, weak_ssl 여부
SITES = {
'괴산군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\1.괴산군\충청북도_괴산군.xlsx',
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
'단양군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\2.단양군\충청북도_단양군.xlsx',
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
'보은군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\3.보은군\충청북도_보은군.xlsx',
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
'영동군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\4.영동군\충청북도_영동군.xlsx',
'body_sel': ['#txt', '#contents', 'main'], 'weak_ssl': True},
'옥천군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\5.옥천군\충청북도_옥천군.xlsx',
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
'음성군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\6.음성군\충청북도_음성군.xlsx',
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
'제천시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\7.제천시\충청북도_제천시.xlsx',
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
'증평군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\8.증평군\충청북도_증평군.xlsx',
'body_sel': ['#txt', '#contents', 'main'], 'weak_ssl': False},
'진천군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\9.진천군\충청북도_진천군.xlsx',
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
'청주시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\10.청주시\충청북도_청주시.xlsx',
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
'충주시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\11.충주시\충청북도_충주시.xlsx',
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
}
def make_session(weak_ssl=False):
s = requests.Session()
s.headers.update(H)
if weak_ssl:
s.mount('https://', WeakSSLAdapter())
return s
def fetch(session, url, timeout=12):
try:
r = session.get(url, timeout=timeout, verify=False, allow_redirects=True)
# 메타 charset 우선
meta_charset = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
if meta_charset:
r.encoding = meta_charset.group(1).decode('ascii', errors='ignore')
else:
r.encoding = r.apparent_encoding
if r.status_code == 200:
return BeautifulSoup(r.text, 'html.parser')
except Exception:
pass
return None
def get_body(soup, selectors):
for sel in selectors:
el = soup.select_one(sel)
if el:
return el
return soup
def detect_form(body):
has_paging = bool(body.select('.pagination, .paging, nav.paging, .page_nav'))
text_inputs = [i for i in body.find_all('input')
if (i.get('type') or 'text').lower() in ('text', 'search')]
has_search = len(text_inputs) >= 1
txt = body.get_text(' ', strip=True)
m = TOTAL_PAT.search(txt) or TOTAL_PAT_LOOSE.search(txt)
total = None
if m:
digits = m.group(1).replace(',', '')
if digits.isdigit():
total = int(digits)
is_board = has_paging or has_search or (total is not None)
if is_board:
return '게시판', total if total is not None else 0
return '페이지', 1
def extract_detail_urls(body, base_url, limit=5):
urls = []
seen = set()
for a in body.find_all('a', href=True):
h = a['href']
if not h or h.startswith('#'):
continue
if DETAIL_PAT.search(h):
full = urljoin(base_url, h)
if full not in seen:
seen.add(full)
urls.append(full)
if len(urls) >= limit:
break
return urls
def detect_media(body):
has_text = len(body.get_text(strip=True)) > 30
has_image = False
for img in body.find_all('img'):
src = img.get('src', '')
if KOGL_IMG_PAT.search(src):
continue
if not src:
continue
has_image = True
break
has_video = False
for iframe in body.find_all('iframe'):
if YOUTUBE_PAT.search(iframe.get('src', '')):
has_video = True
break
if not has_video:
for a in body.find_all('a', href=True):
if YOUTUBE_PAT.search(a['href']):
has_video = True
break
if not has_video and body.find_all('video'):
has_video = True
if not has_video and VIDEO_EXT.search(str(body)):
has_video = True
return has_image, has_video, has_text
def n_string(has_text, has_image, has_video):
parts = []
if has_text: parts.append('어문')
if has_image: parts.append('이미지')
if has_video: parts.append('영상')
return ','.join(parts) if parts else '없음'
def img_has_valid_anchor(img):
p = img.parent
while p is not None:
if p.name == 'a':
href = p.get('href', '')
if href and not href.startswith('#') and not href.lower().startswith('javascript:'):
return True
return False
p = p.parent
return False
def detect_kogl(body):
types = set()
q_any_y = False
q_any_n = False
for a in body.find_all('a', href=True):
m = KOGL_LINK_PAT.search(a['href'])
if m:
types.add(int(m.group(1)))
q_any_y = True
for img in body.find_all('img'):
src = img.get('src', '')
m = KOGL_IMG_PAT.search(src)
if m:
types.add(int(m.group(1)))
if img_has_valid_anchor(img):
q_any_y = True
else:
q_any_n = True
for el in body.find_all(style=True):
m = KOGL_IMG_PAT.search(el.get('style', ''))
if m:
types.add(int(m.group(1)))
q_any_n = True
if not types:
return set(), None
return types, ('Y' if q_any_y else 'N')
def process_row(session, url, body_selectors):
out = {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
soup = fetch(session, url)
if soup is None:
out['note'] = '접근 실패'
return out
body = get_body(soup, body_selectors)
form, count = detect_form(body)
out['L'] = form
out['M'] = count if form == '게시판' else 1
has_img, has_vid, has_txt = detect_media(body)
types_main, q_main = detect_kogl(body)
P = '게시판' if types_main else ''
types_all = set(types_main)
q_flags = []
if q_main:
q_flags.append(q_main)
if form == '게시판':
detail_urls = extract_detail_urls(body, url, limit=5)
for du in detail_urls:
d_soup = fetch(session, du, timeout=10)
if not d_soup:
continue
d_body = get_body(d_soup, body_selectors)
di, dv, dt = detect_media(d_body)
has_img = has_img or di
has_vid = has_vid or dv
has_txt = has_txt or dt
dt_types, dt_q = detect_kogl(d_body)
if dt_types and not types_main and not P:
P = '게시물'
types_all |= dt_types
if dt_q:
q_flags.append(dt_q)
out['N'] = n_string(has_txt, has_img, has_vid)
if not types_all:
out['O'] = '미부착'
else:
sorted_types = sorted(types_all)
out['O'] = ','.join(f'{n}유형' for n in sorted_types)
out['P'] = P if P else '게시판'
out['Q'] = 'Y' if 'Y' in q_flags else 'N'
return out
def run_site(name, xlsx, body_selectors, weak_ssl=False, workers=10):
print(f'\n{"="*60}\n[{name}] {xlsx}\n{"="*60}')
wb = openpyxl.load_workbook(xlsx)
ws = wb.active
START = 3
END = START - 1
for r in range(START, ws.max_row + 1):
if ws.cell(r, 2).value is None:
break
END = r
tasks = []
for r in range(START, END + 1):
url = ws.cell(r, 11).value
is_ext = (ws.cell(r, 19).value == '외부링크')
tasks.append((r, url, is_ext))
n_ext = sum(1 for t in tasks if t[2])
print(f' 처리 대상: 총 {len(tasks)}행 (외부링크 {n_ext})')
t0 = time.time()
results = {}
session = make_session(weak_ssl=weak_ssl)
def worker(task):
row, url, is_ext = task
if is_ext:
return row, {'L': '사이트', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
if not url or not isinstance(url, str):
return row, {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': 'URL 없음'}
return row, process_row(session, url, body_selectors)
done = 0
with ThreadPoolExecutor(max_workers=workers) as ex:
futs = [ex.submit(worker, t) for t in tasks]
for fut in as_completed(futs):
row, res = fut.result()
results[row] = res
done += 1
if done % 50 == 0 or done == len(tasks):
print(f' 진행 {done}/{len(tasks)} ({time.time()-t0:.0f}s)')
for r in range(START, END + 1):
res = results.get(r, {})
if not res:
continue
if res.get('L'): ws.cell(r, 12).value = res['L']
if res.get('M') != '': ws.cell(r, 13).value = res['M']
if res.get('N'): ws.cell(r, 14).value = res['N']
if res.get('O'): ws.cell(r, 15).value = res['O']
if res.get('P'): ws.cell(r, 16).value = res['P']
if res.get('Q'): ws.cell(r, 17).value = res['Q']
if res.get('note'):
existing = ws.cell(r, 19).value
if not existing:
ws.cell(r, 19).value = res['note']
wb.save(xlsx)
forms = {}
attach = {'미부착': 0, '부착': 0, '기타': 0}
q_dist = {'Y': 0, 'N': 0, '': 0}
for r, res in results.items():
forms[res.get('L', '')] = forms.get(res.get('L', ''), 0) + 1
o = res.get('O', '')
if o == '미부착': attach['미부착'] += 1
elif o and '유형' in o: attach['부착'] += 1
else: attach['기타'] += 1
q = res.get('Q', '')
q_dist[q] = q_dist.get(q, 0) + 1
print(f' [{name}] L분포 {forms} | 부착 {attach} | Q {q_dist} | 시간 {time.time()-t0:.0f}s')
def main():
targets = sys.argv[1:] if len(sys.argv) > 1 else list(SITES.keys())
total_t0 = time.time()
for name in targets:
if name not in SITES:
print(f' 알 수 없음: {name}')
continue
cfg = SITES[name]
try:
run_site(name, cfg['xlsx'], cfg['body_sel'], weak_ssl=cfg.get('weak_ssl', False))
except Exception as e:
print(f' [{name}] 실패: {e}')
import traceback
traceback.print_exc()
print(f'\n총 소요: {time.time()-total_t0:.0f}s')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,734 @@
"""충청남도 11개 시·군 Phase 1 일괄 처리.
매뉴얼: D:\\01.프로젝트\\DB수집\\사이트맵_수집_매뉴얼.md
출력: 폴더의 {기관명}.xlsx
"""
import os
import re
import shutil
import sys
import warnings
from copy import copy
from urllib.parse import urljoin
import openpyxl
import requests
from bs4 import BeautifulSoup
from openpyxl.styles import Alignment, Font
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
TEMPLATE = r'D:\01.프로젝트\DB수집\자료_취합_예시.xlsx'
# ================================================================
# 공통 유틸
# ================================================================
def fetch_html(url, timeout=20):
r = requests.get(url, headers=H, timeout=timeout, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
return r.text
def clean_text(s):
return re.sub(r'\s+', ' ', s or '').strip().replace('\xa0', '')
def extract_href(a):
if a is None:
return ''
href = (a.get('href') or '').strip()
if not href or href.startswith('#') or href.lower().startswith('javascript:'):
# Try onclick encodeURI extraction
return ''
return href
ENCODE_URI_PAT = re.compile(r"""(?:location\.href|window\.location(?:\.href)?)\s*=\s*encodeURI\(\s*['"]([^'"]+)['"]\s*\)""")
ONCLICK_HREF_PAT = re.compile(r"""(?:location\.href|window\.open)\s*\(?\s*['"]([^'"]+)['"]""")
def extract_href_with_onclick(a):
"""Extract href; if href is dummy (#...), check onclick for encodeURI/location.href."""
if a is None:
return ''
href = (a.get('href') or '').strip()
if href and not href.startswith('#') and not href.lower().startswith('javascript:'):
return href
# Try onclick
onclick = a.get('onclick', '')
if onclick:
m = ENCODE_URI_PAT.search(onclick)
if m:
return m.group(1)
m = ONCLICK_HREF_PAT.search(onclick)
if m:
return m.group(1)
return ''
# ================================================================
# 사이트별 파서 (각각 raw_rows 리스트 반환)
# raw_rows: [{'D': str, 'E': str, 'F': str, 'G': str, 'H': str, 'I': str, 'J': str, 'href': str}, ...]
# ================================================================
def parse_eGov_type1(soup, base):
"""논산시·아산시 형: div.sitemap.type1 > [div.s_1th + div.inner > div.s_2th + ul...]."""
sitemap = soup.select_one('div.sitemap.type1') or soup.select_one('div.sitemap')
rows = []
if not sitemap:
return rows
current_D = ''
def walk(ul, depth, base_path, out):
for li in ul.find_all('li', recursive=False):
a = li.find('a', recursive=False)
if not a:
continue
text = clean_text(a.get_text())
href = extract_href(a)
path = base_path[:depth] + [(text, href)]
nested = li.find('ul', recursive=False)
if nested:
out.append({'path': list(path), 'href': href})
walk(nested, depth + 1, path, out)
else:
out.append({'path': list(path), 'href': href})
for child in sitemap.find_all('div', recursive=False):
cls = child.get('class', [])
if 's_1th' in cls:
a = child.find('a')
current_D = clean_text(a.get_text()) if a else ''
elif 'inner' in cls:
s2 = child.find('div', class_='s_2th')
mid_a = s2.find('a') if s2 else None
mid_name = clean_text(mid_a.get_text()) if mid_a else ''
mid_href = extract_href(mid_a) if mid_a else ''
uls = child.find_all('ul', recursive=False)
if not uls:
rows.append({'D': current_D, 'E': mid_name, 'href': mid_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
for ul in uls:
tmp = []
walk(ul, 0, [], tmp)
for item in tmp:
p = item['path']
row = {'D': current_D, 'E': mid_name, 'href': item['href'],
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''}
for di, (t, _) in enumerate(p):
col = 'FGHIJ'[di] if di < 5 else 'J'
row[col] = t
rows.append(row)
return rows
def parse_amThum(soup, base):
"""당진시·태안군 형: div.amThum > h2.siteNN + div.sitemap_grep > ul.sitemap_list > li > a.first + ul > li > a (+ ul > li > a)."""
rows = []
for amthum in soup.select('div.amThum'):
h2 = amthum.find(['h2', 'h3'], class_=re.compile(r'site\d+'))
D = clean_text(h2.find('span').get_text()) if h2 and h2.find('span') else (clean_text(h2.get_text()) if h2 else '')
grep = amthum.find('div', class_='sitemap_grep') or amthum
for sl in grep.find_all('ul', class_='sitemap_list'):
# Each ul.sitemap_list contains li > a.first + ul > li > a
for top_li in sl.find_all('li', recursive=False):
a_first = top_li.find('a', class_='first', recursive=False)
if not a_first:
continue
E = clean_text(a_first.get_text())
E_href = extract_href(a_first)
inner_ul = top_li.find('ul', recursive=False)
if not inner_ul:
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
for mid_li in inner_ul.find_all('li', recursive=False):
mid_a = mid_li.find('a', recursive=False)
if not mid_a:
continue
F = clean_text(mid_a.get_text())
F_href = extract_href(mid_a)
deeper = mid_li.find('ul', recursive=False)
if not deeper:
rows.append({'D': D, 'E': E, 'F': F, 'href': F_href,
'G': '', 'H': '', 'I': '', 'J': ''})
continue
rows.append({'D': D, 'E': E, 'F': F, 'href': F_href,
'G': '', 'H': '', 'I': '', 'J': ''})
for deep_li in deeper.find_all('li', recursive=False):
deep_a = deep_li.find('a', recursive=False)
if not deep_a:
continue
G = clean_text(deep_a.get_text())
rows.append({'D': D, 'E': E, 'F': F, 'G': G,
'href': extract_href(deep_a),
'H': '', 'I': '', 'J': ''})
return rows
def parse_ul_sitemap_h4(soup, base):
"""보령시·서천군 형: ul.sitemap > li (대분류) > h4.siteNN > span + ul > li > h5 > a + ul > li.list > a."""
rows = []
container = soup.select_one('ul.sitemap') or soup.select_one('div#contents ul.sitemap')
if not container:
return rows
for top_li in container.find_all('li', recursive=False):
h4 = top_li.find(['h4', 'h3'], recursive=False)
D = ''
if h4:
span = h4.find('span')
D = clean_text(span.get_text() if span else h4.get_text())
# Each direct ul under top_li is a sub-group
for sub_ul in top_li.find_all('ul', recursive=False):
for sub_li in sub_ul.find_all('li', recursive=False):
h5 = sub_li.find(['h5', 'h6'], recursive=False)
if h5:
h5_a = h5.find('a')
E = clean_text(h5_a.get_text()) if h5_a else clean_text(h5.get_text())
E_href = extract_href(h5_a) if h5_a else ''
else:
a = sub_li.find('a', recursive=False)
E = clean_text(a.get_text()) if a else ''
E_href = extract_href(a) if a else ''
deeper = sub_li.find('ul', recursive=False)
if not deeper:
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
for leaf_li in deeper.find_all('li', recursive=False):
leaf_a = leaf_li.find('a', recursive=False)
if not leaf_a:
continue
F = clean_text(leaf_a.get_text())
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(leaf_a),
'G': '', 'H': '', 'I': '', 'J': ''})
# 4-level (rare)
sub_ul2 = leaf_li.find('ul', recursive=False)
if sub_ul2:
for ll2 in sub_ul2.find_all('li', recursive=False):
la2 = ll2.find('a', recursive=False)
if la2:
G = clean_text(la2.get_text())
rows.append({'D': D, 'E': E, 'F': F, 'G': G,
'href': extract_href(la2),
'H': '', 'I': '', 'J': ''})
return rows
def parse_dl_dt_dd(soup, base, use_onclick=False):
"""홍성군·예산군·천안시·금산군 dl 형: div.sitemap (옵션 .type2) > dl > dt + dd > b > a + ul > li > a (+ ul > li > a).
use_onclick=True 경우 onclick의 encodeURI/location.href에서 URL을 추출.
"""
rows = []
sm = soup.select_one('div.sitemap[class*=type2]') or soup.select_one('div.sitemap')
if not sm:
return rows
href_fn = extract_href_with_onclick if use_onclick else extract_href
def walk(ul, depth, base_path, out):
for li in ul.find_all('li', recursive=False):
a = li.find('a', recursive=False)
if not a:
continue
text = clean_text(a.get_text())
href = href_fn(a)
path = base_path[:depth] + [(text, href)]
nested = li.find('ul', recursive=False)
if nested:
out.append({'path': list(path), 'href': href})
walk(nested, depth + 1, path, out)
else:
out.append({'path': list(path), 'href': href})
for dl in sm.find_all('dl', recursive=False):
dt = dl.find('dt')
dt_a = dt.find('a') if dt else None
D = clean_text(dt_a.get_text() if dt_a else (dt.get_text() if dt else ''))
for dd in dl.find_all('dd', recursive=False):
b = dd.find('b')
b_a = b.find('a') if b else None
E = clean_text(b_a.get_text()) if b_a else ''
E_href = href_fn(b_a) if b_a else ''
nested = dd.find('ul', recursive=False)
if nested:
tmp = []
walk(nested, 0, [], tmp)
for item in tmp:
p = item['path']
row = {'D': D, 'E': E, 'href': item['href'],
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''}
for di, (t, _) in enumerate(p):
col = 'FGHIJ'[di] if di < 5 else 'J'
row[col] = t
rows.append(row)
else:
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
return rows
def parse_seosan_top_menu(soup, base):
"""서산시: ul.top_menu > li.depth1 > a.depth1_ti (D) + div...div.depth2_wrap > ul.depth2 > li > a (E) + ul.depth3 > li > a (F)."""
rows = []
tm = soup.select_one('ul.top_menu')
if not tm:
return rows
for top_li in tm.find_all('li', class_='depth1', recursive=False):
a1 = top_li.find('a', class_='depth1_ti')
D = clean_text(a1.get_text()) if a1 else ''
depth2 = top_li.select_one('ul.depth2')
if not depth2:
continue
for d2_li in depth2.find_all('li', recursive=False):
d2_a = d2_li.find('a', recursive=False)
if not d2_a:
continue
E = clean_text(d2_a.get_text())
E_href = extract_href(d2_a)
d3 = d2_li.find('ul', class_='depth3', recursive=False)
if not d3:
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
for d3_li in d3.find_all('li', recursive=False):
d3_a = d3_li.find('a', recursive=False)
if not d3_a:
continue
F = clean_text(d3_a.get_text())
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(d3_a),
'G': '', 'H': '', 'I': '', 'J': ''})
return rows
def parse_yesan_depth(soup, base):
"""예산군: #gnb > ul.depth1_ul > li > a.th_1st + div.item > ul.depth2_ul > li > a + ul.depth3_ul > li > a."""
rows = []
container = soup.select_one('#gnb ul.depth1_ul') or soup.select_one('ul.depth1_ul')
if not container:
return rows
for top_li in container.find_all('li', recursive=False):
a1 = top_li.find('a', class_='th_1st', recursive=False)
D = clean_text(a1.get_text()) if a1 else ''
item = top_li.find('div', class_='item')
if not item:
continue
d2_ul = item.find('ul', class_='depth2_ul')
if not d2_ul:
continue
for d2_li in d2_ul.find_all('li', recursive=False):
d2_a = d2_li.find('a', recursive=False)
if not d2_a:
continue
E = clean_text(d2_a.get_text())
E_href = extract_href(d2_a)
d3_ul = d2_li.find('ul', class_='depth3_ul', recursive=False)
if not d3_ul:
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
for d3_li in d3_ul.find_all('li', recursive=False):
d3_a = d3_li.find('a', recursive=False)
if not d3_a:
continue
F = clean_text(d3_a.get_text())
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(d3_a),
'G': '', 'H': '', 'I': '', 'J': ''})
return rows
def parse_cheonan_depth(soup, base):
"""천안시: ul.depth1-ul > li > a.depth1-btn + div.depth1-content > div.layout > ul.depth2-ul > li > a.depth2-btn + div.depth2-content > ul.depth3-ul > li > a.depth3-btn."""
rows = []
container = soup.select_one('ul.depth1-ul')
if not container:
return rows
for top_li in container.find_all('li', recursive=False):
a1 = top_li.find('a', class_='depth1-btn', recursive=False)
D = clean_text(a1.get_text()) if a1 else ''
D_href = extract_href(a1) if a1 else ''
d2_ul = top_li.select_one('ul.depth2-ul')
if not d2_ul:
rows.append({'D': D, 'E': '', 'href': D_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
for d2_li in d2_ul.find_all('li', recursive=False):
d2_a = d2_li.find('a', class_='depth2-btn', recursive=False)
if not d2_a:
continue
E = clean_text(d2_a.get_text())
E_href = extract_href(d2_a)
d3_ul = d2_li.select_one('ul.depth3-ul')
if not d3_ul:
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
for d3_li in d3_ul.find_all('li', recursive=False):
d3_a = d3_li.find('a', class_='depth3-btn', recursive=False)
if not d3_a:
continue
F = clean_text(d3_a.get_text())
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(d3_a),
'G': '', 'H': '', 'I': '', 'J': ''})
return rows
def parse_cheongyang_depth1(soup, base):
"""청양군: ul.depth1_ul > li > a.th_1st + div.item > ul.depth2_ul > li > a + ul.depth3_ul > li > a (예산군 형)."""
# 청양군은 예산군과 동일한 e-Gov GNB 구조 — yesan 파서 재사용
return parse_yesan_depth(soup, base)
def parse_asan_gnb(soup, base):
"""아산시: div#mGnb-anchor{n}.gnb-sub-list > ul > li > a.gnb-sub-trigger + ul.sub-ul > li > a.subm."""
rows = []
# 대분류 이름 — mobile-nav의 gnb-main-trigger 텍스트 + href(#mGnb-anchorN) 매핑
nav_main = soup.select_one('nav#mobile-nav')
main_categories = [] # [(D, anchor_id)]
if nav_main:
for trig in nav_main.find_all(['a', 'button'], class_='gnb-main-trigger'):
text = clean_text(trig.get_text())
target = trig.get('href') or trig.get('data-target') or ''
if target.startswith('#mGnb-anchor'):
main_categories.append((text, target.lstrip('#')))
# Fallback: sections without name mapping
if not main_categories:
for sec in soup.select('div[id^=mGnb-anchor]'):
main_categories.append((sec.get('id'), sec.get('id')))
for D, anchor_id in main_categories:
section = soup.find('div', id=anchor_id)
if not section:
continue
# section > ul > li > a.gnb-sub-trigger + ul.sub-ul > li > a.subm
for ul in section.find_all('ul', recursive=False):
for li in ul.find_all('li', recursive=False):
a_E = li.find('a', class_='gnb-sub-trigger', recursive=False) or li.find('a', recursive=False)
if not a_E:
continue
E = clean_text(a_E.get_text())
E_href = extract_href(a_E)
sub_ul = li.find('ul', class_='sub-ul', recursive=False) or li.find('ul', recursive=False)
if not sub_ul:
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
for sub_li in sub_ul.find_all('li', recursive=False):
sub_a = sub_li.find('a', recursive=False)
if not sub_a:
continue
F = clean_text(sub_a.get_text())
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(sub_a),
'G': '', 'H': '', 'I': '', 'J': ''})
return rows
def parse_buyeo_topmenu(soup, base):
"""부여군: ul#tm > li.th1 > a.th1_lnk (D) + div.summry > ul.th2 > li > a.th2_lnk (E) + ul.th3 > li > a (F)."""
rows = []
tm = soup.select_one('ul#tm')
if not tm:
return rows
for top_li in tm.find_all('li', class_=re.compile(r'th1?'), recursive=False):
a1 = top_li.find('a', class_='th1_lnk', recursive=False)
if not a1:
a1 = top_li.find('a', recursive=False)
if not a1:
continue
D = clean_text(a1.get_text())
D_href = extract_href(a1)
# Find ul.th2 inside div.summry
summry = top_li.find('div', class_=re.compile(r'summry'), recursive=False)
th2 = (summry.find('ul', class_='th2') if summry else None) or top_li.find('ul', class_='th2')
if not th2:
rows.append({'D': D, 'E': '', 'href': D_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
for li2 in th2.find_all('li', recursive=False):
a2 = li2.find('a', class_='th2_lnk', recursive=False) or li2.find('a', recursive=False)
if not a2:
continue
E = clean_text(a2.get_text())
E_href = extract_href(a2)
th3 = li2.find('ul', class_='th3', recursive=False)
if not th3:
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
continue
rows.append({'D': D, 'E': E, 'href': E_href,
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
for li3 in th3.find_all('li', recursive=False):
a3 = li3.find('a', recursive=False)
if not a3:
continue
F = clean_text(a3.get_text())
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(a3),
'G': '', 'H': '', 'I': '', 'J': ''})
return rows
# ================================================================
# 사이트 설정
# ================================================================
SITES = [
{
'idx': 5, 'name': '당진시', 'base': 'https://www.dangjin.go.kr',
'sitemap': 'https://www.dangjin.go.kr/kor/sitemap_11.do',
'sheet': '05_당진시', 'parser': parse_amThum,
'domain': 'dangjin.go.kr',
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\5.당진시',
},
{
'idx': 6, 'name': '보령시', 'base': 'https://www.brcn.go.kr',
'sitemap': 'https://www.brcn.go.kr/kor/sitemap_11.do',
'sheet': '06_보령시', 'parser': parse_ul_sitemap_h4,
'domain': 'brcn.go.kr',
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\6.보령시',
},
{
'idx': 7, 'name': '부여군', 'base': 'https://www.buyeo.go.kr',
'sitemap': 'https://www.buyeo.go.kr/html/kr/',
'sheet': '07_부여군', 'parser': parse_buyeo_topmenu,
'domain': 'buyeo.go.kr',
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\7.부여군',
},
{
'idx': 8, 'name': '서산시', 'base': 'https://www.seosan.go.kr',
'sitemap': 'https://www.seosan.go.kr/www/index.do',
'sheet': '08_서산시', 'parser': parse_seosan_top_menu,
'domain': 'seosan.go.kr',
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\8.서산시',
},
{
'idx': 9, 'name': '서천군', 'base': 'https://www.seocheon.go.kr',
'sitemap': 'https://www.seocheon.go.kr/kor/sitemap_11.do',
'sheet': '09_서천군', 'parser': parse_ul_sitemap_h4,
'domain': 'seocheon.go.kr',
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\9.서천군',
},
{
'idx': 10, 'name': '아산시', 'base': 'https://www.asan.go.kr',
'sitemap': 'https://www.asan.go.kr/main/',
'sheet': '10_아산시', 'parser': parse_asan_gnb,
'domain': 'asan.go.kr',
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\10.아산시',
},
{
'idx': 11, 'name': '예산군', 'base': 'https://www.yesan.go.kr',
'sitemap': 'https://www.yesan.go.kr/kor/sitemap.do',
'sheet': '11_예산군', 'parser': parse_yesan_depth,
'domain': 'yesan.go.kr',
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\11.예산군',
},
{
'idx': 12, 'name': '천안시', 'base': 'https://www.cheonan.go.kr',
'sitemap': 'https://www.cheonan.go.kr/kor/sitemap.do',
'sheet': '12_천안시', 'parser': parse_cheonan_depth,
'domain': 'cheonan.go.kr',
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\12.천안시',
},
{
'idx': 13, 'name': '청양군', 'base': 'https://www.cheongyang.go.kr',
'sitemap': 'https://www.cheongyang.go.kr/kor/sitemap_11.do',
'sheet': '13_청양군', 'parser': parse_cheongyang_depth1,
'domain': 'cheongyang.go.kr',
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\13.청양군',
},
{
'idx': 14, 'name': '태안군', 'base': 'https://www.taean.go.kr',
'sitemap': 'https://www.taean.go.kr/kor/sitemap_11.do',
'sheet': '14_태안군', 'parser': parse_amThum,
'domain': 'taean.go.kr',
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\14.태안군',
},
{
'idx': 15, 'name': '홍성군', 'base': 'https://www.hongseong.go.kr',
'sitemap': 'https://www.hongseong.go.kr/kor/sitemap.do',
'sheet': '15_홍성군',
'parser': lambda soup, base: parse_dl_dt_dd(soup, base, use_onclick=True),
'domain': 'hongseong.go.kr',
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\15.홍성군',
},
]
# ================================================================
# 엑셀 생성 (공통)
# ================================================================
def write_excel(site, raw_rows):
name = site['name']
base = site['base']
domain = site['domain']
output = f"{site['folder']}\\{os.path.basename(os.path.dirname(site['folder']))}_{name}.xlsx"
def abs_url(href):
if not href:
return ''
href = href.strip()
if href.startswith(('javascript:', '#')):
return ''
if href.startswith(('http://', 'https://')):
return href
return urljoin(base + '/', href)
def is_external(url):
return url.startswith(('http://', 'https://')) and domain not in url
# 부모-자식 URL 중복 제거
final_rows = []
i = 0
removed = 0
while i < len(raw_rows):
row = raw_rows[i]
if (i + 1 < len(raw_rows)
and row.get('G', '') == ''
and raw_rows[i + 1].get('D') == row.get('D')
and raw_rows[i + 1].get('E') == row.get('E')
and raw_rows[i + 1].get('F') == row.get('F')
and raw_rows[i + 1].get('G', '') != ''
and raw_rows[i + 1].get('href') == row.get('href')):
removed += 1
i += 1
continue
final_rows.append(row)
i += 1
print(f' [{name}] 원본 {len(raw_rows)} → 중복 {removed} → 최종 {len(final_rows)}')
if not final_rows:
print(f' [{name}] !! 행 0개 — 파서 점검 필요. 엑셀 생성 스킵.')
return False
shutil.copy(TEMPLATE, output)
wb = openpyxl.load_workbook(output)
ws = wb.active
ws.title = site['sheet']
HEADER = {'B1:R1', 'S1:W1', 'Y1:AA1'}
for rng in [str(m) for m in ws.merged_cells.ranges if str(m) not in HEADER]:
ws.unmerge_cells(rng)
for row in ws.iter_rows(min_row=3, max_row=ws.max_row, min_col=1, max_col=ws.max_column):
for cell in row:
cell.value = None
START = 3
template_r = 3
cur_max = ws.max_row
for idx, item in enumerate(final_rows, start=START):
if idx > cur_max:
for c in range(1, ws.max_column + 1):
srcc = ws.cell(template_r, c)
tgt = ws.cell(idx, c)
if srcc.has_style:
tgt.font = copy(srcc.font)
tgt.fill = copy(srcc.fill)
tgt.border = copy(srcc.border)
tgt.alignment = copy(srcc.alignment)
tgt.number_format = srcc.number_format
tgt.protection = copy(srcc.protection)
url = abs_url(item.get('href', ''))
ws.cell(idx, 2).value = idx - 2
ws.cell(idx, 3).value = name
ws.cell(idx, 4).value = item.get('D', '')
ws.cell(idx, 5).value = item.get('E', '')
ws.cell(idx, 6).value = item.get('F', '')
ws.cell(idx, 7).value = item.get('G', '')
ws.cell(idx, 8).value = item.get('H', '')
ws.cell(idx, 9).value = item.get('I', '')
ws.cell(idx, 10).value = item.get('J', '')
ws.cell(idx, 11).value = url
if is_external(url):
ws.cell(idx, 19).value = '외부링크'
END = START + len(final_rows) - 1
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
def merge_runs(col_letter, col_idx, group_cols=()):
runs = []
cur_val = ws.cell(START, col_idx).value
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
run_start = START
for r in range(START + 1, END + 1):
v = ws.cell(r, col_idx).value
g = tuple(ws.cell(r, gg).value for gg in group_cols)
if v == cur_val and g == cur_grp:
continue
if cur_val not in (None, '') and r - 1 > run_start:
runs.append((run_start, r - 1))
cur_val, cur_grp, run_start = v, g, r
if cur_val not in (None, '') and END > run_start:
runs.append((run_start, END))
for s, e in runs:
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
ws.cell(s, col_idx).alignment = center
return len(runs)
# 병합 순서 F→E→D (D를 먼저 병합하면 2행부터 D=None이 되어 E 그룹키가 깨짐)
n_f = merge_runs('F', 6, group_cols=(4, 5))
n_e = merge_runs('E', 5, group_cols=(4,))
n_d = merge_runs('D', 4)
for r in range(START, END + 1):
for c in (4, 5, 6, 7):
if ws.cell(r, c).value is not None:
ws.cell(r, c).alignment = center
for r in range(1, END + 1):
ws.row_dimensions[r].height = 15
link_n = 0
for r in range(START, END + 1):
cell = ws.cell(r, 11)
u = cell.value
if u and isinstance(u, str) and u.startswith(('http://', 'https://')):
cell.hyperlink = u
old = cell.font
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
bold=old.bold, italic=old.italic,
color='0000FF', underline='single')
cell.alignment = left
link_n += 1
wb.save(output)
ext_n = sum(1 for r in range(START, END + 1) if ws.cell(r, 19).value == '외부링크')
print(f' [{name}] 병합 D:{n_d} E:{n_e} F:{n_f} | 외부링크 {ext_n} | K링크 {link_n}{output}')
return True
def main():
targets = sys.argv[1:] if len(sys.argv) > 1 else [s['name'] for s in SITES]
for site in SITES:
if site['name'] not in targets:
continue
print(f"\n{'='*60}\n[{site['idx']}.{site['name']}] {site['sitemap']}\n{'='*60}")
try:
html = fetch_html(site['sitemap'])
soup = BeautifulSoup(html, 'html.parser')
raw_rows = site['parser'](soup, site['base'])
write_excel(site, raw_rows)
except Exception as e:
print(f' [{site["name"]}] !! 실패: {e}')
import traceback
traceback.print_exc()
if __name__ == '__main__':
main()

View File

@ -0,0 +1,378 @@
"""충청남도 11개 시·군 Phase 2~4 일괄 처리 (L/M/N/O/P/Q).
매뉴얼: D:\\01.프로젝트\\DB수집\\사이트맵_수집_매뉴얼.md
입력: 폴더의 {기관명}.xlsx
처리: K열 URL 접근 L(게시판형태), M(수량), N(저작물 유형), O/P/Q(공공누리)
"""
import re
import sys
import time
import warnings
from urllib.parse import urljoin
from concurrent.futures import ThreadPoolExecutor, as_completed
import openpyxl
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
# ================================================================
# 정규식
# ================================================================
TOTAL_PAT = re.compile(r'\s*(?:게시물\s*)?(\d[\d,]*)\s*(?:건|개|page|페이지)', re.I)
TOTAL_PAT_LOOSE = re.compile(r'(?:전체|총)\s*[:\-]?\s*(\d[\d,]*)\s*건', re.I)
KOGL_IMG_PAT = re.compile(r'(?:new_)?img_opentype(\d{2})\.png', re.I)
KOGL_LINK_PAT = re.compile(r'kogl\.or\.kr/info/licenseType(\d)', re.I)
YOUTUBE_PAT = re.compile(r'(?:youtube\.com|youtu\.be)', re.I)
VIDEO_EXT = re.compile(r'\.(mp4|webm|mov|avi)(?:\?|$)', re.I)
# ================================================================
# 사이트 설정
# ================================================================
SITES = {
'논산시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\4.논산시\충청남도_논산시.xlsx',
'body_sel': ['#txt', '#contents', 'main']},
'당진시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\5.당진시\충청남도_당진시.xlsx',
'body_sel': ['#txt', '#contents', 'main']},
'보령시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\6.보령시\충청남도_보령시.xlsx',
'body_sel': ['#txt', '#contents', 'main']},
'부여군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\7.부여군\충청남도_부여군.xlsx',
'body_sel': ['#txt', '#contents', 'main']},
'서산시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\8.서산시\충청남도_서산시.xlsx',
'body_sel': ['#contents', '#txt', 'main']},
'서천군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\9.서천군\충청남도_서천군.xlsx',
'body_sel': ['#txt', '#contents', 'main']},
'아산시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\10.아산시\충청남도_아산시.xlsx',
'body_sel': ['.contents', '#contents', '#txt', 'main']},
'예산군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\11.예산군\충청남도_예산군.xlsx',
'body_sel': ['#txt', '#contents', 'main']},
'천안시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\12.천안시\충청남도_천안시.xlsx',
'body_sel': ['#txt', '#contents', 'main']},
'청양군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\13.청양군\충청남도_청양군.xlsx',
'body_sel': ['#txt', '#contents', 'main']},
'태안군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\14.태안군\충청남도_태안군.xlsx',
'body_sel': ['#txt', '#contents', 'main']},
'홍성군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\15.홍성군\충청남도_홍성군.xlsx',
'body_sel': ['#txt', '#contents', 'main']},
}
# ================================================================
# 크롤링 공통 함수
# ================================================================
def fetch(url, timeout=12):
try:
r = requests.get(url, headers=H, timeout=timeout, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
if r.status_code == 200:
return BeautifulSoup(r.text, 'html.parser')
except Exception:
pass
return None
def get_body(soup, selectors):
for sel in selectors:
el = soup.select_one(sel)
if el:
return el
return soup
def detect_form(body):
"""L 판별. 페이징·검색·총건수 3요소 중 하나라도 있으면 '게시판'."""
has_paging = bool(body.select('.pagination, .paging, nav.paging, .page_nav'))
text_inputs = [i for i in body.find_all('input')
if (i.get('type') or 'text').lower() in ('text', 'search')]
has_search = len(text_inputs) >= 1
txt = body.get_text(' ', strip=True)
m = TOTAL_PAT.search(txt)
if not m:
m = TOTAL_PAT_LOOSE.search(txt)
total = None
if m:
digits = m.group(1).replace(',', '')
if digits.isdigit():
total = int(digits)
is_board = has_paging or has_search or (total is not None)
if is_board:
return '게시판', total if total is not None else 0
return '페이지', 1
DETAIL_PAT = re.compile(r'(mode=V|view\.do|bbtSn=|seqRepeat=|nttId=|articleNo=|boardSeq=)', re.I)
def extract_detail_urls(body, base_url, limit=5):
urls = []
seen = set()
for a in body.find_all('a', href=True):
h = a['href']
if not h or h.startswith('#'):
continue
if DETAIL_PAT.search(h):
full = urljoin(base_url, h)
if full not in seen:
seen.add(full)
urls.append(full)
if len(urls) >= limit:
break
# Also try fn_search_detail JS pattern (공주시 형)
if len(urls) < limit:
for a in body.find_all('a'):
onclick = a.get('onclick', '')
m = re.search(r"fn_(?:search_)?detail\(['\"]([^'\"]+)['\"]", onclick)
if m:
ntt_id = m.group(1)
# Build URL by replacing list.do with view.do?nttId=
view_url = re.sub(r'list\.do[^\'"]*', f'view.do?nttId={ntt_id}', base_url)
if view_url not in seen:
seen.add(view_url)
urls.append(view_url)
if len(urls) >= limit:
break
return urls
def detect_media(body):
has_text = len(body.get_text(strip=True)) > 30
has_image = False
for img in body.find_all('img'):
src = img.get('src', '')
if KOGL_IMG_PAT.search(src):
continue
if not src:
continue
has_image = True
break
has_video = False
for iframe in body.find_all('iframe'):
if YOUTUBE_PAT.search(iframe.get('src', '')):
has_video = True
break
if not has_video:
for a in body.find_all('a', href=True):
if YOUTUBE_PAT.search(a['href']):
has_video = True
break
if not has_video and body.find_all('video'):
has_video = True
if not has_video and VIDEO_EXT.search(str(body)):
has_video = True
return has_image, has_video, has_text
def n_string(has_text, has_image, has_video):
parts = []
if has_text:
parts.append('어문')
if has_image:
parts.append('이미지')
if has_video:
parts.append('영상')
return ','.join(parts) if parts else '없음'
def img_has_valid_anchor(img):
p = img.parent
while p is not None:
if p.name == 'a':
href = p.get('href', '')
if href and not href.startswith('#') and not href.lower().startswith('javascript:'):
return True
return False
p = p.parent
return False
def detect_kogl(body):
types = set()
q_any_y = False
q_any_n = False
for a in body.find_all('a', href=True):
m = KOGL_LINK_PAT.search(a['href'])
if m:
types.add(int(m.group(1)))
q_any_y = True
for img in body.find_all('img'):
src = img.get('src', '')
m = KOGL_IMG_PAT.search(src)
if m:
types.add(int(m.group(1)))
if img_has_valid_anchor(img):
q_any_y = True
else:
q_any_n = True
for el in body.find_all(style=True):
m = KOGL_IMG_PAT.search(el.get('style', ''))
if m:
types.add(int(m.group(1)))
q_any_n = True
if not types:
return set(), None
return types, ('Y' if q_any_y else 'N')
def process_row(url, body_selectors):
out = {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
soup = fetch(url)
if soup is None:
out['note'] = '접근 실패'
return out
body = get_body(soup, body_selectors)
form, count = detect_form(body)
out['L'] = form
out['M'] = count if form == '게시판' else 1
has_img, has_vid, has_txt = detect_media(body)
types_main, q_main = detect_kogl(body)
P = '게시판' if types_main else ''
types_all = set(types_main)
q_flags = []
if q_main:
q_flags.append(q_main)
if form == '게시판':
detail_urls = extract_detail_urls(body, url, limit=5)
for du in detail_urls:
d_soup = fetch(du, timeout=10)
if not d_soup:
continue
d_body = get_body(d_soup, body_selectors)
di, dv, dt = detect_media(d_body)
has_img = has_img or di
has_vid = has_vid or dv
has_txt = has_txt or dt
dt_types, dt_q = detect_kogl(d_body)
if dt_types and not types_main and not P:
P = '게시물'
types_all |= dt_types
if dt_q:
q_flags.append(dt_q)
out['N'] = n_string(has_txt, has_img, has_vid)
if not types_all:
out['O'] = '미부착'
else:
sorted_types = sorted(types_all)
out['O'] = ','.join(f'{n}유형' for n in sorted_types)
out['P'] = P if P else '게시판'
out['Q'] = 'Y' if 'Y' in q_flags else 'N'
return out
# ================================================================
# 사이트 단위 실행
# ================================================================
def run_site(name, xlsx, body_selectors, workers=10):
print(f'\n{"="*60}\n[{name}] {xlsx}\n{"="*60}')
wb = openpyxl.load_workbook(xlsx)
ws = wb.active
START = 3
# 진짜 데이터 행만 가져옴 (B열에 순번 있어야)
END = START - 1
for r in range(START, ws.max_row + 1):
if ws.cell(r, 2).value is None:
break
END = r
tasks = []
for r in range(START, END + 1):
url = ws.cell(r, 11).value
is_ext = (ws.cell(r, 19).value == '외부링크')
tasks.append((r, url, is_ext))
n_ext = sum(1 for t in tasks if t[2])
print(f' 처리 대상: 총 {len(tasks)}행 (외부링크 {n_ext})')
t0 = time.time()
results = {}
def worker(task):
row, url, is_ext = task
if is_ext:
return row, {'L': '사이트', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
if not url or not isinstance(url, str):
return row, {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': 'URL 없음'}
return row, process_row(url, body_selectors)
done = 0
with ThreadPoolExecutor(max_workers=workers) as ex:
futs = [ex.submit(worker, t) for t in tasks]
for fut in as_completed(futs):
row, res = fut.result()
results[row] = res
done += 1
if done % 50 == 0 or done == len(tasks):
print(f' 진행 {done}/{len(tasks)} ({time.time()-t0:.0f}s)')
for r in range(START, END + 1):
res = results.get(r, {})
if not res:
continue
if res.get('L'):
ws.cell(r, 12).value = res['L']
if res.get('M') != '':
ws.cell(r, 13).value = res['M']
if res.get('N'):
ws.cell(r, 14).value = res['N']
if res.get('O'):
ws.cell(r, 15).value = res['O']
if res.get('P'):
ws.cell(r, 16).value = res['P']
if res.get('Q'):
ws.cell(r, 17).value = res['Q']
if res.get('note'):
existing = ws.cell(r, 19).value
if not existing:
ws.cell(r, 19).value = res['note']
wb.save(xlsx)
forms = {}
attach = {'미부착': 0, '부착': 0, '기타': 0}
q_dist = {'Y': 0, 'N': 0, '': 0}
for r, res in results.items():
forms[res.get('L', '')] = forms.get(res.get('L', ''), 0) + 1
o = res.get('O', '')
if o == '미부착':
attach['미부착'] += 1
elif o and '유형' in o:
attach['부착'] += 1
else:
attach['기타'] += 1
q = res.get('Q', '')
q_dist[q] = q_dist.get(q, 0) + 1
print(f' [{name}] L분포 {forms} | 부착 {attach} | Q {q_dist} | 시간 {time.time()-t0:.0f}s')
def main():
targets = sys.argv[1:] if len(sys.argv) > 1 else list(SITES.keys())
total_t0 = time.time()
for name in targets:
if name not in SITES:
print(f' 알 수 없음: {name}')
continue
cfg = SITES[name]
try:
run_site(name, cfg['xlsx'], cfg['body_sel'])
except Exception as e:
print(f' [{name}] 실패: {e}')
import traceback
traceback.print_exc()
print(f'\n총 소요: {time.time()-total_t0:.0f}s')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,127 @@
# -*- coding: utf-8 -*-
"""규칙 B 전 사이트 적용: F 하나에 G가 여러 개이고 그 G들이 전부 L=페이지면
1행으로 합치고(G/H/I/J 비움), K= G URL, 수량(M)=G들의 M . (공주시와 동일)
안전장치:
· G 하나라도 페이지가 아니면(게시판/사이트) 합치지 않음.
· 합칠 공공누리: O유형=그룹 합집합(숫자 오름차순,콤마,공백없음), Q=하나라도 Y면 Y,
P=게시판 우선(없으면 게시물). 부착 정보 손실 방지하며 합침.
· G가 잎일 때만(H/I/J 비어있음). 깊은 중첩은 건드리지 않음.
공주시(완료)·계룡시(검수완료) 제외.
사용: python -X utf8 _collapse_all.py dry [기관...]
python -X utf8 _collapse_all.py run [기관...]
"""
import os, re, sys, shutil, importlib.util, warnings
import openpyxl
warnings.filterwarnings('ignore')
TYPE_PAT = re.compile(r'(\d)\s*유형')
def merge_kogl(grp):
"""그룹 G행들의 O/P/Q 통합. 반환 (O, P, Q)."""
types = set()
for g in grp:
o = g['vals'].get(15)
if isinstance(o, str):
for m in TYPE_PAT.finditer(o):
types.add(int(m.group(1)))
if types:
O = ','.join(f'{n}유형' for n in sorted(types))
else:
O = '미부착'
Ps = [g['vals'].get(16) for g in grp]
P = '게시판' if '게시판' in Ps else ('게시물' if '게시물' in Ps else None)
Qs = [g['vals'].get(17) for g in grp]
Q = 'Y' if 'Y' in Qs else ('N' if 'N' in Qs else None)
return O, (P if O != '미부착' else None), (Q if O != '미부착' else None)
spec = importlib.util.spec_from_file_location('dd', __file__.replace('_collapse_all.py', '_dedup_all.py'))
dd = importlib.util.module_from_spec(spec); spec.loader.exec_module(dd)
EXCLUDE = {'공주시', '계룡시'}
def mval(x):
try:
return int(x)
except Exception:
return 1
def collapse(rows):
out, groups, held = [], [], []
i, n = 0, len(rows)
while i < n:
v = rows[i]['vals']
D, E, F, G = v.get(4), v.get(5), v.get(6), v.get(7)
if F not in (None, '') and G not in (None, ''):
j = i
while (j < n and rows[j]['vals'].get(4) == D and rows[j]['vals'].get(5) == E
and rows[j]['vals'].get(6) == F and rows[j]['vals'].get(7) not in (None, '')):
j += 1
grp = rows[i:j]
all_page = all(g['vals'].get(12) == '페이지' for g in grp)
leaf = all(all(g['vals'].get(c) in (None, '') for c in (8, 9, 10)) for g in grp)
has_attach = any(g['vals'].get(15) not in (None, '', '미부착') for g in grp)
if len(grp) >= 2 and all_page and leaf:
msum = sum(mval(g['vals'].get(13)) for g in grp)
O, P, Q = merge_kogl(grp)
first = dict(grp[0]); nv = dict(first['vals'])
for c in (7, 8, 9, 10):
nv[c] = None
nv[13] = msum
nv[15] = O; nv[16] = P; nv[17] = Q
first['vals'] = nv
out.append(first)
groups.append({'E': E, 'F': F, 'n': len(grp), 'sum': msum, 'O': O})
if has_attach:
held.append({'E': E, 'F': F, 'O': O})
i = j
continue
out.extend(grp); i = j; continue
out.append(rows[i]); i += 1
return out, groups, held
def main():
mode = sys.argv[1] if len(sys.argv) > 1 else 'dry'
only = set(a for a in sys.argv[2:] if not a.startswith('--'))
sites = [s for s in dd.SITES if s[2] not in EXCLUDE and (not only or s[2] in only)]
print(f'대상 기관: {len(sites)}개 (공주·계룡 제외) 모드={mode}\n')
print(f'{"기관":<8}{"현재":>6}{"합친그룹":>7}{"제거행":>6}{"→남음":>7}{"부착합침":>7}')
print('-' * 55)
grand_g = grand_r = grand_h = 0
attach_detail = []
for prov, idx, name in sites:
xp = dd.xpath(prov, idx, name)
if not os.path.exists(xp):
print(f'{name:<8} 엑셀 없음'); continue
wb = openpyxl.load_workbook(xp); ws = wb.active
rows = dd.load_flat(ws)
out, groups, held = collapse(rows)
removed = len(rows) - len(out)
grand_g += len(groups); grand_r += removed; grand_h += len(held)
print(f'{name:<8}{len(rows):>6}{len(groups):>7}{removed:>6}{len(out):>7}{len(held):>7}')
for h in held:
attach_detail.append((name, h['E'], h['F'], h['O']))
if mode == 'run' and groups:
bak = xp.replace('.xlsx', '_backup_collapse전.xlsx')
shutil.copy(xp, bak)
dd.write_back(ws, out)
try:
wb.save(xp)
except PermissionError:
wb.save(xp.replace('.xlsx', '_LP.xlsx')); print(f' !! {name} 잠김→_LP')
print('-' * 55)
print(f'합계: 합친그룹 {grand_g} / 제거행 {grand_r} / 부착합침 {grand_h} ({mode})')
if attach_detail:
print(f'\n=== 부착 유형 합쳐진 그룹 {len(attach_detail)}개 (O 결과 확인용) ===')
for nm, E, F, O in attach_detail:
print(f' [{nm}] {E} > {F} → O={O}')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,126 @@
# -*- coding: utf-8 -*-
"""중분류/소분류 '랜딩행' 제거 — 매뉴얼 1-5 확장(③ K동일 조건 제거).
_dedup_all.py 부모행 URL == 자식 URL 때만 부모행을 삭제한다().
eGov '누리집지도'(고창·임실·정읍·진안 ) 중분류 랜딩( 002000000 '민원안내')
소분류(002007000) URL이 달라 걸리고 랜딩행이 남아 있다.
사용자 요청(2026-05-31): 조건을 빼고, 자식을 거느린 부모 랜딩행을 전부 삭제.
· 부모행의 가장 깊은 카테고리 컬럼 lc, lc+1 빈칸(부모는 깊이의 잎이 아님)
· 직후 자식행이 D~lc 동일 & lc+1 채워짐 부모행 삭제(자식이 병합으로 라벨 승계)
URL 무관. 카테고리 텍스트(D/E/F) 재병합으로 보존.
평탄화/재병합/순번/하이퍼링크는 _dedup_all 재사용. 컬럼(B~AA, L~Q) 보존.
사용:
python -X utf8 _collapse_landing.py dry <기관명...|경로>
python -X utf8 _collapse_landing.py run <기관명...|경로>
--maxlc N : lc<=N 깊이까지만 삭제(기본 9=I, D~I 랜딩 모두). E만 원하면 --maxlc 5.
"""
import os, sys, shutil, importlib.util, warnings
import openpyxl
warnings.filterwarnings('ignore')
spec = importlib.util.spec_from_file_location('dd', __file__.replace('_collapse_landing.py', '_dedup_all.py'))
dd = importlib.util.module_from_spec(spec); spec.loader.exec_module(dd)
CAT_COLS = dd.CAT_COLS # D E F G H I J = 4..10
EXCLUDE = {'계룡시'} # 검수완료
def is_safe_shell(v):
"""순수 메뉴 셸인가: L=페이지(또는 빈칸) & M<=1 & 공공누리 미부착."""
L = v.get(12); M = v.get(13); O = v.get(15)
if isinstance(O, str) and '유형' in O:
return False
if L not in (None, '', '페이지'):
return False
if isinstance(M, (int, float)) and M and M > 1:
return False
return True
def collapse(rows, maxlc=9, safe=False):
"""랜딩행 제거. safe=True면 순수 셸(데이터 무보유)만. 반환 (남은행, 제거목록)."""
out, removed = [], []
i = 0
while i < len(rows):
v = rows[i]['vals']
if i + 1 < len(rows):
nv = rows[i + 1]['vals']
lc = dd.leaf_depth(v) # 부모의 가장 깊은 카테고리 컬럼
if 4 <= lc <= maxlc:
same_upper = all((v.get(c) or '') == (nv.get(c) or '')
for c in CAT_COLS if c <= lc)
child_next = nv.get(lc + 1) not in (None, '')
parent_is_landing = v.get(lc + 1) in (None, '')
if safe and not is_safe_shell(v):
same_upper = False # 데이터 보유 랜딩은 보존
if same_upper and child_next and parent_is_landing:
removed.append({'src': rows[i]['src'], 'lc': lc,
'label': v.get(lc), 'child': nv.get(lc + 1),
'url': v.get(11), 'curl': nv.get(11)})
i += 1
continue
out.append(rows[i]); i += 1
return out, removed
LV = {4: 'D', 5: 'E', 6: 'F', 7: 'G', 8: 'H', 9: 'I', 10: 'J'}
def resolve(args):
paths = []
for a in args:
if a.lower().endswith('.xlsx') or os.path.sep in a:
paths.append((os.path.splitext(os.path.basename(a))[0], a))
names = [a for a in args if not (a.lower().endswith('.xlsx') or os.path.sep in a)]
for prov, idx, name in dd.SITES:
if names and name in names:
paths.append((name, dd.xpath(prov, idx, name)))
return paths
def main():
mode = sys.argv[1] if len(sys.argv) > 1 else 'dry'
rest = [a for a in sys.argv[2:] if not a.startswith('--')]
maxlc = 9
if '--maxlc' in sys.argv:
maxlc = int(sys.argv[sys.argv.index('--maxlc') + 1])
safe = '--safe' in sys.argv
targets = resolve(rest)
if not targets:
print('대상 없음.'); return
grand = 0
for name, xp in targets:
if name in EXCLUDE:
print(f'{name}: 검수완료 제외'); continue
if not os.path.exists(xp):
print(f'{name}: 엑셀 없음'); continue
wb = openpyxl.load_workbook(xp); ws = wb.active
rows = dd.load_flat(ws)
out, removed = collapse(rows, maxlc, safe)
grand += len(removed)
bylv = {}
for r in removed:
bylv[r['lc']] = bylv.get(r['lc'], 0) + 1
lvstr = ' '.join(f'{LV[k]}:{v}' for k, v in sorted(bylv.items()))
print(f'\n=== {name} === {len(rows)}{len(out)} (제거 {len(removed)}) [{lvstr}]')
for r in removed[:6]:
print(f" [{LV[r['lc']]}] {r['label']} ({r['url']}) → 자식 첫행 {r['child']} ({r['curl']})")
if len(removed) > 6:
print(f' ... 외 {len(removed)-6}')
if mode == 'run' and removed:
bak = xp.replace('.xlsx', '_backup_랜딩제거전.xlsx')
shutil.copy(xp, bak)
dd.write_back(ws, out)
try:
wb.save(xp)
except PermissionError:
wb.save(xp.replace('.xlsx', '_LP.xlsx')); print(' !! 잠김→_LP')
else:
print(f' 저장 완료. 백업: {os.path.basename(bak)}')
print(f'\n총 제거 대상: {grand}행 ({mode}, maxlc={maxlc})')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,166 @@
# -*- coding: utf-8 -*-
"""규칙 B 재귀판: 탭 레벨(G→H→I…)에서 '자식이 전부 페이지'면 1행으로 합치고 M=합.
메뉴 카테고리(D·E·F) 보존 합치기는 lc>=7(G 이하 레벨)에서만 수행.
레벨 lc 에서:
부모(D..lc-1) 동일 + lc 비어있지 않음 으로 형제 묶음 묶음이
· 전부 L=페이지, · lc보다 깊은 카테고리열 모두 비어있음()
이면 1행으로 합침: lc 이하 카테고리열 비움, M=, K= URL,
공공누리 O=유형 합집합(오름차순,콤마,공백없음)/P=게시판우선/Q=하나라도 Y.
깊은얕은 순으로 반복(안정될 때까지) GH 합친 FG까지 자연 연쇄.
평탄화/재병합(D/E/F)/순번/하이퍼링크는 _dedup_all.write_back 재사용.
사용:
python -X utf8 _collapse_recursive.py dry <엑셀경로 | 기관명...>
python -X utf8 _collapse_recursive.py run <엑셀경로 | 기관명...>
(기관명 없이 경로 1개만 줘도 )
"""
import os, re, sys, shutil, importlib.util, warnings
import openpyxl
warnings.filterwarnings('ignore')
spec = importlib.util.spec_from_file_location('dd', __file__.replace('_collapse_recursive.py', '_dedup_all.py'))
dd = importlib.util.module_from_spec(spec); spec.loader.exec_module(dd)
CAT = [4, 5, 6, 7, 8, 9, 10] # D E F G H I J
COLLAPSE_LEVELS = [10, 9, 8, 7] # 탭 레벨만(깊은→얕은). E(6)·D(5)는 제외=카테고리 보존
TYPE_PAT = re.compile(r'(\d)\s*유형')
EXCLUDE = {'계룡시'} # 검수완료
def mval(x):
try:
return int(x)
except Exception:
return 1
def merge_kogl(grp):
types = set()
for g in grp:
o = g['vals'].get(15)
if isinstance(o, str):
for m in TYPE_PAT.finditer(o):
types.add(int(m.group(1)))
if types:
O = ','.join(f'{n}유형' for n in sorted(types))
else:
O = '미부착'
Ps = [g['vals'].get(16) for g in grp]
P = '게시판' if '게시판' in Ps else ('게시물' if '게시물' in Ps else None)
Qs = [g['vals'].get(17) for g in grp]
Q = 'Y' if 'Y' in Qs else ('N' if 'N' in Qs else None)
return O, (P if O != '미부착' else None), (Q if O != '미부착' else None)
def collapse_level(rows, lc):
parent = [c for c in CAT if c < lc]
deeper = [c for c in CAT if c > lc]
out, groups = [], []
i, n = 0, len(rows)
while i < n:
v = rows[i]['vals']
if v.get(lc) not in (None, ''):
j = i
while (j < n and all(rows[j]['vals'].get(c) == v.get(c) for c in parent)
and rows[j]['vals'].get(lc) not in (None, '')):
j += 1
grp = rows[i:j]
all_page = all(g['vals'].get(12) == '페이지' for g in grp)
leaf = all(all(g['vals'].get(c) in (None, '') for c in deeper) for g in grp)
has_attach = any(g['vals'].get(15) not in (None, '', '미부착') for g in grp)
if len(grp) >= 2 and all_page and leaf:
msum = sum(mval(g['vals'].get(13)) for g in grp)
O, P, Q = merge_kogl(grp)
first = dict(grp[0]); nv = dict(first['vals'])
for c in [lc] + deeper:
nv[c] = None
nv[13] = msum; nv[15] = O; nv[16] = P; nv[17] = Q
first['vals'] = nv
out.append(first)
groups.append({'lc': lc, 'path': [v.get(c) for c in parent],
'label': v.get(lc), 'n': len(grp),
'ms': [mval(g['vals'].get(13)) for g in grp],
'sum': msum, 'O': O, 'attach': has_attach,
'names': [g['vals'].get(lc) for g in grp]})
i = j; continue
out.extend(grp); i = j; continue
out.append(rows[i]); i += 1
return out, groups
def collapse_all_levels(rows):
all_groups = []
changed = True
while changed:
changed = False
for lc in COLLAPSE_LEVELS:
rows, groups = collapse_level(rows, lc)
if groups:
changed = True
all_groups.extend(groups)
return rows, all_groups
def resolve_paths(args):
"""경로 또는 기관명 목록 → [(name, xlsx경로)]"""
paths = []
for a in args:
if a.lower().endswith('.xlsx') or os.path.sep in a:
paths.append((os.path.splitext(os.path.basename(a))[0], a))
names = [a for a in args if not (a.lower().endswith('.xlsx') or os.path.sep in a)]
for prov, idx, name in dd.SITES:
if names and name not in names:
continue
if not names:
continue
paths.append((name, dd.xpath(prov, idx, name)))
return paths
LV = {7: 'F→G', 8: 'G→H', 9: 'H→I', 10: 'I→J'}
def main():
mode = sys.argv[1] if len(sys.argv) > 1 else 'dry'
rest = [a for a in sys.argv[2:] if not a.startswith('--')]
targets = resolve_paths(rest)
if not targets:
print('대상 없음. 경로 또는 기관명을 지정하세요.'); return
for name, xp in targets:
if name in EXCLUDE:
print(f'{name}: 검수완료 제외'); continue
if not os.path.exists(xp):
print(f'{name}: 엑셀 없음 ({xp})'); continue
wb = openpyxl.load_workbook(xp); ws = wb.active
rows = dd.load_flat(ws)
out, groups = collapse_all_levels(rows)
removed = len(rows) - len(out)
bylv = {}
for g in groups:
bylv.setdefault(g['lc'], 0)
bylv[g['lc']] += 1
lvstr = ' '.join(f"{LV.get(k,k)}:{v}" for k, v in sorted(bylv.items()))
print(f'\n=== {name} === {len(rows)}{len(out)} (제거 {removed}) [{lvstr}]')
for g in groups:
ms = '+'.join(str(m) for m in g['ms'])
star = ' ★부착' if g['attach'] else ''
print(f" [{LV.get(g['lc'],g['lc'])}] {' > '.join(str(x) for x in g['path'] if x)} > {g['label']}"
f" ({g['n']}개) → M={ms}={g['sum']} O={g['O']}{star}")
if mode == 'run' and groups:
bak = xp.replace('.xlsx', '_backup_재귀합치기전.xlsx')
shutil.copy(xp, bak)
dd.write_back(ws, out)
try:
wb.save(xp)
except PermissionError:
wb.save(xp.replace('.xlsx', '_LP.xlsx')); print(f' !! {name} 잠김→_LP')
else:
print(f' 저장 완료. 백업: {os.path.basename(bak)}')
elif mode != 'run':
print(' (DRY)')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,160 @@
대상 기관: 40개 (공주·계룡 제외) 모드=run
기관 현재 합친그룹 제거행 →남음 부착합침
-------------------------------------------------------
금산군 533 54 167 366 0
논산시 758 75 155 603 0
당진시 309 0 0 309 0
보령시 547 0 0 547 0
부여군 222 0 0 222 0
서산시 306 0 0 306 0
서천군 275 0 0 275 0
아산시 261 0 0 261 0
예산군 566 51 95 471 0
천안시 640 55 152 488 0
청양군 257 0 0 257 0
태안군 154 0 0 154 0
홍성군 524 47 138 386 0
괴산군 240 2 6 234 0
단양군 428 20 109 319 10
보은군 512 26 102 410 0
영동군 764 25 66 698 0
옥천군 521 37 165 356 0
음성군 632 45 244 388 0
제천시 594 41 224 370 1
증평군 404 0 0 404 0
진천군 297 0 0 297 0
청주시 406 0 0 406 0
충주시 376 0 0 376 0
고창군 379 19 53 326 0
군산시 637 9 35 602 9
김제시 452 0 0 452 0
남원시 331 38 90 241 38
무주군 83 0 0 83 0
부안군 201 0 0 201 0
순창군 452 28 152 300 28
완주군 328 0 0 328 0
익산시 535 26 131 404 26
임실군 313 16 127 186 0
장수군 78 0 0 78 0
전주시 263 0 0 263 0
정읍시 245 0 0 245 0
진안군 307 0 0 307 0
서귀포시 388 1 1 387 0
제주시 310 0 0 310 0
-------------------------------------------------------
합계: 합친그룹 615 / 제거행 2212 / 부착합침 112 (run)
=== 부착 유형 합쳐진 그룹 112개 (O 결과 확인용) ===
[단양군] 상징 > 브랜드 → O=2유형
[단양군] 군정안내 > 청사배치도 → O=2유형
[단양군] 아동복지 > 아동복지정책 → O=2유형
[단양군] 청소년복지 > 청소년복지시설 → O=2유형
[단양군] 여성·가족 복지 > 여성복지시설 → O=2유형
[단양군] 여성·가족 복지 > 여성·가족복지정책 → O=2유형
[단양군] 노인복지 > 노인복지시설 → O=2유형
[단양군] 노인복지 > 노인복지정책 → O=2유형
[단양군] 장애인복지 > 장애인복지시설 → O=2유형
[단양군] 장애인복지 > 장애인복지정책 → O=2유형
[제천시] 제천시소개 > 공공저작물 → O=1유형,2유형,3유형
[군산시] 산업인프라 > 항만/여객/공항/철도/컨벤션 → O=4유형
[군산시] 농업/축산업 > 농산물 유통 → O=4유형
[군산시] 건설 > 자전거 → O=4유형
[군산시] 에너지 > 태양광 → O=4유형
[군산시] 에너지 > 가스/석유 → O=4유형
[군산시] 군산시 소개 > 행정구역/행정지도 → O=4유형
[군산시] 군산시 소개 > 자매결연/국제협력 도시 → O=4유형
[군산시] 군산시 소개 > 군산의 상징 → O=4유형
[군산시] 시청안내 > 전화번호안내 → O=4유형
[남원시] 행복민원실 > 민원발급안내 → O=4유형
[남원시] 부동산정보 > 지적재조사 → O=4유형
[남원시] 행정정보공개 > 정보공개제도안내 → O=4유형
[남원시] 예산공개 > 2026년도 → O=4유형
[남원시] 예산공개 > 2025년도 → O=4유형
[남원시] 예산공개 > 2024년도 → O=4유형
[남원시] 예산공개 > 2023년도 → O=4유형
[남원시] 예산공개 > 2022년도 → O=4유형
[남원시] 예산공개 > 2021년도 → O=4유형
[남원시] 예산공개 > 2020년도 → O=4유형
[남원시] 예산공개 > 2019년도 → O=4유형
[남원시] 예산공개 > 2018년도 → O=4유형
[남원시] 예산공개 > 2017년도 → O=4유형
[남원시] 예산공개 > 2016년도 → O=4유형
[남원시] 예산공개 > 2015년도 → O=4유형
[남원시] 예산공개 > 2014년도 → O=4유형
[남원시] 예산공개 > 2013년도 → O=4유형
[남원시] 예산공개 > 2012년도 → O=4유형
[남원시] 예산공개 > 2011년도 → O=4유형
[남원시] 예산공개 > 2009년도 → O=4유형
[남원시] 재정공시 > 2025년 → O=4유형
[남원시] 재정공시 > 2024년 → O=4유형
[남원시] 재정공시 > 2023년 → O=4유형
[남원시] 재정공시 > 2022년 → O=4유형
[남원시] 재정공시 > 2021년 → O=4유형
[남원시] 재정공시 > 2020년 → O=4유형
[남원시] 재정공시 > 2019년 → O=4유형
[남원시] 재정공시 > 2018년 → O=4유형
[남원시] 재정공시 > 2017년 → O=4유형
[남원시] 재정공시 > 2016년 → O=4유형
[남원시] 재정공시 > 2015년 → O=4유형
[남원시] 열린행정 > 행정서비스헌장 → O=4유형
[남원시] 남원의역사 > 시대별 → O=4유형
[남원시] 남원의상징 > 기관상징 → O=4유형
[남원시] 자매 우호 결연 > 자매결연 → O=4유형
[남원시] 자매 우호 결연 > 우호결연 → O=4유형
[남원시] 시청안내 > 청사(시설물)안내 → O=4유형
[남원시] 시청안내 > 찾아오시는길 → O=4유형
[순창군] 주요민원안내 > 자동차등록안내 → O=4유형
[순창군] 인허가제도 > 축사시설 허가 → O=4유형
[순창군] 인허가제도 > 정보통신설비 허가 → O=4유형
[순창군] 여성/가족 > 여성친화도시 → O=4유형
[순창군] 여성/가족 > 아동/청소년 → O=4유형
[순창군] 행복순창! 인구정책 길라잡이 > 임신/출산 → O=4유형
[순창군] 행복순창! 인구정책 길라잡이 > 영유아 → O=4유형
[순창군] 행복순창! 인구정책 길라잡이 > 아동/청소년 → O=4유형
[순창군] 행복순창! 인구정책 길라잡이 > 결혼 → O=4유형
[순창군] 행복순창! 인구정책 길라잡이 > 청년/중장년 → O=4유형
[순창군] 행복순창! 인구정책 길라잡이 > 노년 → O=4유형
[순창군] 행복순창! 인구정책 길라잡이 > 다문화 → O=4유형
[순창군] 행복순창! 인구정책 길라잡이 > 공통사항 → O=4유형
[순창군] 군민생활 > 복지·주거 → O=4유형
[순창군] 군민생활 > 식품·공중위생 → O=4유형
[순창군] 군민생활 > 군민안전보험 → O=4유형
[순창군] 군민생활 > 교육·인재양성 → O=4유형
[순창군] 군민생활 > 일자리·고용 → O=4유형
[순창군] 경제·산업 > 농공단지 → O=4유형
[순창군] 상하수도·환경 > 상수도 → O=4유형
[순창군] 상하수도·환경 > 생활환경정보 → O=4유형
[순창군] 농촌개발·축산·산림 > 일반농산어촌개발사업 → O=4유형
[순창군] 농촌개발·축산·산림 > 농촌체험마을 → O=4유형
[순창군] 농촌개발·축산·산림 > 축산·산림 → O=4유형
[순창군] 재난·안전·민방위 > 민방위 → O=4유형
[순창군] 순창 장류산업 지역특구 > 주요사업 → O=4유형
[순창군] 순창 장류산업 지역특구 > 주요시설 → O=4유형
[순창군] 재정정보 > 공유재산공개 → O=4유형
[익산시] 기부美 > 명예의전당 → O=4유형
[익산시] 교통 > 시내버스 → O=4유형
[익산시] 교통 > 주정차 → O=4유형
[익산시] 복지 > 아동/청소년 → O=4유형
[익산시] 복지 > 국민생활복지 → O=4유형
[익산시] 복지 > 의료급여제도 → O=4유형
[익산시] 위생 > 식품위생 → O=4유형
[익산시] 환경/보건 > 정신재활시설 → O=4유형
[익산시] 상하수도사업단 > 사업단소개 → O=4유형
[익산시] 상하수도사업단 > 하수도 → O=4유형
[익산시] 상하수도사업단 > 주요시책사업 → O=4유형
[익산시] 산업/경제 > 소상공인 정책 → O=4유형
[익산시] 전입혜택 > 학생지원 → O=4유형
[익산시] 전입혜택 > 일반시민 → O=4유형
[익산시] 전입혜택 > 여성보육 → O=4유형
[익산시] 전입혜택 > 전입청년 → O=4유형
[익산시] 재난안전 > 대피장소 현황 → O=4유형
[익산시] 익산의 상징 > 익산브랜드 → O=4유형
[익산시] 익산의 역사 > 역사와 유래 → O=4유형
[익산시] 익산의 역사 > 시대별 익산의 역사 → O=4유형
[익산시] 익산의통계 > 도표로 보는 통계 → O=4유형
[익산시] 시청안내 > 조직도 → O=4유형
[익산시] 시청안내 > 부서소개 → O=4유형
[익산시] 시청안내 > 시청사소개 → O=4유형
[익산시] 시청안내 > 전화번호 → O=4유형
[익산시] 상호결연·우호도시 > 국제도시 → O=4유형

215
_스크립트/_dedup_all.py Normal file
View File

@ -0,0 +1,215 @@
# -*- coding: utf-8 -*-
"""부모-자식 URL 중복 제거 — 전 깊이 일반화 (매뉴얼 1-5 확장).
기존 phase1 dedup은 FG 깊이만 처리. 서산시처럼 중분류(E) 클릭 랜딩이고
URL이 소분류(F) 같으면(EF 중복) 지워졌다. 도구는 깊이 무관:
부모행의 가장 깊은 카테고리 컬럼이 lc 이고, 직후 자식행이
· D~lc 까지 값이 모두 동일, · lc+1 컬럼이 채워짐, · K(URL) 부모와 '정확히' 동일
이면 부모행을 삭제(자식이 카테고리 텍스트를 병합으로 승계). URL이 다르면 보존.
평탄화재병합(D/E/F)순번/행높이/하이퍼링크 재설정은 _tab_expand 동일 로직.
모든 컬럼(B~AA, L~Q Phase2~4 데이터 포함) 보존.
사용:
python -X utf8 _dedup_all.py dry [기관...] # 제거대상 집계만(읽기전용)
python -X utf8 _dedup_all.py run [기관...] # 백업(*_backup_dedup전.xlsx) 후 제거
"""
import os
import sys
import shutil
import warnings
from copy import copy
import openpyxl
from openpyxl.styles import Alignment, Font
warnings.filterwarnings('ignore')
MAXCOL = 27
CAT_COLS = [4, 5, 6, 7, 8, 9, 10] # D E F G H I J
HEADER_MERGES = {'B1:R1', 'S1:W1', 'Y1:AA1'}
HERE = os.path.dirname(os.path.abspath(__file__))
MAP = os.path.join(os.path.dirname(HERE), '작업파일', '광역_사이트맵')
# 전 42개 시·군 (광역, idx, 기관)
SITES = (
[('충청남도', i, n) for i, n in [
(1, '계룡시'), (2, '공주시'), (3, '금산군'), (4, '논산시'), (5, '당진시'),
(6, '보령시'), (7, '부여군'), (8, '서산시'), (9, '서천군'), (10, '아산시'),
(11, '예산군'), (12, '천안시'), (13, '청양군'), (14, '태안군'), (15, '홍성군')]]
+ [('충청북도', i, n) for i, n in [
(1, '괴산군'), (2, '단양군'), (3, '보은군'), (4, '영동군'), (5, '옥천군'),
(6, '음성군'), (7, '제천시'), (8, '증평군'), (9, '진천군'), (10, '청주시'), (11, '충주시')]]
+ [('전북특별자치도', i, n) for i, n in [
(1, '고창군'), (2, '군산시'), (3, '김제시'), (4, '남원시'), (5, '무주군'),
(6, '부안군'), (7, '순창군'), (8, '완주군'), (9, '익산시'), (10, '임실군'),
(11, '장수군'), (12, '전주시'), (13, '정읍시'), (14, '진안군')]]
+ [('제주특별자치도', i, n) for i, n in [(1, '서귀포시'), (2, '제주시')]]
)
def xpath(prov, idx, name):
return os.path.join(MAP, prov, f'{idx}.{name}', f'{prov}_{name}.xlsx')
def load_flat(ws):
for mr in list(ws.merged_cells.ranges):
s = str(mr)
if s in HEADER_MERGES:
continue
top = ws.cell(mr.min_row, mr.min_col).value
ws.unmerge_cells(s)
for rr in range(mr.min_row, mr.max_row + 1):
for cc in range(mr.min_col, mr.max_col + 1):
if ws.cell(rr, cc).value in (None, ''):
ws.cell(rr, cc).value = top
rows = []
for r in range(3, ws.max_row + 1):
if all(ws.cell(r, c).value in (None, '') for c in range(2, MAXCOL + 1)):
continue
vals = {c: ws.cell(r, c).value for c in range(1, MAXCOL + 1)}
styles = {}
for c in range(1, MAXCOL + 1):
sc = ws.cell(r, c)
styles[c] = (copy(sc.font), copy(sc.fill), copy(sc.border),
copy(sc.alignment), sc.number_format, copy(sc.protection))
hl = ws.cell(r, 11).hyperlink
rows.append({'src': r, 'vals': vals, 'styles': styles,
'hyperlink': hl.target if hl else None})
return rows
def leaf_depth(vals):
deep = 4
for c in CAT_COLS:
if vals.get(c) not in (None, ''):
deep = c
return deep
def dedup(rows):
"""전 깊이 부모-자식 URL 중복 제거. 반환: (남은행, 제거목록)."""
out, removed = [], []
i = 0
while i < len(rows):
v = rows[i]['vals']
if i + 1 < len(rows):
nv = rows[i + 1]['vals']
lc = leaf_depth(v)
ku, kc = v.get(11), nv.get(11)
if lc < 10 and isinstance(ku, str) and ku.startswith('http') and ku == kc:
same_upper = all((v.get(c) or '') == (nv.get(c) or '')
for c in CAT_COLS if c <= lc)
child_next = nv.get(lc + 1) not in (None, '')
if same_upper and child_next:
removed.append({'src': rows[i]['src'], 'lc': lc,
'E': v.get(5), 'F_child': nv.get(lc + 1),
'url': ku})
i += 1
continue
out.append(rows[i])
i += 1
return out, removed
def write_back(ws, out_rows):
"""남은 행으로 데이터영역 재작성 + D/E/F 재병합 + 순번/행높이/하이퍼링크."""
for r in range(3, ws.max_row + 1):
for c in range(1, MAXCOL + 1):
ws.cell(r, c).value = None
ws.cell(r, c).hyperlink = None
START = 3
for idx, orow in enumerate(out_rows):
r = START + idx
sty = orow['styles']
for c in range(1, MAXCOL + 1):
cell = ws.cell(r, c)
f, fl, bd, al, nf, pr = sty[c]
cell.font = copy(f); cell.fill = copy(fl); cell.border = copy(bd)
cell.alignment = copy(al); cell.number_format = nf; cell.protection = copy(pr)
v = orow['vals']
for c in range(1, MAXCOL + 1):
ws.cell(r, c).value = v.get(c)
ws.cell(r, 2).value = idx + 1
END = START + len(out_rows) - 1
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
def merge_runs(letter, col, group_cols=()):
cur = ws.cell(START, col).value
grp = tuple(ws.cell(START, g).value for g in group_cols)
run = START
runs = []
for r in range(START + 1, END + 1):
val = ws.cell(r, col).value
g = tuple(ws.cell(r, gg).value for gg in group_cols)
if val == cur and g == grp:
continue
if cur not in (None, '') and r - 1 > run:
runs.append((run, r - 1))
cur, grp, run = val, g, r
if cur not in (None, '') and END > run:
runs.append((run, END))
for s, e in runs:
ws.merge_cells(f'{letter}{s}:{letter}{e}')
ws.cell(s, col).alignment = center
merge_runs('F', 6, (4, 5))
merge_runs('E', 5, (4,))
merge_runs('D', 4)
for r in range(START, END + 1):
for c in (4, 5, 6, 7):
if ws.cell(r, c).value is not None:
ws.cell(r, c).alignment = center
for r in range(1, END + 1):
ws.row_dimensions[r].height = 15
for r in range(START, END + 1):
cell = ws.cell(r, 11)
u = cell.value
if isinstance(u, str) and u.startswith(('http://', 'https://')):
cell.hyperlink = u
old = cell.font
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
bold=old.bold, italic=old.italic,
color='0000FF', underline='single')
cell.alignment = left
def main():
mode = sys.argv[1] if len(sys.argv) > 1 else 'dry'
only = set(a for a in sys.argv[2:] if not a.startswith('--'))
sites = [s for s in SITES if not only or s[2] in only]
grand = 0
print(f'{"기관":<8}{"현재행":>6}{"제거":>5}{"→남음":>7} 유형(중분류 예시)')
print('-' * 80)
for prov, idx, name in sites:
xp = xpath(prov, idx, name)
if not os.path.exists(xp):
print(f'{name:<8} 엑셀 없음')
continue
wb = openpyxl.load_workbook(xp)
ws = wb.active
rows = load_flat(ws)
out, removed = dedup(rows)
grand += len(removed)
ex = ''
if removed:
sample = removed[0]
ex = f"E={sample['E']} = {sample['F_child']}"
print(f'{name:<8}{len(rows):>6}{len(removed):>5}{len(out):>7} {ex}')
if mode == 'run' and removed:
bak = xp.replace('.xlsx', '_backup_dedup전.xlsx')
shutil.copy(xp, bak)
write_back(ws, out)
try:
wb.save(xp)
except PermissionError:
wb.save(xp.replace('.xlsx', '_LP.xlsx'))
print(f' !! 원본 잠김 → _LP.xlsx 저장')
print('-' * 80)
print(f'총 제거 대상: {grand}행 ({mode})')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,23 @@
"""충청북도 시·군 목록을 마스터 엑셀에서 추출."""
import openpyxl
from pathlib import Path
ROOT = Path(r'D:\01.프로젝트\DB수집')
XLSX = ROOT / '붙임1_공공저작물 개방 대상기관(1160개) 실태조사 목록(신유형개방지원사업) 양식_260518.xlsx'
wb = openpyxl.load_workbook(XLSX, data_only=True)
ws = wb['1단계_홈페이지']
targets = ['괴산군', '단양군', '보은군', '영동군', '옥천군', '음성군',
'제천시', '증평군', '진천군', '청주시', '충주시']
for name in targets:
for r in range(5, ws.max_row + 1):
cell_name = ws.cell(row=r, column=4).value
if cell_name and ('충청북도' in str(cell_name) or '충북' in str(cell_name)) and name in str(cell_name):
daebun = ws.cell(row=r, column=2).value
jungbun = ws.cell(row=r, column=3).value
url = ws.cell(row=r, column=5).value
print(f'Row {r}: 대={daebun!r} 중={jungbun!r} 기관명={cell_name!r} URL={url!r}')
break
else:
print(f'NOT FOUND: {name}')

View File

@ -0,0 +1,27 @@
"""Find how Chungnam cities are categorized in the master file."""
import openpyxl
from pathlib import Path
ROOT = Path(r'D:\01.프로젝트\DB수집')
XLSX = ROOT / '붙임1_공공저작물 개방 대상기관(1160개) 실태조사 목록(신유형개방지원사업) 양식_260518.xlsx'
wb = openpyxl.load_workbook(XLSX, data_only=True)
ws = wb['1단계_홈페이지']
# Known 충청남도 cities to search
targets = ['계룡시', '공주시', '금산군', '논산시', '당진시', '보령시',
'부여군', '서산시', '서천군', '아산시', '예산군', '천안시',
'청양군', '태안군', '홍성군']
for name in targets:
found = False
for r in range(5, ws.max_row + 1):
cell_name = ws.cell(row=r, column=4).value
if cell_name and name in str(cell_name):
daebun = ws.cell(row=r, column=2).value
jungbun = ws.cell(row=r, column=3).value
url = ws.cell(row=r, column=5).value
print(f'Row {r}: 대={daebun!r} 중={jungbun!r} 기관명={cell_name!r} URL={url!r}')
found = True
break
if not found:
print(f'NOT FOUND: {name}')

View File

@ -0,0 +1,52 @@
# -*- coding: utf-8 -*-
"""미검수 27곳 C열(사이트명)을 "광역 기관"(공백구분)으로 채움. 검수완료 제외.
사용: python -X utf8 _fill_C_province.py [--write]
"""
import sys, os, re, shutil, importlib.util
import openpyxl
HERE=os.path.dirname(os.path.abspath(__file__))
MODULES=['_chungnam_phase234_all.py','_chungbuk_phase234_all.py','_jeonbuk_phase234_all.py']
TARGETS={'서산시','아산시','천안시','청양군','홍성군',
'괴산군','단양군','보은군','영동군','옥천군','음성군','제천시','증평군','진천군','청주시','충주시',
'부안군','순창군','완주군','익산시','임실군','장수군','전주시','정읍시','진안군','서귀포시','제주시'}
def load():
sites={}
for f in MODULES:
sp=importlib.util.spec_from_file_location(f[:-3],os.path.join(HERE,f));m=importlib.util.module_from_spec(sp);sp.loader.exec_module(m)
sites.update(m.SITES)
return sites
def province_of(path):
m=re.search(r'광역_사이트맵[\\/]([^\\/]+)[\\/]',path)
return m.group(1) if m else '?'
def main():
write='--write' in sys.argv
sites=load(); locked=[]; done=[]
for org in sorted(TARGETS):
if org not in sites: print(org,'SITES없음'); continue
xlsx=sites[org]['xlsx']
prov=province_of(xlsx); newC=f'{prov} {org}'
wb=openpyxl.load_workbook(xlsx); ws=wb.active
ch=0
for r in range(3,ws.max_row+1):
content = any(ws.cell(r,c).value not in (None,'') for c in (2,4,5,6,7,8,9,10,11,12))
cur=ws.cell(r,3).value
if (cur not in (None,'')) or content:
if cur!=newC:
if write: ws.cell(r,3).value=newC
ch+=1
if write and ch:
try:
bak=xlsx.replace('.xlsx','_backup_C광역전.xlsx')
if not os.path.exists(bak): shutil.copy(xlsx,bak)
wb.save(xlsx)
except PermissionError:
locked.append(org); print(f'{org:6} ❌파일열림 (미적용)'); continue
done.append((org,newC,ch))
print(f'{org:6}"{newC}" 변경 {ch}{"[적용]" if write else "[DRY]"}')
if locked: print('\n⚠️ 파일열림으로 미적용:',', '.join(locked))
if __name__=='__main__': main()

View File

@ -0,0 +1,116 @@
# -*- coding: utf-8 -*-
"""E열(대분류) 병합 누락 보정 — 재크롤링 없이 기존 엑셀의 D/E/F 병합만 바로잡는다.
원인: phase1이 D를 먼저 병합 블록 2행부터 D=None E 병합(같은 D 안에서만)
D블록 행에서 끊겨 E가 단독으로 남음.
해결: 현재 병합에서 복원 D/E/F 병합 해제 행에 채움
F E D 순서로 재병합(컬럼을 병합하면 컬럼 2행부터 None이 되므로 순서가 중요).
데이터(D/E/F 텍스트) 그대로. 헤더 병합(B1:R1 ) 건드리지 않음.
파일별 *_backup_emerge전.xlsx 백업.
"""
import sys, os, glob, shutil, warnings, openpyxl
from openpyxl.styles import Alignment
from openpyxl.utils import get_column_letter
sys.stdout.reconfigure(encoding='utf-8')
warnings.filterwarnings('ignore')
START = 3
COLS = (4, 5, 6) # D, E, F
CENTER = Alignment(horizontal='center', vertical='center', wrap_text=False)
def filled_values(ws, col):
"""현재 병합 상태에서 각 행의 실제 값을 복원(병합 top값으로 채움)."""
end = ws.max_row
vals = {r: ws.cell(r, col).value for r in range(START, end + 1)}
for mr in ws.merged_cells.ranges:
if mr.min_col == col == mr.max_col and mr.min_row >= START:
top = ws.cell(mr.min_row, col).value
for r in range(mr.min_row, mr.max_row + 1):
vals[r] = top
return vals, end
def merge_runs(ws, col, end, group_keys, fills):
"""fills[col][r] 기준으로 연속 동일 구간 병합. group_keys: 상위 컬럼 튜플."""
letter = get_column_letter(col)
runs = []
def key(r):
return tuple(fills[g][r] for g in group_keys)
cur_val = fills[col][START]
cur_grp = key(START)
run_start = START
for r in range(START + 1, end + 1):
v, g = fills[col][r], key(r)
if v == cur_val and g == cur_grp:
continue
if cur_val not in (None, '') and r - 1 > run_start:
runs.append((run_start, r - 1))
cur_val, cur_grp, run_start = v, g, r
if cur_val not in (None, '') and end > run_start:
runs.append((run_start, end))
for s, e in runs:
ws.merge_cells(f'{letter}{s}:{letter}{e}')
ws.cell(s, col).alignment = CENTER
return len(runs)
def fix(xlsx):
wb = openpyxl.load_workbook(xlsx)
ws = wb.active
# 1) 값 복원
fills, end = {}, None
for c in COLS:
fills[c], end = filled_values(ws, c)
# 2) D/E/F 데이터 병합만 해제 (단일컬럼·row>=START)
to_unmerge = [str(mr) for mr in ws.merged_cells.ranges
if mr.min_col == mr.max_col and mr.min_col in COLS and mr.min_row >= START]
for rng in to_unmerge:
ws.unmerge_cells(rng)
# 3) 전 행에 값 다시 기입
for c in COLS:
for r in range(START, end + 1):
ws.cell(r, c).value = fills[c][r]
# 4) F → E → D 순서 재병합
n_f = merge_runs(ws, 6, end, (4, 5), fills)
n_e = merge_runs(ws, 5, end, (4,), fills)
n_d = merge_runs(ws, 4, end, (), fills)
return wb, (n_d, n_e, n_f, len(to_unmerge))
def main():
sel = [a for a in sys.argv[1:] if not a.startswith('-')]
files = []
for region in ['충청남도', '충청북도']:
for d in sorted(glob.glob(region + r'\*')):
if not os.path.isdir(d):
continue
base = os.path.basename(d)
nm = base.split('.', 1)[1] if '.' in base else base
prov = os.path.basename(os.path.dirname(d))
f = os.path.join(d, f'{prov}_{nm}.xlsx')
if os.path.exists(f) and (not sel or nm in sel):
files.append((nm, f))
print(f'대상 {len(files)}개: {[n for n, _ in files]}')
ok = []
for nm, f in files:
try:
bak = f.replace('.xlsx', '') + '_backup_emerge전.xlsx'
if not os.path.exists(bak):
shutil.copy2(f, bak)
wb, (nd, ne, nf, un) = fix(f)
wb.save(f)
print(f'{nm:<6} 병합해제 {un:>3} → 재병합 D:{nd} E:{ne} F:{nf}')
ok.append(nm)
except PermissionError:
print(f' ⚠️ {nm:<6} 파일 잠김(열려있음) → 건너뜀')
except Exception as e:
import traceback; traceback.print_exc()
print(f' !! {nm:<6} 오류: {e}')
print(f'\n완료 {len(ok)}/{len(files)}')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,91 @@
# -*- coding: utf-8 -*-
"""금산군 메뉴 랜딩행 M = 서브트리 전체 leaf 합 (2026-05-31, 사용자 규칙)
여성가족=10 패턴: 메뉴(6자리 prefix) 모든 자식페이지 X01..X0N에 대해
leaf = 라이브 ui-nav_tabs 개수 ( 없으면 1)
M(랜딩행) = Σ leaf. 메뉴의 자식 게시판이 있으면 SKIP(분리 필요, 수동).
대상: 랜딩행 >= 143(여성가족) 이고 prefix의 시트행이 1(sheet=1) 메뉴.
sheet>1(이미 G확장됨: 군민복지050307·하수처리050601·주민참여060603) SKIP.
"""
import sys, re, json, openpyxl
import urllib.request, ssl
from bs4 import BeautifulSoup
from concurrent.futures import ThreadPoolExecutor
PATH = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\충청남도_금산군.xlsx'
WRITE = '--write' in sys.argv
ctx = ssl.create_default_context(); ctx.check_hostname = False; ctx.verify_mode = ssl.CERT_NONE
def fetch(u):
req = urllib.request.Request(u, headers={'User-Agent': 'Mozilla/5.0'})
try:
return urllib.request.urlopen(req, context=ctx, timeout=15).read().decode('utf-8', 'replace')
except Exception:
return None
def analyze(code):
sub = 'sub06' if code.startswith('06') else 'sub05'
h = fetch(f'https://www.geumsan.go.kr/kr/html/{sub}/{code}.html')
if h is None:
return None
soup = BeautifulSoup(h, 'html.parser')
t = soup.select_one('.location_wrap li:last-child, h2.h2')
title = t.get_text(strip=True) if t else '?'
ntab = len(soup.select('ul.ui-nav_tabs a.ui-tabs_link'))
board = bool(soup.select('table.bbs,.board_list,.bbs_list,.pagination')) or bool(re.search(r'\s*[\d,]+\s*건', str(soup)))
return dict(code=code, title=title, ntab=ntab, board=board, leaf=max(1, ntab))
wb = openpyxl.load_workbook(PATH); ws = wb.active
# prefix -> sheet rows (sub05/sub06 8-digit)
pref_rows = {}
for r in range(3, ws.max_row + 1):
k = ws.cell(r, 11).value
if not k:
continue
m = re.search(r'/sub0[56]/(\d{6})(\d{2})\.html', str(k))
if m:
pref_rows.setdefault(m.group(1), []).append((r, m.group(0)))
targets = {}
for pfx, rows in pref_rows.items():
landing = min(r for r, _ in rows)
if landing < 143:
continue # 여성가족(143) 위는 제외
if len(rows) != 1:
print(f'SKIP {pfx} (sheet={len(rows)} rows={[r for r,_ in rows]}) — 이미 확장됨')
continue
targets[pfx] = landing
print(f'대상 sheet=1 메뉴: {len(targets)}')
plan = []
for pfx, landing in sorted(targets.items(), key=lambda x: x[1]):
codes = [f'{pfx}{xx:02d}' for xx in range(1, 16)]
kids = [k for k in ThreadPoolExecutor(8).map(analyze, codes) if k]
leaf_sum = sum(k['leaf'] for k in kids)
nboard = sum(1 for k in kids if k['board'])
cur_m = ws.cell(landing, 13).value
cur_f = ws.cell(landing, 6).value or ws.cell(landing, 7).value
flag = ' ⚠GESIPAN' if nboard else ''
plan.append(dict(pfx=pfx, row=landing, m_old=cur_m, m_new=leaf_sum, nkids=len(kids), nboard=nboard, kids=kids))
print(f'row{landing} {pfx} [{cur_f}] M {cur_m}->{leaf_sum} (자식{len(kids)},board{nboard}){flag}')
if nboard:
for k in kids:
print(f' {k["code"]} leaf{k["leaf"]} board={k["board"]} {k["title"][:20]}')
# 적용: 게시판 없는 메뉴만 M 갱신
applied = 0
for p in plan:
if p['nboard'] == 0:
ws.cell(p['row'], 13).value = p['m_new']
applied += 1
else:
print(f' ⚠ row{p["row"]} {p["pfx"]} 게시판 포함 → M 미적용(수동 검토)')
print(f'적용 대상(게시판없음): {applied}/{len(plan)}')
json.dump([{k: v for k, v in p.items() if k != 'kids'} for p in plan],
open(r'D:\01.프로젝트\DB수집\_temp\_geumsan_subtreeM_plan.json', 'w', encoding='utf8'), ensure_ascii=False)
if WRITE:
wb.save(PATH); print('SAVED')
else:
print('DRY-RUN (--write 로 저장)')

View File

@ -0,0 +1,102 @@
# -*- coding: utf-8 -*-
"""금산군 ui-nav_tabs 재검토 + 군민복지 cross-product 정리 (2026-05-31)
- PHASE1: ui-nav_tabs 인페이지 페이지 1 유지, M=라이브 개수
여성가족 05030301=3, 어르신 05030401=3, 사회복지 05030601=5(무변경),
· 행정복지센터 06X0101 ×10 =2
- PHASE2: 군민복지(05030701) tab-ul type1 4 cross-product(147~159, 13)
4행으로 축소(공지사항/금산군자원봉사센터/복지시설/관련사이트). 151~159 삭제.
"""
import sys, openpyxl
from copy import copy
from openpyxl.utils import range_boundaries
PATH = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\충청남도_금산군.xlsx'
WRITE = '--write' in sys.argv
wb = openpyxl.load_workbook(PATH)
ws = wb.active
MAXC = ws.max_column
# ---------- PHASE 1: M 정규화 (URL 매칭) ----------
MTARGETS = {
'05030301': 3, '05030401': 3, '05030601': 5,
'06010101': 2, '06020101': 2, '06030101': 2, '06040101': 2, '06050101': 2,
'06060101': 2, '06070101': 2, '06080101': 2, '06090101': 2, '06100101': 2,
}
print('=== PHASE1: M 정규화 ===')
for r in range(3, ws.max_row + 1):
k = ws.cell(r, 11).value
if not k:
continue
ks = str(k)
for code, mval in MTARGETS.items():
if code + '.html' in ks:
old_m = ws.cell(r, 13).value
old_l = ws.cell(r, 12).value
ws.cell(r, 13).value = mval
print(f' row{r} {code} L={old_l} M:{old_m}->{mval}')
break
# ---------- PHASE 2: 군민복지 restructure ----------
print('=== PHASE2: 군민복지 정리 ===')
# Step A: 147~150 G<-H, clear H
for r in range(147, 151):
g_old = ws.cell(r, 7).value
h = ws.cell(r, 8).value
ws.cell(r, 7).value = h
ws.cell(r, 8).value = None
print(f' row{r} G:{g_old}->{h} (H clear)')
DEL_START, DEL_COUNT = 151, 9
# Step B: capture & clear merges
merges = [str(mc) for mc in ws.merged_cells.ranges]
for mc in list(ws.merged_cells.ranges):
ws.unmerge_cells(str(mc))
# Step C: manual shift up (delete 151~159)
max_row = ws.max_row
for r in range(DEL_START, max_row - DEL_COUNT + 1):
src = r + DEL_COUNT
for c in range(1, MAXC + 1):
s = ws.cell(src, c); d = ws.cell(r, c)
d.value = s.value
if s.has_style:
d._style = copy(s._style)
d.number_format = s.number_format
d.hyperlink = None
if s.hyperlink:
d.hyperlink = copy(s.hyperlink)
d.hyperlink.ref = d.coordinate
for r in range(max_row - DEL_COUNT + 1, max_row + 1):
for c in range(1, MAXC + 1):
d = ws.cell(r, c); d.value = None; d.hyperlink = None
# Step D: rebuild merges with row mapping
def surv(x):
return x < DEL_START or x >= DEL_START + DEL_COUNT
def mp(x):
return x if x < DEL_START else x - DEL_COUNT
for mc in merges:
c1, r1, c2, r2 = range_boundaries(mc)
s = [x for x in range(r1, r2 + 1) if surv(x)]
if not s:
continue
nr1, nr2 = mp(min(s)), mp(max(s))
if nr1 == nr2 and c1 == c2:
continue # collapsed to single cell, no merge
ws.merge_cells(start_row=nr1, start_column=c1, end_row=nr2, end_column=c2)
# Step E: 순번 재번호 (151행 이후 -9 → 연속성 유지)
for r in range(DEL_START, ws.max_row + 1):
v = ws.cell(r, 2).value
if isinstance(v, int):
ws.cell(r, 2).value = v - DEL_COUNT
print(f' deleted rows {DEL_START}~{DEL_START+DEL_COUNT-1} (9 rows)')
if WRITE:
wb.save(PATH)
print('SAVED', PATH)
else:
print('DRY-RUN (use --write to save)')

View File

@ -0,0 +1,87 @@
# -*- coding: utf-8 -*-
"""공주시 227~414행 재검수 적용.
규칙(226 검수 캘리브레이션 결과):
R1 관광지 dongList(tursmCn): L 페이지->게시판, M=썸네일수
R2 게시판 글수 재계산: M=현재 크롤 글수 (board_count 있는 , 값이 다를 때만)
R3 찾아오시는길/오시는길: N=어문,이미지
DRY 기본. --write 백업 저장.
"""
import sys, json, shutil
import openpyxl
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
EV = r'D:\01.프로젝트\DB수집\_스크립트\_evidence.json'
LO, HI = 227, 414
def main():
write = '--write' in sys.argv
ev = json.load(open(EV, encoding='utf-8'))
wb = openpyxl.load_workbook(XLSX)
ws = wb.active
changes = [] # (row, col, colname, old, new, rule)
drift = [] # 미세 드리프트(원본 유지) 보고용
for r in range(LO, HI + 1):
e = ev.get(str(r))
if not e:
continue
url = (ws.cell(r, 11).value or '')
L = ws.cell(r, 12).value
M = ws.cell(r, 13).value
N = ws.cell(r, 14).value
Gname = ws.cell(r, 7).value or ''
Fname = ws.cell(r, 6).value or ''
name = '%s %s' % (Fname, Gname)
# R1 관광지 dongList
if 'tursmCn' in url and e.get('tursm_count'):
tc = e['tursm_count']
if L != '게시판':
changes.append((r, 12, 'L', L, '게시판', 'R1관광지'))
if M != tc:
changes.append((r, 13, 'M', M, tc, 'R1관광지'))
continue # 관광지는 M 규칙2 적용 안함
# R2 게시판 글수 재계산 (0/빈 값만 복구; 유효 비0값의 미세 드리프트는 원본 유지)
if L == '게시판' and e.get('board_count') is not None:
bc = e['board_count']
if (M in (None, 0, '0')) and bc:
changes.append((r, 13, 'M', M, bc, 'R2게시판글수복구'))
elif M != bc:
drift.append((r, M, bc, ws.cell(r, 7).value or ws.cell(r, 6).value or ''))
# R3 찾아오시는길 N
if ('찾아오시는' in name) or ('오시는길' in name) or ('오시는 길' in name):
if N == '어문':
changes.append((r, 14, 'N', N, '어문,이미지', 'R3오시는길'))
# report
print('=== 제안 변경 %d건 (DRY) ===' % len(changes))
by = {}
for r, c, cn, old, new, rule in changes:
by.setdefault(rule, []).append((r, cn, old, new))
for rule in sorted(by):
print('\n[%s] %d' % (rule, len(by[rule])))
for r, cn, old, new in by[rule]:
g = ws.cell(r, 7).value or ws.cell(r, 6).value or ws.cell(r, 5).value or ''
print(' r%d %s: %r -> %r (%s)' % (r, cn, old, new, g))
if drift:
print('\n[미세 드리프트: 원본 유지, 참고용] %d' % len(drift))
for r, old, new, g in drift:
print(' r%d 게시판글수 원본 %s (현재 크롤 %s, 차이 %+d) %s' % (r, old, new, new - (old or 0), g))
if write:
bak = XLSX.replace('.xlsx', '_backup_재검수227전.xlsx')
shutil.copy(XLSX, bak)
for r, c, cn, old, new, rule in changes:
ws.cell(r, c).value = new
wb.save(XLSX)
print('\n저장 완료. 백업:', bak)
else:
print('\n(DRY — 적용하려면 --write)')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,82 @@
# -*- coding: utf-8 -*-
"""공주시 전용 규칙 B: F 하나에 G가 여러 개이고 그 G들이 전부 L=페이지면
1행으로 합치고(G/H/I/J 비움), K= G의 URL, 수량(M)=G들의 M .
G 하나라도 페이지가 아니면(게시판/사이트 ) 그룹은 합치지 않음.
평탄화/재병합/스타일은 _dedup_all 재사용.
사용: python -X utf8 _gongju_collapse.py [--write]
"""
import sys, shutil, importlib.util, warnings
import openpyxl
warnings.filterwarnings('ignore')
spec = importlib.util.spec_from_file_location('dd', __file__.replace('_gongju_collapse.py', '_dedup_all.py'))
dd = importlib.util.module_from_spec(spec); spec.loader.exec_module(dd)
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
def mval(x):
try:
return int(x)
except Exception:
return 1
def collapse(rows):
out, groups = [], []
i, n = 0, len(rows)
while i < n:
v = rows[i]['vals']
D, E, F, G = v.get(4), v.get(5), v.get(6), v.get(7)
if F not in (None, '') and G not in (None, ''):
j = i
while (j < n and rows[j]['vals'].get(4) == D and rows[j]['vals'].get(5) == E
and rows[j]['vals'].get(6) == F and rows[j]['vals'].get(7) not in (None, '')):
j += 1
grp = rows[i:j]
all_page = all(g['vals'].get(12) == '페이지' for g in grp)
leaf = all(all(g['vals'].get(c) in (None, '') for c in (8, 9, 10)) for g in grp)
if len(grp) >= 2 and all_page and leaf:
msum = sum(mval(g['vals'].get(13)) for g in grp)
first = dict(grp[0]); nv = dict(first['vals'])
for c in (7, 8, 9, 10):
nv[c] = None
nv[13] = msum
first['vals'] = nv
out.append(first)
groups.append({'D': D, 'E': E, 'F': F, 'n': len(grp),
'ms': [mval(g['vals'].get(13)) for g in grp],
'gs': [g['vals'].get(7) for g in grp],
'sum': msum, 'k': nv.get(11)})
i = j
continue
out.extend(grp); i = j; continue
out.append(rows[i]); i += 1
return out, groups
def main():
write = '--write' in sys.argv
wb = openpyxl.load_workbook(XLSX)
ws = wb.active
rows = dd.load_flat(ws)
out, groups = collapse(rows)
removed = len(rows) - len(out)
print(f'합칠 F그룹: {len(groups)}개 / 제거행 {removed} ({len(rows)}{len(out)})\n')
for g in groups:
ms = '+'.join(str(m) for m in g['ms'])
print(f" [{g['E']} > {g['F']}] G {g['n']}개({'/'.join(str(x) for x in g['gs'])[:50]})")
print(f" → 1행, M={ms}={g['sum']}, K={g['k']}")
if write and groups:
bak = XLSX.replace('.xlsx', '_backup_collapse전.xlsx')
shutil.copy(XLSX, bak)
dd.write_back(ws, out)
wb.save(XLSX)
print(f'\n저장 완료. 백업: {bak}')
elif not write:
print('\n(DRY — 실제 적용하려면 --write)')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,120 @@
# -*- coding: utf-8 -*-
"""공주시 227~414행 재검수용 증거 수집(쓰기 없음).
URL을 크롤링해 게시판 글수/관광지 썸네일수/외부링크/공공누리마크/#nav탭수/본문이미지·영상 신호를 수집.
검증용으로 226행까지 검수에서 바뀐 일부도 같이 수집.
출력: _스크립트/_evidence.json + 콘솔 요약
"""
import sys, json, re, warnings
from urllib.parse import urlparse
from concurrent.futures import ThreadPoolExecutor, as_completed
import requests
from bs4 import BeautifulSoup
import openpyxl
warnings.filterwarnings('ignore')
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
H = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/120 Safari/537.36'}
COMMON_IMG = re.compile(r'(flag\.jpg|slogan|/common/|/template/|btn_|/btn|icon|blank|_mark\.png|sns|share|loading)', re.I)
def fetch(row, url):
try:
r = requests.get(url, headers=H, timeout=25, verify=False)
r.encoding = r.apparent_encoding or 'utf-8'
return row, url, r.status_code, r.text
except Exception as e:
return row, url, None, ('ERR:%s' % e)
def analyze(url, html):
ev = {}
host = urlparse(url).netloc
ev['host'] = host
ev['external'] = ('gongju.go.kr' not in host)
if not html or html.startswith('ERR:'):
ev['error'] = html
return ev
s = BeautifulSoup(html, 'html.parser')
# board total count
cnt_el = s.select_one('.program--count strong')
if cnt_el:
m = re.sub(r'[^0-9]', '', cnt_el.get_text())
ev['board_count'] = int(m) if m else None
ev['is_board'] = True
else:
ev['is_board'] = False
# tursmCn thumbnails (관광지)
tt = len(re.findall(r'/thumbnail/tursmCn/', html))
if tt:
ev['tursm_count'] = tt
# content images (static html, excludes template/common) — JS maps missed
cimgs = []
for im in s.find_all('img'):
src = im.get('src') or im.get('data-src') or ''
if src and not COMMON_IMG.search(src):
cimgs.append(src)
ev['content_img'] = len(cimgs)
ev['content_img_sample'] = cimgs[:4]
# video signals
low = html.lower()
ev['video'] = bool(re.search(r'youtube\.com/embed|youtu\.be/|player\.vimeo|<video|\.mp4|data-video', low))
# KOGL / 공공누리
kogl_hits = re.findall(r'(opentype0?[1-4]|kogl[_\-]?[1-4]?|공공누리)', html, re.I)
ev['kogl'] = bool(kogl_hits)
ev['kogl_sample'] = list(dict.fromkeys(kogl_hits))[:6]
# explicit opentype number
ot = re.findall(r'opentype0?([1-4])', html, re.I)
if ot:
ev['kogl_types'] = sorted(set(int(x) for x in ot))
# #nav tab count (규칙 A)
best = 0
for ul in s.find_all('ul'):
cls = ' '.join(ul.get('class') or []).lower()
if 'tab-ul' in cls:
anchors = [a for a in ul.find_all('a')
if (a.get('href') or '').strip().startswith('#') and a.get_text(strip=True)]
if len(anchors) >= 2:
best = max(best, len(anchors))
if best:
ev['nav_tabs'] = best
return ev
def main():
lo, hi = 227, 414
if len(sys.argv) > 2:
lo, hi = int(sys.argv[1]), int(sys.argv[2])
wb = openpyxl.load_workbook(XLSX)
ws = wb.active
targets = []
rowinfo = {}
for r in range(lo, hi + 1):
u = ws.cell(r, 11).value
if isinstance(u, str) and u.strip().startswith('http'):
targets.append((r, u.strip()))
rowinfo[r] = {
'D': ws.cell(r, 4).value, 'E': ws.cell(r, 5).value, 'F': ws.cell(r, 6).value,
'G': ws.cell(r, 7).value, 'K': u.strip(), 'L': ws.cell(r, 12).value,
'M': ws.cell(r, 13).value, 'N': ws.cell(r, 14).value, 'O': ws.cell(r, 15).value,
}
print('크롤 대상:', len(targets), '', lo, '~', hi)
out = {}
done = 0
with ThreadPoolExecutor(max_workers=10) as ex:
futs = [ex.submit(fetch, r, u) for r, u in targets]
for f in as_completed(futs):
row, url, st, html = f.result()
ev = analyze(url, html)
ev['status'] = st
ev['row'] = row
out[row] = {**rowinfo[row], **ev}
done += 1
if done % 25 == 0:
print(' ...', done, '/', len(targets))
json.dump(out, open(r'D:\01.프로젝트\DB수집\_스크립트\_evidence.json', 'w', encoding='utf-8'),
ensure_ascii=False, indent=1)
print('저장: _evidence.json (', len(out), '행 )')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,91 @@
# -*- coding: utf-8 -*-
"""공주시 전용: 본문에 인페이지 앵커 탭(ul.tab-ul 안 a[href^="#"], 예: #nav1~4)이 있는
페이지는 1 유지하되 수량(M, 13) 개수로 기입.
판별: class 'tab-ul' 포함한 ul 안에서 href '#' 시작하는 <a> 개수(>=2).
페이지에 그런 탭그룹이 여럿이면 가장 그룹의 .
사용: python -X utf8 _gongju_navcount.py [--write]
"""
import sys, warnings
from concurrent.futures import ThreadPoolExecutor, as_completed
import openpyxl, requests, shutil
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
DOMAIN = 'gongju.go.kr'
H = {'User-Agent': 'Mozilla/5.0 Chrome/120 Safari/537.36'}
def nav_count(html):
soup = BeautifulSoup(html, 'html.parser')
best = 0
for ul in soup.find_all('ul'):
cls = ' '.join(ul.get('class') or []).lower()
if 'tab-ul' not in cls:
continue
anchors = [a for a in ul.find_all('a')
if (a.get('href') or '').strip().startswith('#') and a.get_text(strip=True)]
if len(anchors) >= 2:
best = max(best, len(anchors))
return best
def fetch(r, u):
try:
x = requests.get(u, headers=H, timeout=15, verify=False)
return r, u, x.content
except Exception:
return r, u, None
def main():
write = '--write' in sys.argv
wb = openpyxl.load_workbook(XLSX)
ws = wb.active
targets = []
for r in range(3, ws.max_row + 1):
u = ws.cell(r, 11).value
if isinstance(u, str) and DOMAIN in u:
targets.append((r, u))
print(f'스캔 대상(동일도메인): {len(targets)}')
res = {}
with ThreadPoolExecutor(max_workers=8) as ex:
for f in as_completed([ex.submit(fetch, r, u) for r, u in targets]):
r, u, html = f.result()
if html:
c = nav_count(html)
if c >= 2:
res[r] = (u, c)
print(f'\n인페이지 탭(#) 보유 페이지: {len(res)}')
print(f"{'':>4} {'현L':<5}{'현M':>4}{'새M':>4} F/카테고리 | URL")
print('-' * 90)
chg = 0
for r in sorted(res):
u, c = res[r]
L = ws.cell(r, 12).value or ''
M = ws.cell(r, 13).value
cat = ws.cell(r, 6).value or ws.cell(r, 5).value or ws.cell(r, 4).value or ''
mark = '' if str(M) == str(c) else ''
if str(M) != str(c):
chg += 1
print(f"{r:>4} {str(L):<5}{str(M):>4}{c:>4}{mark} {cat} | {u}")
if write:
ws.cell(r, 13).value = c
print('-' * 90)
print(f'변경 대상: {chg}')
if write and chg:
bak = XLSX.replace('.xlsx', '_backup_navcount전.xlsx')
shutil.copy(XLSX, bak)
wb.save(XLSX)
print(f'저장 완료. 백업: {bak}')
elif not write:
print('(DRY — 실제 기입하려면 --write)')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,109 @@
# -*- coding: utf-8 -*-
"""공주시: 인페이지(#nav) 탭을 가진 페이지의 각 탭 패널 안에 '게시판'이 있는지 점검.
대상: _gongju_navcount 에서 잡힌 14 페이지(동일도메인 자동 재탐지).
<a href="#navN"> 패널 element(id=navN) 내부에서 게시판 신호 탐지:
· 목록 table(td 링크 다수) · 페이징 · '총 N건' · 상세링크(view.do/mode=V/nttId/BBSMSTR) · iframe
판정: 신호 1+ 탭은 '게시판 포함' 가능.
사용: python -X utf8 _gongju_navtab_board.py
"""
import re, warnings
from concurrent.futures import ThreadPoolExecutor, as_completed
import openpyxl, requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
DOMAIN = 'gongju.go.kr'
H = {'User-Agent': 'Mozilla/5.0 Chrome/120 Safari/537.36'}
DETAIL = re.compile(r'(view\.do|mode=V|nttId|BBSMSTR|selectBoard|selectBbs)', re.I)
TOTAL = re.compile(r'\s*[\d,]+\s*(건|개)')
def get_navtabs(soup):
"""type3(#nav) 탭그룹의 [(label, panel_id)] 반환 (가장 큰 그룹)."""
best = []
for ul in soup.find_all('ul'):
cls = ' '.join(ul.get('class') or []).lower()
if 'tab-ul' not in cls:
continue
tabs = []
for a in ul.find_all('a'):
h = (a.get('href') or '').strip()
t = a.get_text(strip=True)
if h.startswith('#') and len(h) > 1 and t:
tabs.append((t, h[1:]))
if len(tabs) >= 2 and len(tabs) > len(best):
best = tabs
return best
def board_signals(panel):
if panel is None:
return []
sig = []
# 목록 테이블: 행 3+ 이고 링크 포함
for tb in panel.find_all('table'):
rows = tb.find_all('tr')
links = tb.find_all('a', href=True)
if len(rows) >= 3 and len(links) >= 3:
sig.append('목록table')
break
if panel.select('.paging,.pagination,.board_paging,.bbs_paging'):
sig.append('페이징')
if TOTAL.search(panel.get_text(' ', strip=True)):
sig.append('총건수')
if any(DETAIL.search(a['href']) for a in panel.find_all('a', href=True)):
sig.append('상세링크')
if panel.find('iframe'):
sig.append('iframe')
return sig
def fetch(r, u):
try:
return r, u, requests.get(u, headers=H, timeout=15, verify=False).content
except Exception:
return r, u, None
def main():
ws = openpyxl.load_workbook(XLSX).active
targets = [(r, ws.cell(r, 11).value) for r in range(3, ws.max_row + 1)
if isinstance(ws.cell(r, 11).value, str) and DOMAIN in ws.cell(r, 11).value]
htmls = {}
with ThreadPoolExecutor(max_workers=8) as ex:
for f in as_completed([ex.submit(fetch, r, u) for r, u in targets]):
r, u, h = f.result()
if h:
htmls[r] = (u, h)
found_pages = 0
board_hits = 0
for r in sorted(htmls):
u, h = htmls[r]
soup = BeautifulSoup(h, 'html.parser')
tabs = get_navtabs(soup)
if len(tabs) < 2:
continue
found_pages += 1
cat = ws.cell(r, 6).value or ws.cell(r, 5).value or ws.cell(r, 4).value or ''
per = []
for label, pid in tabs:
sig = board_signals(soup.find(id=pid))
per.append((label, sig))
any_board = any(s for _, s in per)
flag = '🟢게시판 발견' if any_board else '— 게시판 없음(정적 콘텐츠)'
print(f'\n[행{r}] {cat}{len(tabs)}{flag}')
print(f' {u}')
for label, sig in per:
mk = ('게시판? ' + ','.join(sig)) if sig else '정적'
print(f' · {label[:30]:<30} {mk}')
if sig:
board_hits += 1
print(f'\n===== 요약: {found_pages}개 페이지 / 게시판 신호 탭 {board_hits}개 =====')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,171 @@
# -*- coding: utf-8 -*-
"""공주시 전용: type1 탭 메뉴를 재귀 크롤링해 각 시트 행의 '진짜 수량'(잎 페이지 수)을
끝까지 세서 재계산. weight = 인페이지 #nav 개수(>=2) 또는 1.
하위 subtree에 게시판이 하나라도 있으면 행은 '합치지 않음'으로 플래그(M 미변경).
판별:
· type1 메뉴 = ul.tab-ul( type3/#nav 제외) 안 같은도메인 .do 링크 집합 중 self 포함하는 것
· #nav 개수 = ul.tab-ul 안 href^='#' 탭 수
· 게시판 = 목록table/페이징/총N건/상세링크(view.do·mode=V·nttId·BBSMSTR)
사용: python -X utf8 _gongju_recurse.py (DRY 리포트)
python -X utf8 _gongju_recurse.py --write (M 갱신)
"""
import re, sys, shutil, warnings
from urllib.parse import urljoin, urlsplit
from concurrent.futures import ThreadPoolExecutor, as_completed
import openpyxl, requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
DOMAIN = 'gongju.go.kr'
H = {'User-Agent': 'Mozilla/5.0 Chrome/120 Safari/537.36'}
DETAIL = re.compile(r'(view\.do|mode=V|nttId|BBSMSTR|selectBoard|selectBbs)', re.I)
TOTAL = re.compile(r'\s*[\d,]+\s*건')
S = requests.Session(); S.headers.update(H)
CACHE = {} # url -> (menu_set, menu_list, nav, is_board)
def norm(u):
s = urlsplit(u); return urljoin('http://x/', s.path).split('//', 1)[-1].rstrip('/').lower()
def analyze(url):
try:
html = S.get(url, timeout=15, verify=False).content
except Exception:
return (frozenset(), [], 0, False)
soup = BeautifulSoup(html, 'html.parser')
self_n = norm(url)
nav = 0
type1_groups = []
for ul in soup.find_all('ul'):
cls = ' '.join(ul.get('class') or []).lower()
if 'tab-ul' not in cls:
continue
do_links = []
navc = 0
for a in ul.find_all('a'):
h = (a.get('href') or '').strip()
t = a.get_text(strip=True)
if not t:
continue
if h.startswith('#'):
navc += 1
elif h and not h.startswith('javascript:'):
au = urljoin(url, h)
if DOMAIN in au and au.lower().endswith('.do'):
do_links.append((t, au))
if navc >= 2:
nav = max(nav, navc)
if len(do_links) >= 2:
type1_groups.append(do_links)
# self 포함하는 type1 그룹 우선, 없으면 최대
menu = []
for g in type1_groups:
if any(norm(u) == self_n for _, u in g):
menu = g; break
if not menu and type1_groups:
menu = max(type1_groups, key=len)
menu_set = frozenset(norm(u) for _, u in menu)
# 게시판 판정
body = soup.select_one('#txt') or soup
txt = body.get_text(' ', strip=True)
is_board = bool(TOTAL.search(txt)) or bool(body.select('.paging,.pagination,.board_paging')) \
or any(DETAIL.search(a['href']) for a in body.find_all('a', href=True))
return (menu_set, menu, nav, is_board)
def get(url):
n = norm(url)
if n not in CACHE:
CACHE[n] = analyze(url)
return CACHE[n]
def crawl(seed_urls):
seen = set()
queue = list(seed_urls)
while queue:
batch = [u for u in queue if norm(u) not in seen]
for u in batch:
seen.add(norm(u))
queue = []
with ThreadPoolExecutor(max_workers=8) as ex:
futs = {ex.submit(get, u): u for u in batch}
for f in as_completed(futs):
_ms, menu, _nav, _b = f.result()
for _, cu in menu:
if norm(cu) not in seen:
queue.append(cu)
def weight(url):
_ms, _menu, nav, _b = get(url)
return nav if nav >= 2 else 1
def expand(url, parent_set, path):
n = norm(url)
if n in path:
return 1, False
ms, menu, nav, is_board = get(url)
if not menu or ms == parent_set:
return (nav if nav >= 2 else 1), is_board
tot = 0; board = is_board
for _, cu in menu:
w, b = expand(cu, ms, path | {n})
tot += w; board = board or b
return tot, board
def main():
write = '--write' in sys.argv
wb = openpyxl.load_workbook(XLSX); ws = wb.active
rows = []
seeds = []
for r in range(3, ws.max_row + 1):
u = ws.cell(r, 11).value
if isinstance(u, str) and DOMAIN in u and u.lower().endswith('.do'):
rows.append((r, u, ws.cell(r, 12).value, ws.cell(r, 13).value,
ws.cell(r, 6).value or ws.cell(r, 5).value or ws.cell(r, 4).value or ''))
seeds.append(u)
print(f'시트 .do 페이지행: {len(rows)} — 재귀 크롤 시작...')
crawl(seeds)
print(f'크롤한 고유 페이지: {len(CACHE)}\n')
changes = []; boards = []
for r, u, L, M, cat in rows:
if L != '페이지':
continue
newM, hasboard = expand(u, frozenset({norm(u)}), set())
if hasboard:
boards.append((r, cat, u, M, newM))
continue
if str(newM) != str(M):
changes.append((r, cat, M, newM, u))
changes.sort(key=lambda x: -(x[3] - (x[2] or 0)))
print(f'=== 수량 변경(증가/감소) 대상: {len(changes)}행 ===')
for r, cat, oldM, newM, u in changes:
print(f'{r} {cat[:24]:<24} M {oldM}{newM} {u}')
if boards:
print(f'\n=== ⚠ 하위에 게시판 있어 합치지 않음(M 보류): {len(boards)}행 ===')
for r, cat, u, M, nm in boards:
print(f'{r} {cat[:24]:<24} (현 M={M}, 재귀={nm}) {u}')
if write and changes:
bak = XLSX.replace('.xlsx', '_backup_recurse전.xlsx')
shutil.copy(XLSX, bak)
for r, cat, oldM, newM, u in changes:
ws.cell(r, 13).value = newM
wb.save(XLSX)
print(f'\n저장 완료({len(changes)}행 M갱신). 백업: {bak}')
elif not write:
print('\n(DRY — 적용하려면 --write)')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,164 @@
# -*- coding: utf-8 -*-
"""335~443행 일반 다단계 탭 분해. _split_plan.json 의 29개 잡을 적용.
- : 소스행의 가장 깊은 카테고리열(base) 찾아, 자식을 base+1, 손자를 base+2 배치
- 잡은 행번호 내림차순(아래)으로 처리해 미처리 인덱스 불변
- 검증된 시프트 로직(+_style+hyperlink, 병합 +delta 시프트, base열 신규 병합)
- 순번(B) 마지막에 305~ 일괄 재번호
--write 적용.
"""
import sys, json, shutil
from copy import copy
import openpyxl
from openpyxl.utils.cell import range_boundaries
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
PLAN = r'D:\01.프로젝트\DB수집\_스크립트\_split_plan.json'
B_START_ROW = 305
B_START_VAL = 311
CATCOLS = [4, 5, 6, 7, 8, 9, 10] # D..J
def resolve_path(ws, row):
"""소스행의 카테고리 경로(병합 반영). {col: value}"""
path = {}
for c in CATCOLS:
v = ws.cell(row, c).value
if v is None:
# 병합 top-left 찾기
for mr in ws.merged_cells.ranges:
if mr.min_col <= c <= mr.max_col and mr.min_row <= row <= mr.max_row:
v = ws.cell(mr.min_row, mr.min_col).value
break
if v is not None:
path[c] = v
return path
def flatten_leaves(job):
"""잡 children 트리 -> 잎 리스트 [(lvl1name, lvl2name|None, url, L, M)]"""
out = []
for ch in job['children']:
if ch.get('children'):
for gc in ch['children']:
out.append((ch['name'], gc['name'], gc['url'], gc.get('L', '페이지'), gc.get('M', 1)))
else:
out.append((ch['name'], None, ch['url'], ch.get('L', '페이지'), ch.get('M', 1)))
return out
def apply_job(ws, source, leaves, tmpl, last_data):
N = len(leaves)
delta = N - 1
base = max(resolve_path(ws, source).keys()) # 가장 깊은 카테고리열
base_val = ws.cell(source, base).value
if base_val is None:
p = resolve_path(ws, source); base_val = p[base]
# 1) 병합 해제(전체) — 시프트 위해
old_merges = [str(mr) for mr in list(ws.merged_cells.ranges)]
for mr in old_merges:
ws.unmerge_cells(mr)
# 2) 시프트 source+1..last_data -> +delta (아래에서 위로)
if delta > 0:
for sr in range(last_data, source, -1):
dr = sr + delta
for c in range(1, 28):
s = ws.cell(sr, c); d = ws.cell(dr, c)
d.value = s.value
if s.has_style:
d._style = copy(s._style)
if s.hyperlink is not None:
d.hyperlink = copy(s.hyperlink); d.hyperlink.ref = d.coordinate
s.hyperlink = None
else:
d.hyperlink = None
# 3) source..source+N-1 클리어 + 템플릿 스타일
for r in range(source, source + N):
for c in range(1, 28):
cell = ws.cell(r, c)
cell.value = None; cell.hyperlink = None
if c in tmpl:
cell._style = copy(tmpl[c])
# 4) 잎 기입
l1col, l2col = base + 1, base + 2
prev_l1 = None
for i, (l1, l2, url, L, M) in enumerate(leaves):
r = source + i
if l1 != prev_l1:
ws.cell(r, l1col).value = l1
prev_l1 = l1
if l2:
ws.cell(r, l2col).value = l2
kc = ws.cell(r, 11); kc.value = url; kc.hyperlink = url
ws.cell(r, 12).value = L
ws.cell(r, 13).value = M
ws.cell(r, 14).value = '어문'
ws.cell(r, 15).value = '미부착'
# base 값 최상단
ws.cell(source, base).value = base_val
# 5) 병합 재생성: 기존(>=source+1행 +delta) + 신규
def shift(rr):
return rr + delta if rr >= source + 1 else rr
for rng in old_merges:
c1, r1, c2, r2 = range_boundaries(rng)
ws.merge_cells(start_row=shift(r1), start_column=c1, end_row=shift(r2), end_column=c2)
# base열 신규 병합(잡 전체)
if N > 1:
ws.merge_cells(start_row=source, start_column=base, end_row=source + N - 1, end_column=base)
# l1col 손자그룹 병합(같은 l1 연속 run)
i = 0
while i < N:
j = i
while j + 1 < N and leaves[j + 1][0] == leaves[i][0]:
j += 1
if j > i: # 2개 이상 -> 병합
ws.merge_cells(start_row=source + i, start_column=l1col, end_row=source + j, end_column=l1col)
i = j + 1
return last_data + delta
def main():
write = '--write' in sys.argv
jobs = json.load(open(PLAN, encoding='utf-8'))
jobs.sort(key=lambda j: j['row'], reverse=True) # 내림차순(아래→위)
wb = openpyxl.load_workbook(XLSX)
ws = wb.active
# 템플릿 스타일(클린 데이터행 305)
tmpl = {c: copy(ws.cell(305, c)._style) for c in range(1, 28) if ws.cell(305, c).has_style}
# 현재 마지막 데이터행
last_data = 3
for r in range(3, ws.max_row + 1):
if ws.cell(r, 11).value:
last_data = r
print('%d개, 시작 last_data=%d' % (len(jobs), last_data))
if not write:
for j in jobs:
print(' r%d %s -> 잎 %d' % (j['row'], j['label'], len(flatten_leaves(j))))
print('(DRY)')
return
bak = XLSX.replace('.xlsx', '_backup_분야별분해전.xlsx')
shutil.copy(XLSX, bak)
for j in jobs:
leaves = flatten_leaves(j)
last_data = apply_job(ws, j['row'], leaves, tmpl, last_data)
print(' 적용 r%d %s (+%d) last_data=%d' % (j['row'], j['label'][:24], len(leaves) - 1, last_data))
# 순번 B 305~last 일괄 재번호
for r in range(B_START_ROW, last_data + 1):
ws.cell(r, 2).value = B_START_VAL + (r - B_START_ROW)
wb.save(XLSX)
print('저장 완료. 백업:', bak, '| 최종 데이터행', last_data)
if __name__ == '__main__':
main()

View File

@ -0,0 +1,154 @@
# -*- coding: utf-8 -*-
"""공주시 305행(장애인 단일행, 순번311, sub06_01_02_01.do)을
크롤링한 트리(1단계10잎30) 분해.
- 305행을 30(305~334)으로 확장, 306~414행은 +29 시프트(335~443)
- 순번(B) 305~443 연속 재번호(311..), 305 이전(검수본·삭제갭) 불변
- 병합셀 재조정(>=306 +29) + 장애인 F/G 병합 신규
- K열 하이퍼링크 유지/재설정
사용: --write 적용(미지정 계획만 출력)
"""
import sys, shutil
from copy import copy
import openpyxl
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
SPLIT_ROW = 305 # 장애인 행
N = 30 # 새 행 수
SHIFT = N - 1 # 29
DATA_LAST = 414 # 현재 마지막 데이터행
B_START = 311 # 305행 순번
BASE = 'https://www.gongju.go.kr'
# (G, H, url, L, M) H=None 이면 단일(leaf)
ROWS = [
('장애인인권헌장', None, '/kr/sub06_01_02_01.do', '페이지', 1),
('장애인등록안내', '장애인이란', '/kr/sub06_01_02_03_01.do', '페이지', 1),
('장애인등록안내', '장애인등록/심사제도', '/kr/sub06_01_02_03_02.do', '페이지', 1),
('장애인등록안내', '장애인등록현황', '/kr/sub06_01_02_03_03.do', '페이지', 1),
('장애인등록안내', '장애인관련법률', '/kr/sub06_01_02_03_04.do', '페이지', 1),
('장애인복지카드안내', None, '/kr/sub06_01_02_04_01.do', '페이지', 1),
('장애인생활안정지원', '장애수당지급', '/kr/sub06_01_02_05_01.do', '페이지', 1),
('장애인생활안정지원', '장애인의료비지원', '/kr/sub06_01_02_05_02.do', '페이지', 1),
('장애인생활안정지원', '자립자금대여', '/kr/sub06_01_02_05_03.do', '페이지', 1),
('장애인생활안정지원', '재활보조기구 무료교부', '/kr/sub06_01_02_05_04.do', '페이지', 1),
('장애인생활안정지원', '장애아동수당', '/kr/sub06_01_02_05_05.do', '페이지', 1),
('장애인생활안정지원', '장애인일자리', '/kr/sub06_01_02_05_06.do', '페이지', 1),
('자동차관련시책', '장애인자동차표지 발급', '/kr/sub06_01_02_06_01.do', '페이지', 1),
('자동차관련시책', '고속도로통행료 할인', '/kr/sub06_01_02_06_02.do', '페이지', 1),
('자동차관련시책', '자동차특별소비세 면제', '/kr/sub06_01_02_06_03.do', '페이지', 1),
('세금감면시책', '소득세공제', '/kr/sub06_01_02_07_01.do', '페이지', 1),
('세금감면시책', '상속세공제', '/kr/sub06_01_02_07_02.do', '페이지', 1),
('각종요금할인', '전화요금할인', '/kr/sub06_01_02_08_01.do', '페이지', 1),
('각종요금할인', 'TV수신료면제', '/kr/sub06_01_02_08_02.do', '페이지', 1),
('각종요금할인', '이동통신 요금할인', '/kr/sub06_01_02_08_03.do', '페이지', 1),
('각종요금할인', '교통요금 할인', '/kr/sub06_01_02_08_04.do', '페이지', 1),
('각종요금할인', '공공시설 이용요금 감면', '/kr/sub06_01_02_08_05.do', '페이지', 1),
('각종요금할인', '장애인 전기요금 감면', '/kr/sub06_01_02_08_06.do', '페이지', 1),
('각종요금할인', '초고속인터넷 요금할인', '/kr/sub06_01_02_08_07.do', '페이지', 1),
('장애인전화상담', '장애인전화상담소', '/kr/sub06_01_02_09_01.do', '페이지', 1),
('장애인전화상담', '이용안내', '/kr/sub06_01_02_09_02.do', '페이지', 1),
('장애인 전동휠체어 급속충전기 설치장소', None, '/kr/sub06_01_02_10.do', '페이지', 1),
('한눈에 보는 장애인복지 서비스', '시설소개', '/kr/sub06_01_02_11_01.do', '페이지', 1),
('한눈에 보는 장애인복지 서비스', '사업홍보', 'https://www.gongju.go.kr/bbs/BBSMSTR_000000001671/list.do?mno=sub06_01_02_11_02', '게시판', 96),
('한눈에 보는 장애인복지 서비스', '채용정보', 'https://www.gongju.go.kr/bbs/BBSMSTR_000000001672/list.do?mno=sub06_01_02_11_03', '게시판', 11),
]
assert len(ROWS) == N
def full(u):
return u if u.startswith('http') else BASE + u
def main():
write = '--write' in sys.argv
wb = openpyxl.load_workbook(XLSX)
ws = wb.active
if write:
# 0) 모든 병합 먼저 해제(시프트 시 MergedCell 쓰기불가 방지)
old_merges = [str(mr) for mr in list(ws.merged_cells.ranges)]
for mr in old_merges:
ws.unmerge_cells(mr)
# 1) 시프트: source 414..306 -> dest +29 (아래에서 위로)
for sr in range(DATA_LAST, SPLIT_ROW, -1): # 414..306
dr = sr + SHIFT
for c in range(1, 28):
src = ws.cell(sr, c)
dst = ws.cell(dr, c)
dst.value = src.value
if src.has_style:
dst._style = copy(src._style)
# hyperlink 이동
if src.hyperlink is not None:
dst.hyperlink = copy(src.hyperlink)
dst.hyperlink.ref = dst.coordinate
src.hyperlink = None
else:
dst.hyperlink = None
# 2) 305~334 영역 클리어 + 스타일 템플릿(원래 305행) 적용
# 원래 305행 스타일은 시프트 안했으므로 그대로 305에 남아있음(템플릿)
tmpl = {c: copy(ws.cell(SPLIT_ROW, c)._style) for c in range(1, 28) if ws.cell(SPLIT_ROW, c).has_style}
for r in range(SPLIT_ROW, SPLIT_ROW + N):
for c in range(1, 28):
cell = ws.cell(r, c)
cell.value = None
cell.hyperlink = None
if c in tmpl:
cell._style = copy(tmpl[c])
# 3) 30행 데이터 기입 (G는 그룹 최상단에만)
prev_g = None
for i, (g, h, u, L, M) in enumerate(ROWS):
r = SPLIT_ROW + i
if g != prev_g:
ws.cell(r, 7).value = g # G 소 (그룹 최상단)
prev_g = g
if h:
ws.cell(r, 8).value = h # H 세부
fu = full(u)
kc = ws.cell(r, 11)
kc.value = fu
kc.hyperlink = fu
ws.cell(r, 12).value = L # L
ws.cell(r, 13).value = M # M
ws.cell(r, 14).value = '어문' # N
ws.cell(r, 15).value = '미부착' # O
# F=장애인 (병합 최상단)
ws.cell(SPLIT_ROW, 6).value = '장애인'
# 4) 순번(B) 305~443 연속 재번호
for r in range(SPLIT_ROW, DATA_LAST + SHIFT + 1):
ws.cell(r, 2).value = B_START + (r - SPLIT_ROW)
# 5) 병합셀 재조정 (0단계서 해제한 old_merges를 시프트해 재생성)
from openpyxl.utils.cell import range_boundaries
def shift(rr):
return rr + SHIFT if rr >= SPLIT_ROW + 1 else rr
for rng in old_merges:
c1, r1, c2, r2 = range_boundaries(rng)
nr1, nr2 = shift(r1), shift(r2)
ws.merge_cells(start_row=nr1, start_column=c1, end_row=nr2, end_column=c2)
# 신규 병합: F305:F334
ws.merge_cells(start_row=305, start_column=6, end_row=334, end_column=6)
# G 병합(다중 H 그룹)
gmerges = [(306, 309), (311, 316), (317, 319), (320, 321), (322, 328), (329, 330), (332, 334)]
for a, b in gmerges:
ws.merge_cells(start_row=a, start_column=7, end_row=b, end_column=7)
bak = XLSX.replace('.xlsx', '_backup_장애인분해전.xlsx')
shutil.copy(XLSX, bak)
wb.save(XLSX)
print('저장 완료. 백업:', bak)
else:
print('=== 계획(DRY) ===')
print('305행(장애인) → 30행 확장, 306~414 → +29 시프트')
print('순번 305~443 = 311..%d' % (B_START + (DATA_LAST + SHIFT - SPLIT_ROW)))
print('F305:F334 병합, G병합 7개')
for i, (g, h, u, L, M) in enumerate(ROWS):
print(' r%d G=%s H=%s L=%s M=%s' % (SPLIT_ROW + i, g, h or '', L, M))
if __name__ == '__main__':
main()

View File

@ -0,0 +1,707 @@
"""전북특별자치도 14개 시·군 + 제주특별자치도 2개 시 Phase 1 일괄 처리.
매뉴얼: D:\\01.프로젝트\\DB수집\\사이트맵_수집_매뉴얼.md
출력: 폴더의 {기관명}.xlsx (D~K열)
파서 매핑은 _probe_jeonbuk*.py 탐색 결과 기반.
"""
import json
import os
import re
import shutil
import ssl
import sys
import warnings
from copy import copy
from urllib.parse import urljoin
import openpyxl
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.ssl_ import create_urllib3_context
from bs4 import BeautifulSoup
from openpyxl.styles import Alignment, Font
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
ROOT = r'D:\01.프로젝트\DB수집'
TEMPLATE = ROOT + r'\자료_취합_예시.xlsx'
class WeakSSLAdapter(HTTPAdapter):
def init_poolmanager(self, *args, **kwargs):
ctx = create_urllib3_context()
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
ctx.options |= 0x4
ctx.check_hostname = False
ctx.verify_mode = ssl.CERT_NONE
kwargs['ssl_context'] = ctx
return super().init_poolmanager(*args, **kwargs)
def make_session(weak_ssl=False):
s = requests.Session()
s.headers.update(H)
if weak_ssl:
s.mount('https://', WeakSSLAdapter())
return s
def fetch_html(url, session=None, timeout=25):
s = session or requests.Session()
if not session:
s.headers.update(H)
r = s.get(url, timeout=timeout, verify=False, allow_redirects=True)
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
r.encoding = meta.group(1).decode('ascii', errors='ignore') if meta else r.apparent_encoding
return r.text
def clean_text(s):
return re.sub(r'\s+', ' ', s or '').strip().replace('\xa0', '').lstrip('-').strip()
def extract_href(a):
if a is None:
return ''
href = (a.get('href') or '').strip()
if not href or href.startswith('#') or href.lower().startswith('javascript:'):
return ''
return href
COLS = 'DEFGHIJ'
def rows_from_paths(tmp):
"""[{'path':[(text,href)...], 'href':..}] → D~J dict 행 리스트."""
rows = []
for item in tmp:
p = item['path']
row = {c: '' for c in COLS}
row['href'] = item['href']
for i, (t, _) in enumerate(p):
row[COLS[i] if i < len(COLS) else 'J'] = t
rows.append(row)
return rows
def walk_ul(ul, base_path, out, recursive_li=True):
"""ul > li > a (+ 중첩 ul) 재귀. 망가진 마크업(li 안 li)도 허용."""
for li in ul.find_all('li', recursive=False):
a = li.find('a', recursive=False)
if not a:
# div>a 형태(전주) 허용
d = li.find('div', recursive=False)
a = d.find('a', recursive=False) if d else None
if not a:
continue
text = clean_text(a.get_text())
href = extract_href(a)
path = base_path + [(text, href)]
if not text:
continue
# 자식 ul (정상) 또는 li 직접 중첩(망가진 마크업)
child_uls = li.find_all('ul', recursive=False)
child_lis = [c for c in li.find_all('li', recursive=False)]
if child_uls:
out.append({'path': list(path), 'href': href})
for cul in child_uls:
walk_ul(cul, path, out)
elif recursive_li and child_lis:
out.append({'path': list(path), 'href': href})
for cli in child_lis:
# cli를 단일 li로 감싼 가짜 ul처럼 처리
sub = a.find_parent() # not used
_walk_single_li(cli, path, out)
else:
out.append({'path': list(path), 'href': href})
def _walk_single_li(li, base_path, out):
a = li.find('a', recursive=False)
if not a:
d = li.find('div', recursive=False)
a = d.find('a', recursive=False) if d else None
if not a:
return
text = clean_text(a.get_text())
href = extract_href(a)
path = base_path + [(text, href)]
if not text:
return
child_uls = li.find_all('ul', recursive=False)
child_lis = li.find_all('li', recursive=False)
if child_uls:
out.append({'path': list(path), 'href': href})
for cul in child_uls:
walk_ul(cul, path, out)
elif child_lis:
out.append({'path': list(path), 'href': href})
for cli in child_lis:
_walk_single_li(cli, path, out)
else:
out.append({'path': list(path), 'href': href})
# ================================================================
# 파서들
# ================================================================
def parse_menu_div(soup, base):
"""고창/김제/임실: div.sitemap > div.menuN > h4>a (D) + div > ul > li>a (E) + ul (F) 재귀."""
rows = []
sm = soup.select_one('div.sitemap')
if not sm:
return rows
for block in sm.find_all('div', recursive=False):
h4 = block.find('h4')
D = clean_text(h4.get_text()) if h4 else ''
if not D:
continue
inner = block.find('div', recursive=False)
ul = inner.find('ul', recursive=False) if inner else block.find('ul', recursive=False)
if not ul:
rows.append({'D': D, 'href': '', **{c: '' for c in 'EFGHIJ'}})
continue
tmp = []
walk_ul(ul, [], tmp)
for r in rows_from_paths(tmp):
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
'J': r.get('I', ''), 'href': r.get('href', '')})
return rows
def parse_jeongeup(soup, base):
"""정읍: div.sitemap > div.st_mapNN > p.tit>a (D) + ul > li > b>a (E) + ul > li>a (F)."""
rows = []
sm = soup.select_one('div.sitemap')
if not sm:
return rows
for block in sm.find_all('div', recursive=False):
ptit = block.find('p', class_='tit')
D = clean_text(ptit.get_text()) if ptit else ''
if not D:
continue
ul = block.find('ul', recursive=False)
if not ul:
rows.append({'D': D, 'href': '', **{c: '' for c in 'EFGHIJ'}})
continue
for li in ul.find_all('li', recursive=False):
b = li.find('b', recursive=False)
b_a = b.find('a') if b else li.find('a', recursive=False)
E = clean_text(b_a.get_text()) if b_a else ''
E_href = extract_href(b_a) if b_a else ''
sub = li.find('ul', recursive=False)
if not sub:
rows.append({'D': D, 'E': E, 'href': E_href, **{c: '' for c in 'FGHIJ'}})
continue
rows.append({'D': D, 'E': E, 'href': E_href, **{c: '' for c in 'FGHIJ'}})
tmp = []
walk_ul(sub, [], tmp)
for r in rows_from_paths(tmp):
rows.append({'D': D, 'E': E, 'F': r.get('D', ''), 'G': r.get('E', ''),
'H': r.get('F', ''), 'I': r.get('G', ''), 'J': r.get('H', ''),
'href': r.get('href', '')})
return rows
def parse_group_sitemap(soup, base):
"""남원/익산: div.sitemap_group (여러개) > h4.title (D) + ul.sitemap_2dep > li>a (E) + ul (F) 재귀."""
rows = []
groups = soup.select('div.sitemap_group')
for g in groups:
h4 = g.find('h4', class_='title') or g.find('h4')
D = clean_text(h4.get_text()) if h4 else ''
if not D:
continue
ul = g.find('ul', class_='sitemap_2dep') or g.find('ul', recursive=False)
if not ul:
rows.append({'D': D, 'href': '', **{c: '' for c in 'EFGHIJ'}})
continue
tmp = []
walk_ul(ul, [], tmp)
for r in rows_from_paths(tmp):
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
'J': r.get('I', ''), 'href': r.get('href', '')})
return rows
def parse_namwon(soup, base):
"""남원: div.sitemap > (h4 (D) + ul (E/F 재귀)) 형제 반복."""
rows = []
sm = soup.select_one('div.sitemap')
if not sm:
return rows
curD = ''
for child in sm.find_all(['h4', 'ul'], recursive=False):
if child.name == 'h4':
curD = clean_text(child.get_text())
elif child.name == 'ul' and curD:
tmp = []
walk_ul(child, [], tmp)
for r in rows_from_paths(tmp):
rows.append({'D': curD, 'E': r.get('D', ''), 'F': r.get('E', ''),
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
'J': r.get('I', ''), 'href': r.get('href', '')})
return rows
def parse_muju(soup, base):
"""무주: div#sitemap > div.sitemapN > (h4.sNN>a (D) + ul>li>a (E)) 형제 반복."""
rows = []
cont = soup.select_one('div#sitemap')
if not cont:
return rows
for box in cont.find_all('div', recursive=False):
curD = ''
for child in box.find_all(['h4', 'ul'], recursive=False):
if child.name == 'h4':
a = child.find('a')
curD = clean_text(a.get_text() if a else child.get_text())
elif child.name == 'ul' and curD:
for li in child.find_all('li', recursive=False):
a = li.find('a', recursive=False)
if not a:
continue
E = clean_text(a.get_text())
rows.append({'D': curD, 'E': E, 'href': extract_href(a),
**{c: '' for c in 'FGHIJ'}})
return rows
def parse_buan(soup, base):
"""부안: nav#onmenu > ul > li > div.depth_box > div.depth_boxcon > strong (D) + ul>li>a (E) + ul (F)."""
rows = []
nav = soup.select_one('nav#onmenu')
if not nav:
return rows
top = nav.find('ul')
if not top:
return rows
for li in top.find_all('li', recursive=False):
box = li.find('div', class_='depth_boxcon')
if not box:
continue
strong = box.find('strong')
a0 = li.find('a', recursive=False)
D = clean_text(strong.get_text()) if strong else clean_text(a0.get_text() if a0 else '')
if not D:
continue
ul = box.find('ul')
if not ul:
continue
tmp = []
walk_ul(ul, [], tmp)
for r in rows_from_paths(tmp):
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
'J': r.get('I', ''), 'href': r.get('href', '')})
return rows
def parse_sunchang(soup, base):
"""순창: ul.gnb > li > a (D) + div.box ul.gnb_2dep > li>a (E) + ul.gnb_3dep (F) + ul.gnb_4dep (G)."""
rows = []
gnb = soup.select_one('ul.gnb')
if not gnb:
return rows
for li in gnb.find_all('li', recursive=False):
a0 = li.find('a', recursive=False)
D = clean_text(a0.get_text()) if a0 else ''
if not D:
continue
ul2 = li.find('ul', class_='gnb_2dep')
if not ul2:
rows.append({'D': D, 'href': extract_href(a0), **{c: '' for c in 'EFGHIJ'}})
continue
for li2 in ul2.find_all('li', recursive=False):
a2 = li2.find('a', recursive=False)
if not a2:
continue
E = clean_text(a2.get_text())
E_href = extract_href(a2)
ul3 = li2.find('ul', class_='gnb_3dep')
if not ul3:
rows.append({'D': D, 'E': E, 'href': E_href, **{c: '' for c in 'FGHIJ'}})
continue
rows.append({'D': D, 'E': E, 'href': E_href, **{c: '' for c in 'FGHIJ'}})
for li3 in ul3.find_all('li', recursive=False):
a3 = li3.find('a', recursive=False)
if not a3:
continue
F = clean_text(a3.get_text())
F_href = extract_href(a3)
ul4 = li3.find('ul', class_='gnb_4dep')
if not ul4:
rows.append({'D': D, 'E': E, 'F': F, 'href': F_href, **{c: '' for c in 'GHIJ'}})
continue
rows.append({'D': D, 'E': E, 'F': F, 'href': F_href, **{c: '' for c in 'GHIJ'}})
for li4 in ul4.find_all('li', recursive=False):
a4 = li4.find('a', recursive=False)
if not a4:
continue
rows.append({'D': D, 'E': E, 'F': F, 'G': clean_text(a4.get_text()),
'href': extract_href(a4), **{c: '' for c in 'HIJ'}})
return rows
def parse_jangsu(soup, base):
"""장수: div.sitemap > ul.siteMapList > li.sml_1depth > a.sml_1depthBtn (D) + ul.sml_2depthList > li>a (E) + ul 재귀."""
rows = []
ul = soup.select_one('ul.siteMapList')
if not ul:
return rows
for li in ul.find_all('li', recursive=False):
a0 = li.find('a', recursive=False)
D = clean_text(a0.get_text()) if a0 else ''
if not D:
continue
ul2 = li.find('ul', recursive=False)
if not ul2:
rows.append({'D': D, 'href': extract_href(a0), **{c: '' for c in 'EFGHIJ'}})
continue
tmp = []
walk_ul(ul2, [], tmp)
for r in rows_from_paths(tmp):
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
'J': r.get('I', ''), 'href': r.get('href', '')})
return rows
def parse_jeonju(soup, base):
"""전주: div.sitemap_Warp (여러개) > h4.title_h4 (D) + ul > li > div>a (E) + ul/li (F) 망가진 마크업."""
rows = []
warps = soup.select('div.sitemap_Warp')
for w in warps:
h4 = w.find('h4', class_='title_h4') or w.find('h4')
D = clean_text(h4.get_text()) if h4 else ''
if not D:
continue
ul = w.find('ul', recursive=False)
if not ul:
continue
tmp = []
walk_ul(ul, [], tmp)
for r in rows_from_paths(tmp):
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
'J': r.get('I', ''), 'href': r.get('href', '')})
return rows
def parse_jinan(soup, base):
"""진안: div.sitemap > dl > dt (D) + dd > ul > li>a (E) + ul (F) 재귀."""
rows = []
sm = soup.select_one('div.sitemap')
if not sm:
return rows
for dl in sm.find_all('dl', recursive=False):
dt = dl.find('dt')
D = clean_text(dt.get_text()) if dt else ''
if not D:
continue
dd = dl.find('dd')
ul = dd.find('ul', recursive=False) if dd else None
if not ul:
rows.append({'D': D, 'href': '', **{c: '' for c in 'EFGHIJ'}})
continue
tmp = []
walk_ul(ul, [], tmp)
for r in rows_from_paths(tmp):
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
'J': r.get('I', ''), 'href': r.get('href', '')})
return rows
def parse_box_heading(soup, base):
"""서귀포/제주시: div.sitemap > div.sitemapBox|sitemap_menu > h3|h4 (D) + ul > li>a (E)."""
rows = []
sm = soup.select_one('div.sitemap')
if not sm:
return rows
for box in sm.find_all('div', recursive=False):
h = box.find(['h3', 'h4'])
D = clean_text(h.get_text()) if h else ''
if not D:
continue
ul = box.find('ul')
if not ul:
rows.append({'D': D, 'href': '', **{c: '' for c in 'EFGHIJ'}})
continue
tmp = []
walk_ul(ul, [], tmp)
for r in rows_from_paths(tmp):
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
'J': r.get('I', ''), 'href': r.get('href', '')})
return rows
def parse_wanju_json(soup, base):
"""완주: Playwright로 추출한 _wanju_menu.json 로드."""
rows = []
data = json.load(open(os.path.join(os.path.dirname(os.path.abspath(__file__)), '_wanju_menu.json'), encoding='utf-8'))
for r in data:
rows.append({'D': r.get('D', ''), 'E': r.get('E', ''), 'F': r.get('F', ''),
'href': r.get('href', ''), 'G': '', 'H': '', 'I': '', 'J': ''})
return rows
# ================================================================
# 사이트 설정
# ================================================================
def jb(i, name, folder, base, sitemap, parser, domain, weak=False, fetch_kind='html'):
return {'idx': i, 'name': name, 'base': base, 'sitemap': sitemap,
'sheet': f'{i:02d}_{name}', 'parser': parser, 'domain': domain,
'folder': fr'{ROOT}\작업파일\광역_사이트맵\전북특별자치도\{i}.{name}', 'weak_ssl': weak,
'fetch_kind': fetch_kind}
def jj(i, name, base, sitemap, parser, domain):
return {'idx': i, 'name': name, 'base': base, 'sitemap': sitemap,
'sheet': f'{i:02d}_{name}', 'parser': parser, 'domain': domain,
'folder': fr'{ROOT}\작업파일\광역_사이트맵\제주특별자치도\{i}.{name}', 'weak_ssl': False,
'fetch_kind': 'html'}
SITES = [
jb(1, '고창군', '', 'https://www.gochang.go.kr',
'https://www.gochang.go.kr/index.gochang?menuCd=DOM_000000106006000000', parse_menu_div, 'gochang.go.kr'),
jb(2, '군산시', '', 'https://www.gunsan.go.kr',
'https://www.gunsan.go.kr/main', parse_gunsan if False else None, 'gunsan.go.kr'),
jb(3, '김제시', '', 'https://www.gimje.go.kr',
'https://www.gimje.go.kr/index.gimje?menuCd=DOM_000000107002000000', parse_menu_div, 'gimje.go.kr'),
jb(4, '남원시', '', 'https://www.namwon.go.kr',
'https://www.namwon.go.kr/index.do?menuUid=ff8080818f2717db018f277767500088', parse_namwon, 'namwon.go.kr'),
jb(5, '무주군', '', 'https://www.muju.go.kr',
'https://www.muju.go.kr/index.9is?contentUid=ff8080816db80238016dc8e98fa10ef6', parse_muju, 'muju.go.kr'),
jb(6, '부안군', '', 'https://www.buan.go.kr',
'https://www.buan.go.kr/index.buan?contentsSid=1', parse_buan, 'buan.go.kr'),
jb(7, '순창군', '', 'https://www.sunchang.go.kr',
'https://www.sunchang.go.kr/', parse_sunchang, 'sunchang.go.kr'),
jb(8, '완주군', '', 'https://www.wanju.go.kr',
'JSON', parse_wanju_json, 'wanju.go.kr', fetch_kind='json'),
jb(9, '익산시', '', 'https://www.iksan.go.kr',
'https://www.iksan.go.kr/index.do?menuUid=ff80808199f0d11c019a041b8e35174a', parse_group_sitemap, 'iksan.go.kr'),
jb(10, '임실군', '', 'https://www.imsil.go.kr',
'https://www.imsil.go.kr/index.imsil?menuCd=DOM_000000107003000000', parse_menu_div, 'imsil.go.kr'),
jb(11, '장수군', '', 'https://www.jangsu.go.kr',
'https://www.jangsu.go.kr/index.jangsu?menuCd=DOM_000000107001000000', parse_jangsu, 'jangsu.go.kr'),
jb(12, '전주시', '', 'https://www.jeonju.go.kr',
'https://www.jeonju.go.kr/index.9is?contentUid=ff8080818c7c2e8e018c7fd67b9a03b6', parse_jeonju, 'jeonju.go.kr', fetch_kind='html5'),
jb(13, '정읍시', '', 'https://www.jeongeup.go.kr',
'https://www.jeongeup.go.kr/index.jeongeup?menuCd=DOM_000000106002000000', parse_jeongeup, 'jeongeup.go.kr'),
jb(14, '진안군', '', 'https://www.jinan.go.kr',
'https://www.jinan.go.kr/index.jinan?menuCd=DOM_000000110002000000', parse_jinan, 'jinan.go.kr'),
jj(1, '서귀포시', 'https://www.seogwipo.go.kr',
'https://www.seogwipo.go.kr/help/sitemap.htm', parse_box_heading, 'seogwipo.go.kr'),
jj(2, '제주시', 'https://www.jejusi.go.kr',
'https://www.jejusi.go.kr/guide/sitemap.do', parse_box_heading, 'jejusi.go.kr'),
]
def parse_gunsan(soup, base):
"""군산: div#all_pcmenu > div.allmenubox > a.Bmenu (D) + div.allmw ul.sub_pcmenu > li>a (E) + ul.dep3 (F) 재귀."""
rows = []
pc = soup.select_one('div#all_pcmenu')
if not pc:
return rows
for box in pc.select('div.allmenubox'):
bm = box.find('a', class_='Bmenu') or box.find('a')
D = clean_text(bm.get_text()) if bm else ''
if not D:
continue
ul = box.find('ul', class_='sub_pcmenu')
if not ul:
rows.append({'D': D, 'href': extract_href(bm), **{c: '' for c in 'EFGHIJ'}})
continue
# ★ ul.dep4 = 모바일 아코디언 잔재(F형제 전체를 '- '접두로 복제). 진짜 자식 아님 → 제거.
# 안 지우면 각 F의 G자식으로 재귀돼 카르테시안 폭발(2026-05-31 군산 208행 버그 수정).
for d4 in ul.select('ul.dep4'):
d4.decompose()
tmp = []
walk_ul(ul, [], tmp)
for r in rows_from_paths(tmp):
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
'J': r.get('I', ''), 'href': r.get('href', '')})
return rows
# 군산 파서 바인딩(전방참조 해결)
for _s in SITES:
if _s['name'] == '군산시':
_s['parser'] = parse_gunsan
# ================================================================
# 엑셀 생성 (충북 스크립트와 동일 로직)
# ================================================================
def write_excel(site, raw_rows):
name = site['name']
base = site['base']
domain = site['domain']
output = f"{site['folder']}\\{os.path.basename(os.path.dirname(site['folder']))}_{name}.xlsx"
def abs_url(href):
if not href:
return ''
href = href.strip()
if href.startswith(('javascript:', '#')):
return ''
if href.startswith(('http://', 'https://')):
return href
return urljoin(base + '/', href)
def is_external(url):
return url.startswith(('http://', 'https://')) and domain not in url
final_rows = []
i = 0
removed = 0
while i < len(raw_rows):
row = raw_rows[i]
if (i + 1 < len(raw_rows)
and row.get('G', '') == ''
and raw_rows[i + 1].get('D') == row.get('D')
and raw_rows[i + 1].get('E') == row.get('E')
and raw_rows[i + 1].get('F') == row.get('F')
and raw_rows[i + 1].get('G', '') != ''
and raw_rows[i + 1].get('href') == row.get('href')):
removed += 1
i += 1
continue
final_rows.append(row)
i += 1
print(f' [{name}] 원본 {len(raw_rows)} → 중복 {removed} → 최종 {len(final_rows)}')
if not final_rows:
print(f' [{name}] !! 행 0개 — 파서 점검 필요. 엑셀 생성 스킵.')
return False
shutil.copy(TEMPLATE, output)
wb = openpyxl.load_workbook(output)
ws = wb.active
ws.title = site['sheet']
HEADER = {'B1:R1', 'S1:W1', 'Y1:AA1'}
for rng in [str(m) for m in ws.merged_cells.ranges if str(m) not in HEADER]:
ws.unmerge_cells(rng)
for row in ws.iter_rows(min_row=3, max_row=ws.max_row, min_col=1, max_col=ws.max_column):
for cell in row:
cell.value = None
START = 3
template_r = 3
cur_max = ws.max_row
for idx, item in enumerate(final_rows, start=START):
if idx > cur_max:
for c in range(1, ws.max_column + 1):
srcc = ws.cell(template_r, c)
tgt = ws.cell(idx, c)
if srcc.has_style:
tgt.font = copy(srcc.font)
tgt.fill = copy(srcc.fill)
tgt.border = copy(srcc.border)
tgt.alignment = copy(srcc.alignment)
tgt.number_format = srcc.number_format
tgt.protection = copy(srcc.protection)
url = abs_url(item.get('href', ''))
ws.cell(idx, 2).value = idx - 2
ws.cell(idx, 3).value = name
ws.cell(idx, 4).value = item.get('D', '')
ws.cell(idx, 5).value = item.get('E', '')
ws.cell(idx, 6).value = item.get('F', '')
ws.cell(idx, 7).value = item.get('G', '')
ws.cell(idx, 8).value = item.get('H', '')
ws.cell(idx, 9).value = item.get('I', '')
ws.cell(idx, 10).value = item.get('J', '')
ws.cell(idx, 11).value = url
if is_external(url):
ws.cell(idx, 19).value = '외부링크'
END = START + len(final_rows) - 1
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
def merge_runs(col_letter, col_idx, group_cols=()):
runs = []
cur_val = ws.cell(START, col_idx).value
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
run_start = START
for r in range(START + 1, END + 1):
v = ws.cell(r, col_idx).value
g = tuple(ws.cell(r, gg).value for gg in group_cols)
if v == cur_val and g == cur_grp:
continue
if cur_val not in (None, '') and r - 1 > run_start:
runs.append((run_start, r - 1))
cur_val, cur_grp, run_start = v, g, r
if cur_val not in (None, '') and END > run_start:
runs.append((run_start, END))
for s, e in runs:
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
ws.cell(s, col_idx).alignment = center
return len(runs)
n_f = merge_runs('F', 6, group_cols=(4, 5))
n_e = merge_runs('E', 5, group_cols=(4,))
n_d = merge_runs('D', 4)
for r in range(START, END + 1):
for c in (4, 5, 6, 7):
if ws.cell(r, c).value is not None:
ws.cell(r, c).alignment = center
for r in range(1, END + 1):
ws.row_dimensions[r].height = 15
link_n = 0
for r in range(START, END + 1):
cell = ws.cell(r, 11)
u = cell.value
if u and isinstance(u, str) and u.startswith(('http://', 'https://')):
cell.hyperlink = u
old = cell.font
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
bold=old.bold, italic=old.italic,
color='0000FF', underline='single')
cell.alignment = left
link_n += 1
wb.save(output)
ext_n = sum(1 for r in range(START, END + 1) if ws.cell(r, 19).value == '외부링크')
print(f' [{name}] 병합 D:{n_d} E:{n_e} F:{n_f} | 외부링크 {ext_n} | K링크 {link_n}{output}')
return True
def main():
targets = sys.argv[1:] if len(sys.argv) > 1 else [s['name'] for s in SITES]
for site in SITES:
if site['name'] not in targets:
continue
print(f"\n{'='*60}\n[{site['idx']}.{site['name']}] {site['sitemap']}\n{'='*60}")
try:
if site['fetch_kind'] == 'json':
raw_rows = site['parser'](None, site['base'])
else:
sess = make_session(weak_ssl=site.get('weak_ssl', False))
html = fetch_html(site['sitemap'], session=sess)
engine = 'html5lib' if site.get('fetch_kind') == 'html5' else 'html.parser'
soup = BeautifulSoup(html, engine)
raw_rows = site['parser'](soup, site['base'])
write_excel(site, raw_rows)
except Exception as e:
print(f' [{site["name"]}] !! 실패: {e}')
import traceback
traceback.print_exc()
if __name__ == '__main__':
main()

View File

@ -0,0 +1,360 @@
"""전북특별자치도 14개 시·군 + 제주특별자치도 2개 시 Phase 2~4 일괄 처리.
L(형태)·M(건수)·N(저작물유형)·O(공공누리)·P(부착위치)·Q(링크여부) 자동 채움.
매뉴얼: D:\\01.프로젝트\\DB수집\\사이트맵_수집_매뉴얼.md
KOGL(O열) 권위 판정은 별도 _recheck 단계에서 (feedback_kogl_image_rule).
"""
import re
import ssl
import sys
import time
import warnings
from urllib.parse import urljoin
from concurrent.futures import ThreadPoolExecutor, as_completed
import openpyxl
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.ssl_ import create_urllib3_context
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
try:
sys.stdout.reconfigure(line_buffering=True)
except Exception:
pass
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
ROOT = r'D:\01.프로젝트\DB수집'
class WeakSSLAdapter(HTTPAdapter):
def init_poolmanager(self, *args, **kwargs):
ctx = create_urllib3_context()
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
ctx.options |= 0x4
ctx.check_hostname = False
ctx.verify_mode = ssl.CERT_NONE
kwargs['ssl_context'] = ctx
return super().init_poolmanager(*args, **kwargs)
TOTAL_PAT = re.compile(r'\s*(?:게시물\s*)?(\d[\d,]*)\s*(?:건|개|page|페이지)', re.I)
TOTAL_PAT_LOOSE = re.compile(r'(?:전체|총)\s*[:\-]?\s*(\d[\d,]*)\s*건', re.I)
KOGL_IMG_PAT = re.compile(r'(?:new_)?img_open(?:type|code)(\d{1,2})\.(?:png|jpe?g|gif)', re.I)
KOGL_LINK_PAT = re.compile(r'kogl\.or\.kr/info/licenseType(\d)', re.I)
YOUTUBE_PAT = re.compile(r'(?:youtube\.com|youtu\.be)', re.I)
VIDEO_EXT = re.compile(r'\.(mp4|webm|mov|avi)(?:\?|$)', re.I)
DETAIL_PAT = re.compile(
r'(mode=V|view\.do|view\.9is|/view\b|bbtSn=|dataUid=|dataSid=|seqRepeat=|'
r'nttId=|nttNo=|articleNo=|boardSeq=|bbsSeq=|not_ancmt|menukey=)', re.I)
BODY_SEL = ['#main-contents', '#content', '#contents', '.contents', '#txt',
'main', '#container', '#sub']
def site(i, prov, name, weak=False):
folder = fr'{ROOT}\작업파일\광역_사이트맵\{prov}\{i}.{name}'
return name, {'xlsx': fr'{folder}\{prov}_{name}.xlsx', 'body_sel': BODY_SEL, 'weak_ssl': weak}
SITES = dict([
site(1, '전북특별자치도', '고창군'),
site(2, '전북특별자치도', '군산시'),
site(3, '전북특별자치도', '김제시'),
site(4, '전북특별자치도', '남원시'),
site(5, '전북특별자치도', '무주군'),
site(6, '전북특별자치도', '부안군'),
site(7, '전북특별자치도', '순창군'),
site(8, '전북특별자치도', '완주군'),
site(9, '전북특별자치도', '익산시'),
site(10, '전북특별자치도', '임실군'),
site(11, '전북특별자치도', '장수군'),
site(12, '전북특별자치도', '전주시'),
site(13, '전북특별자치도', '정읍시'),
site(14, '전북특별자치도', '진안군'),
site(1, '제주특별자치도', '서귀포시'),
site(2, '제주특별자치도', '제주시'),
])
def make_session(weak_ssl=False):
s = requests.Session()
s.headers.update(H)
if weak_ssl:
s.mount('https://', WeakSSLAdapter())
return s
def fetch(session, url, timeout=5):
try:
r = session.get(url, timeout=timeout, verify=False, allow_redirects=True)
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
r.encoding = meta.group(1).decode('ascii', errors='ignore') if meta else r.apparent_encoding
if r.status_code == 200:
return BeautifulSoup(r.text, 'html.parser')
except Exception:
pass
return None
def get_body(soup, selectors):
for sel in selectors:
el = soup.select_one(sel)
if el:
return el
return soup
def detect_form(body):
has_paging = bool(body.select('.pagination, .paging, nav.paging, .page_nav, .board_paging'))
text_inputs = [i for i in body.find_all('input')
if (i.get('type') or 'text').lower() in ('text', 'search')]
has_search = len(text_inputs) >= 1
txt = body.get_text(' ', strip=True)
m = TOTAL_PAT.search(txt) or TOTAL_PAT_LOOSE.search(txt)
total = None
if m:
digits = m.group(1).replace(',', '')
if digits.isdigit():
total = int(digits)
is_board = has_paging or has_search or (total is not None)
if is_board:
return '게시판', total if total is not None else 0
return '페이지', 1
def extract_detail_urls(body, base_url, limit=5):
urls = []
seen = set()
for a in body.find_all('a', href=True):
h = a['href']
if not h or h.startswith('#'):
continue
if DETAIL_PAT.search(h):
full = urljoin(base_url, h)
if full not in seen:
seen.add(full)
urls.append(full)
if len(urls) >= limit:
break
return urls
def detect_media(body):
has_text = len(body.get_text(strip=True)) > 30
has_image = False
for img in body.find_all('img'):
src = img.get('src', '')
if KOGL_IMG_PAT.search(src):
continue
if not src:
continue
has_image = True
break
has_video = False
for iframe in body.find_all('iframe'):
if YOUTUBE_PAT.search(iframe.get('src', '')):
has_video = True
break
if not has_video:
for a in body.find_all('a', href=True):
if YOUTUBE_PAT.search(a['href']):
has_video = True
break
if not has_video and body.find_all('video'):
has_video = True
if not has_video and VIDEO_EXT.search(str(body)):
has_video = True
return has_image, has_video, has_text
def n_string(has_text, has_image, has_video):
parts = []
if has_text: parts.append('어문')
if has_image: parts.append('이미지')
if has_video: parts.append('영상')
return ','.join(parts) if parts else '없음'
def img_has_valid_anchor(img):
p = img.parent
while p is not None:
if p.name == 'a':
href = p.get('href', '')
if href and not href.startswith('#') and not href.lower().startswith('javascript:'):
return True
return False
p = p.parent
return False
def detect_kogl(body):
types = set()
q_any_y = False
q_any_n = False
for a in body.find_all('a', href=True):
m = KOGL_LINK_PAT.search(a['href'])
if m:
types.add(int(m.group(1)))
q_any_y = True
for img in body.find_all('img'):
src = img.get('src', '')
m = KOGL_IMG_PAT.search(src)
if m:
types.add(int(m.group(1)))
if img_has_valid_anchor(img):
q_any_y = True
else:
q_any_n = True
for el in body.find_all(style=True):
m = KOGL_IMG_PAT.search(el.get('style', ''))
if m:
types.add(int(m.group(1)))
q_any_n = True
if not types:
return set(), None
return types, ('Y' if q_any_y else 'N')
def process_row(session, url, body_selectors):
out = {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
soup = fetch(session, url)
if soup is None:
out['note'] = '접근 실패'
return out
body = get_body(soup, body_selectors)
form, count = detect_form(body)
out['L'] = form
out['M'] = count if form == '게시판' else 1
has_img, has_vid, has_txt = detect_media(body)
types_main, q_main = detect_kogl(body)
P = '게시판' if types_main else ''
types_all = set(types_main)
q_flags = []
if q_main:
q_flags.append(q_main)
if form == '게시판':
detail_urls = extract_detail_urls(body, url, limit=2)
for du in detail_urls:
d_soup = fetch(session, du, timeout=5)
if not d_soup:
continue
d_body = get_body(d_soup, body_selectors)
di, dv, dt = detect_media(d_body)
has_img = has_img or di
has_vid = has_vid or dv
has_txt = has_txt or dt
dt_types, dt_q = detect_kogl(d_body)
if dt_types and not types_main and not P:
P = '게시물'
types_all |= dt_types
if dt_q:
q_flags.append(dt_q)
out['N'] = n_string(has_txt, has_img, has_vid)
if not types_all:
out['O'] = '미부착'
else:
sorted_types = sorted(types_all)
out['O'] = ','.join(f'{n}유형' for n in sorted_types)
out['P'] = P if P else '게시판'
out['Q'] = 'Y' if 'Y' in q_flags else 'N'
return out
def run_site(name, xlsx, body_selectors, weak_ssl=False, workers=14):
print(f'\n{"="*60}\n[{name}] {xlsx}\n{"="*60}')
wb = openpyxl.load_workbook(xlsx)
ws = wb.active
START = 3
END = START - 1
for r in range(START, ws.max_row + 1):
if ws.cell(r, 2).value is None:
break
END = r
tasks = []
for r in range(START, END + 1):
url = ws.cell(r, 11).value
is_ext = (ws.cell(r, 19).value == '외부링크')
tasks.append((r, url, is_ext))
n_ext = sum(1 for t in tasks if t[2])
print(f' 처리 대상: 총 {len(tasks)}행 (외부링크 {n_ext})')
t0 = time.time()
results = {}
session = make_session(weak_ssl=weak_ssl)
def worker(task):
row, url, is_ext = task
if is_ext:
return row, {'L': '사이트', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
if not url or not isinstance(url, str):
return row, {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': 'URL 없음'}
return row, process_row(session, url, body_selectors)
done = 0
with ThreadPoolExecutor(max_workers=workers) as ex:
futs = [ex.submit(worker, t) for t in tasks]
for fut in as_completed(futs):
row, res = fut.result()
results[row] = res
done += 1
if done % 50 == 0 or done == len(tasks):
print(f' 진행 {done}/{len(tasks)} ({time.time()-t0:.0f}s)')
for r in range(START, END + 1):
res = results.get(r, {})
if not res:
continue
if res.get('L'): ws.cell(r, 12).value = res['L']
if res.get('M') != '': ws.cell(r, 13).value = res['M']
if res.get('N'): ws.cell(r, 14).value = res['N']
if res.get('O'): ws.cell(r, 15).value = res['O']
if res.get('P'): ws.cell(r, 16).value = res['P']
if res.get('Q'): ws.cell(r, 17).value = res['Q']
if res.get('note'):
existing = ws.cell(r, 19).value
if not existing:
ws.cell(r, 19).value = res['note']
wb.save(xlsx)
forms = {}
attach = {'미부착': 0, '부착': 0, '기타': 0}
q_dist = {'Y': 0, 'N': 0, '': 0}
for r, res in results.items():
forms[res.get('L', '')] = forms.get(res.get('L', ''), 0) + 1
o = res.get('O', '')
if o == '미부착': attach['미부착'] += 1
elif o and '유형' in o: attach['부착'] += 1
else: attach['기타'] += 1
q = res.get('Q', '')
q_dist[q] = q_dist.get(q, 0) + 1
print(f' [{name}] L분포 {forms} | 부착 {attach} | Q {q_dist} | 시간 {time.time()-t0:.0f}s')
def main():
targets = sys.argv[1:] if len(sys.argv) > 1 else list(SITES.keys())
total_t0 = time.time()
for name in targets:
if name not in SITES:
print(f' 알 수 없음: {name}')
continue
cfg = SITES[name]
try:
run_site(name, cfg['xlsx'], cfg['body_sel'], weak_ssl=cfg.get('weak_ssl', False))
except Exception as e:
print(f' [{name}] 실패: {e}')
import traceback
traceback.print_exc()
print(f'\n총 소요: {time.time()-total_t0:.0f}s')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,30 @@
# -*- coding: utf-8 -*-
"""공공기관2 몽타주 N 도구 생성(nshot/napply/nbatch 경로 적응)."""
import io, os
D = r"D:\01.프로젝트\DB수집\_스크립트"
def transform(src, dst, repls):
s = io.open(os.path.join(D, src), encoding="utf-8").read()
for a, b in repls:
if a not in s:
print(" [warn] 못찾음:", repr(a[:50]))
s = s.replace(a, b)
io.open(os.path.join(D, dst), "w", encoding="utf-8").write(s)
print("생성:", dst)
# 번호폴더 구조 대응: xlsx 경로를 glob로
XLSX_OLD = "xlsx = os.path.join(OUTDIR, f'{name}.xlsx')"
XLSX_NEW = "xlsx = __import__('glob').glob(os.path.join(OUTDIR, f'*.{name}', f'{name}.xlsx'))[0]"
OUTDIR_OLD = r"OUTDIR = r'D:\01.프로젝트\DB수집\공공기관'"
OUTDIR_NEW = r"OUTDIR = r'D:\01.프로젝트\DB수집\작업파일\공공기관2'"
transform("_공공기관_nshot.py", "_공공기관2_nshot.py", [(OUTDIR_OLD, OUTDIR_NEW), (XLSX_OLD, XLSX_NEW)])
transform("_공공기관_napply.py", "_공공기관2_napply.py", [(OUTDIR_OLD, OUTDIR_NEW), (XLSX_OLD, XLSX_NEW)])
# nbatch: 모듈/probe 경로
transform("_공공기관_nbatch.py", "_공공기관2_nbatch.py", [
("_공공기관_nshot", "_공공기관2_nshot"),
("_공공기관_napply", "_공공기관2_napply"),
("_공공기관_probe.json", "_공공기관2_probe.json"),
])
print("done")

View File

@ -0,0 +1,39 @@
# -*- coding: utf-8 -*-
"""공공기관2 배치용 probe/phase1/phase234 스크립트 생성 (경로만 치환)."""
import io, os
D = r"D:\01.프로젝트\DB수집\_스크립트"
def transform(src_name, dst_name, repls):
src = io.open(os.path.join(D, src_name), encoding="utf-8").read()
for a, b in repls:
if a not in src:
print(" [warn] 패턴 못찾음:", repr(a[:60]))
src = src.replace(a, b)
io.open(os.path.join(D, dst_name), "w", encoding="utf-8").write(src)
print("생성:", dst_name)
# probe: STATUS + 출력 json
transform("_공공기관_probe.py", "_공공기관2_probe.py", [
(r"작업파일\공공기관_작업현황.xlsx", r"작업파일\공공기관2\공공기관2_작업현황.xlsx"),
("_공공기관_probe.json", "_공공기관2_probe.json"),
])
# phase1: OUTDIR(번호폴더구조) + PROBE json + 출력경로 번호폴더
transform("_공공기관_phase1.py", "_공공기관2_phase1.py", [
(r"OUTDIR = r'D:\01.프로젝트\DB수집\공공기관'",
r"OUTDIR = r'D:\01.프로젝트\DB수집\작업파일\공공기관2'"),
("_공공기관_probe.json", "_공공기관2_probe.json"),
("output = os.path.join(OUTDIR, f'{name}.xlsx')",
"output = os.path.join(OUTDIR, f'{num}.{name}', f'{name}.xlsx')"),
])
# phase234: OUTDIR + PROBE json + 입력경로 번호폴더
transform("_공공기관_phase234.py", "_공공기관2_phase234.py", [
(r"OUTDIR = r'D:\01.프로젝트\DB수집\공공기관'",
r"OUTDIR = r'D:\01.프로젝트\DB수집\작업파일\공공기관2'"),
("_공공기관_probe.json", "_공공기관2_probe.json"),
("xlsx = os.path.join(OUTDIR, f'{name}.xlsx')",
"xlsx = os.path.join(OUTDIR, f'{num}.{name}', f'{name}.xlsx')"),
])
print("done")

View File

@ -0,0 +1,179 @@
# -*- coding: utf-8 -*-
"""335~443행 분해 계획 산출(읽기전용).
- 미분해 F섹션(type1 URL이 시트에 1=자기자신만 존재) type1 탭들로 분해
- 이미 분해된 행의 깊은 중첩(type3가 별도 URL, 시트에 없음) type3로 분해
- type3가 #nav 앵커면 분해 안 함(M=앵커수)
재귀로 깊은 단계까지. 결과: 분해 목록 + 증가행수.
"""
import re, json, warnings
from urllib.parse import urlparse
import requests
from bs4 import BeautifulSoup
import openpyxl
warnings.filterwarnings('ignore')
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
BASE = 'https://www.gongju.go.kr'
H = {'User-Agent': 'Mozilla/5.0 Chrome/120 Safari/537.36'}
LO, HI = 335, 443
_cache = {}
def full(h):
return h if h.startswith('http') else BASE + h
def norm(u):
"""시트 멤버십 비교용 정규화: host 제거, path(+bbs id)만"""
p = urlparse(u if u.startswith('http') else BASE + u)
path = p.path
return path
def get(u):
u = full(u)
if u in _cache:
return _cache[u]
try:
x = requests.get(u, headers=H, timeout=25, verify=False)
x.encoding = x.apparent_encoding or 'utf-8'
html = x.text
except Exception:
html = ''
_cache[u] = html
return html
def tab_groups(html):
s = BeautifulSoup(html, 'html.parser')
t1 = []
t3 = []
for ul in s.find_all('ul'):
cls = ' '.join(ul.get('class') or [])
if 'tab-ul' not in cls.lower():
continue
items = [(a.get_text(strip=True), a.get('href', '')) for a in ul.find_all('a')
if a.get_text(strip=True) and a.get('href')]
if not items:
continue
if 'type1' in cls:
t1 = items
elif not t3:
t3 = items
return t1, t3
def board_count(html):
s = BeautifulSoup(html, 'html.parser')
el = s.select_one('.program--count strong')
if el:
m = re.sub(r'[^0-9]', '', el.get_text())
return int(m) if m else None
return None
def is_board(u):
return bool(re.search(r'list\.do|/bbs/|BBSMSTR', u))
def leaf_info(name, u):
"""잎 1개의 (L,M) 결정 + 그 잎의 #nav 앵커수 반영"""
fu = full(u)
html = get(fu)
if is_board(u):
c = board_count(html)
return {'name': name, 'url': fu, 'L': '게시판', 'M': c if c is not None else 0}
# 페이지: 자기 type3가 #nav면 M=앵커수
_, t3 = tab_groups(html)
navs = [h for t, h in t3 if h.startswith('#')]
if len(navs) >= 2:
return {'name': name, 'url': fu, 'L': '페이지', 'M': len(navs)}
return {'name': name, 'url': fu, 'L': '페이지', 'M': 1}
def url_type3(u):
"""페이지의 type3 별도URL 탭들(없으면 [])"""
if is_board(u):
return []
_, t3 = tab_groups(get(full(u)))
return [(t, full(h)) for t, h in t3 if not h.startswith('#')]
def expand_nested(name, u, sheet_urls, depth=0):
"""깊이 2 평탄 분해(공주 탭 최대 2단계). 자식이 url-type3를 가지면 1단계만 더 펼침."""
grand = url_type3(u) if depth == 0 else []
if len(grand) >= 2:
return {'name': name, 'url': full(u),
'children': [leaf_info(t, h) for t, h in grand]}
return leaf_info(name, u)
def main():
wb = openpyxl.load_workbook(XLSX)
ws = wb.active
sheet_urls = set()
for r in range(3, 445):
u = ws.cell(r, 11).value
if isinstance(u, str) and u.startswith('http'):
sheet_urls.add(norm(u))
jobs = []
for r in range(LO, HI + 1):
u = ws.cell(r, 11).value
if not (isinstance(u, str) and u.startswith('http')):
continue
html = get(u)
t1, t3 = tab_groups(html)
t1_urls = [(t, full(h)) for t, h in t1 if not h.startswith('#')]
present = sum(1 for t, h in t1_urls if norm(h) in sheet_urls)
url_t3 = [(t, full(h)) for t, h in t3 if not h.startswith('#')]
label = (ws.cell(r, 6).value or '')
if ws.cell(r, 7).value:
label += ' > ' + str(ws.cell(r, 7).value)
if ws.cell(r, 8).value:
label += ' > ' + str(ws.cell(r, 8).value)
# 미분해 F섹션: type1 URL>=2 이고 시트에 자기 1개만 → type1 자식(각자 type3 1단계 더)
if len(t1_urls) >= 2 and present <= 1:
children = [expand_nested(t, h, sheet_urls, depth=0) for t, h in t1_urls]
jobs.append(('SECTION', r, label, children))
# 이미 분해된 G-잎의 깊은 type3 중첩 → type3 자식(잎)
elif len(url_t3) >= 2:
children = [leaf_info(t, h) for t, h in url_t3]
jobs.append(('NEST', r, label, children))
# 요약
def count_leaves(nodes):
n = 0
for nd in nodes:
if nd.get('children'):
n += count_leaves(nd['children'])
else:
n += 1
return n
total_add = 0
print('=== 분해 계획 (335~443) ===')
for kind, r, label, children in jobs:
leaves = count_leaves(children)
add = leaves - 1 # 기존 1행 대체
total_add += add
print('\n[%s] r%d %s → 잎 %d (기존1, +%d)' % (kind, r, label, leaves, add))
def pr(nodes, ind=1):
for nd in nodes:
tag = ('게시판%s' % nd['M']) if nd.get('L') == '게시판' else ('페이지M%s' % nd.get('M', ''))
if nd.get('children'):
print(' ' * ind + '%s' % nd['name'])
pr(nd['children'], ind + 1)
else:
print(' ' * ind + '- %s [%s]' % (nd['name'], tag))
pr(children)
print('\n총 분해 잡 %d개, 총 증가 행수 +%d (최종 데이터행 ~ %d)' % (len(jobs), total_add, 443 + total_add))
json.dump([{'kind': k, 'row': r, 'label': l, 'children': c} for k, r, l, c in jobs],
open(r'D:\01.프로젝트\DB수집\_스크립트\_split_plan.json', 'w', encoding='utf-8'),
ensure_ascii=False, indent=1)
if __name__ == '__main__':
main()

107
_스크립트/_probe2.py Normal file
View File

@ -0,0 +1,107 @@
"""Detailed structural analysis of representative sitemap pages."""
import warnings
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
SAMPLES = {
# name: (sitemap_url, selector_for_container)
'논산시': ('https://nonsan.go.kr/kor/html/sub07/0701.html', '.sitemap'),
'당진시': ('https://www.dangjin.go.kr/kor/sitemap_11.do', '#sitemap'),
'보령시': ('https://www.brcn.go.kr/kor/sitemap_11.do', '#sitemap'),
'서천군': ('https://www.seocheon.go.kr/kor/sitemap_11.do', '#sitemap'),
'청양군': ('https://www.cheongyang.go.kr/kor/sitemap_11.do', '#sitemap'),
'태안군': ('https://www.taean.go.kr/kor/sitemap_11.do', '#sitemap'),
'아산시': ('https://www.asan.go.kr/main/sitemap.do', '.sitemap'),
'예산군': ('https://www.yesan.go.kr/kor/sitemap.do', 'ul.sitemap'),
'천안시': ('https://www.cheonan.go.kr/kor/sitemap.do', 'ul.sitemap'),
'홍성군': ('https://www.hongseong.go.kr/kor/sitemap.do', 'div.sitemap'),
'부여군': ('https://www.buyeo.go.kr/html/kr/html/sub07/0701.html', '.sitemap'),
'서산시': ('https://www.seosan.go.kr/www/contents.do?key=151', '.sitemap'),
}
def fetch(url):
r = requests.get(url, headers=H, timeout=20, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
return r.url, r.text, r.status_code
def dump_outline(el, indent=0, max_lines=80, lines=None):
if lines is None:
lines = []
if len(lines) >= max_lines:
return lines
name = el.name
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
label = f'{name}'
if eid:
label += f'#{eid}'
if cls:
label += f'.{cls.replace(" ", ".")}'
if name == 'a':
txt = el.get_text(strip=True)[:40]
href = el.get('href', '')[:60]
lines.append(' ' * indent + f'{label} "{txt}"{href}')
else:
lines.append(' ' * indent + label)
for c in el.find_all(recursive=False):
if c.name in ('script', 'style'):
continue
dump_outline(c, indent + 1, max_lines, lines)
if len(lines) >= max_lines:
return lines
return lines
def main():
for name, (url, sel) in SAMPLES.items():
print(f'\n{"="*70}\n{name}: {url}\n{"="*70}')
try:
final, html, code = fetch(url)
except Exception as e:
print(f' ERR: {e}')
continue
if code != 200:
print(f' HTTP {code}')
# Try alternative locations
for alt in [url.replace('sub07', 'sub06'), url.replace('contents.do?key=151', 'sitemap.do')]:
try:
final, html, code = fetch(alt)
if code == 200:
print(f' 대체 OK: {alt}')
break
except Exception:
pass
else:
continue
soup = BeautifulSoup(html, 'html.parser')
el = soup.select_one(sel)
if el is None:
print(f' selector "{sel}" 매칭 실패')
# Find any container with many anchors
for d in soup.find_all(['div', 'section', 'main', 'ul']):
if len(d.find_all('a')) >= 100:
cls = ' '.join(d.get('class', []))
print(f' 대안 발견: {d.name}.{cls} id={d.get("id","")} (a={len(d.find_all("a"))})')
el = d
break
if el is None:
continue
print(f'\n 컨테이너: {el.name} class={el.get("class")} id={el.get("id")}')
print(f' anchors={len(el.find_all("a"))} dl={len(el.find_all("dl"))} ul={len(el.find_all("ul"))} li={len(el.find_all("li"))}')
print(' 구조 (max 80 lines):')
lines = dump_outline(el)
for l in lines:
print(' ', l)
if __name__ == '__main__':
main()

81
_스크립트/_probe3.py Normal file
View File

@ -0,0 +1,81 @@
"""Find sitemap URLs for: 논산시, 아산시, 부여군, 서산시."""
import re
import warnings
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
def fetch(url):
try:
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
return r.status_code, r.url, r.text
except Exception as e:
return 0, str(e), ''
def find_links_to(html, base, keyword='사이트맵'):
soup = BeautifulSoup(html, 'html.parser')
found = []
for a in soup.find_all('a', href=True):
txt = a.get_text(strip=True)
href = a['href']
if keyword in txt or 'sitemap' in href.lower() or 'allMenu' in href:
full = urljoin(base, href)
if full != base:
found.append((txt, full))
return found
CASES = [
('논산시', 'https://nonsan.go.kr/'),
('아산시', 'https://www.asan.go.kr/main/'),
('부여군', 'https://www.buyeo.go.kr/html/kr/'),
('서산시', 'https://www.seosan.go.kr/www/index.do'),
]
for name, base in CASES:
print(f'\n=== {name} {base} ===')
code, url, html = fetch(base)
if code != 200:
print(f' main 실패: {code}')
continue
links = find_links_to(html, url)
# Deduplicate
seen = set()
uniq = []
for t, u in links:
if u not in seen:
seen.add(u)
uniq.append((t, u))
print(f' 사이트맵/allMenu 링크 후보 ({len(uniq)}):')
for t, u in uniq[:10]:
print(f' "{t}"{u}')
# For top candidate, fetch and analyze
for t, u in uniq[:3]:
c2, u2, h2 = fetch(u)
if c2 != 200:
print(f' [{u}] HTTP {c2}')
continue
soup = BeautifulSoup(h2, 'html.parser')
# Find best container
best = (0, None, None)
for sel in ['.sitemap_grep', '.sitemap', '#sitemap', '.allMenu', '#allMenu',
'div[class*=sitemap]', 'div[class*=allMenu]', 'ul.sitemap_list',
'ul.depth1_ul', 'ul.depth1-ul', '#gnb', 'div.menu_all']:
for el in soup.select(sel):
ac = len(el.find_all('a'))
if ac > best[0]:
best = (ac, sel, el)
if best[1]:
cls = ' '.join(best[2].get('class', []))
print(f'{u2}: best={best[1]!r} cls={cls!r} a={best[0]}')
else:
print(f' [{u2}]: no container')

99
_스크립트/_probe4.py Normal file
View File

@ -0,0 +1,99 @@
"""Brute-force common sitemap paths for 아산시, 부여군, 서산시."""
import warnings
from urllib.parse import urlparse
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
def fetch(url):
try:
r = requests.get(url, headers=H, timeout=10, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
return r.status_code, r.url, r.text
except Exception as e:
return 0, str(e), ''
def analyze(html):
soup = BeautifulSoup(html, 'html.parser')
best = (0, '', '')
for sel in ['.sitemap_grep', '.sitemap', '#sitemap', '.allMenu', '#allMenu',
'div[class*=sitemap]', 'div[class*=allMenu]', 'ul.sitemap_list',
'ul.depth1_ul', 'ul.depth1-ul', '#gnb', 'div.menu_all',
'div[class*=menu_all]']:
for el in soup.select(sel):
ac = len(el.find_all('a'))
if ac > best[0]:
best = (ac, sel, ' '.join(el.get('class', [])))
return best
PATHS = [
'/main/sitemap.do',
'/main/sub01_01.do',
'/main/sub.do?key=121',
'/main/sub.do?key=151',
'/sitemap.do',
'/main/contents.do?key=121',
'/main/contents.do?key=151',
'/main/sitemap.html',
'/main/sitemap',
'/main/menu.do',
'/main/allMenu.do',
'/main/totalMenu.do',
'/sub01_01.do',
]
ORIGINS = {
'아산시': 'https://www.asan.go.kr',
'부여군': 'https://www.buyeo.go.kr',
'서산시': 'https://www.seosan.go.kr',
}
# 부여군 prefix
BUYEO_PATHS = [
'/html/kr/sitemap.html',
'/html/kr/html/guide/0701.html',
'/html/kr/html/sub06/0701.html',
'/html/kr/html/sub07/0701.html',
'/html/kr/html/sub01/0101.html',
'/html/kr/sub.do?key=151',
]
# 서산시 prefix
SEOSAN_PATHS = [
'/www/sitemap.do',
'/www/sub.do?key=121',
'/www/sub.do?key=151',
'/www/menu_all.do',
'/www/allMenu.do',
'/www/contents.do?key=121',
'/www/contents.do?key=151',
'/www/contents.do?key=141',
'/www/contents.do?key=131',
]
def try_paths(name, origin, paths):
print(f'\n=== {name} ===')
for p in paths:
url = origin + p
code, real, html = fetch(url)
if code != 200:
continue
info = analyze(html)
if info[0] >= 50:
print(f'{url}: a={info[0]} sel={info[1]!r} cls={info[2]!r}')
else:
print(f' {url}: a={info[0]} (인덱싱 부족)')
try_paths('아산시', ORIGINS['아산시'], PATHS)
try_paths('부여군', ORIGINS['부여군'], BUYEO_PATHS + PATHS)
try_paths('서산시', ORIGINS['서산시'], SEOSAN_PATHS)

37
_스크립트/_probe5.py Normal file
View File

@ -0,0 +1,37 @@
"""Inspect main pages of 아산시, 부여군, 서산시 — grep for 사이트맵 keyword anywhere."""
import re
import warnings
from urllib.parse import urljoin, urlparse
import requests
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
PAGES = {
'아산시': 'https://www.asan.go.kr/main/',
'부여군': 'https://www.buyeo.go.kr/html/kr/',
'서산시': 'https://www.seosan.go.kr/www/index.do',
}
def fetch(url):
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
return r.status_code, r.url, r.text
for name, url in PAGES.items():
print(f'\n=== {name} === {url}')
code, real, html = fetch(url)
print(f' HTTP {code}, len {len(html)}')
# Search for 사이트맵 or sitemap or allMenu
for keyword in ['사이트맵', 'sitemap', 'allMenu', '전체메뉴', 'totalMenu']:
pat = re.compile(r'[\'"][^\'"]*' + keyword + r'[^\'"]*[\'"]', re.I)
matches = pat.findall(html)[:10]
if matches:
print(f' "{keyword}" 매칭:')
for m in matches:
print(f' {m}')

57
_스크립트/_probe6.py Normal file
View File

@ -0,0 +1,57 @@
"""Look for sitemap-like containers directly in main pages."""
import re
import warnings
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
def fetch(url):
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
return r.url, r.text
# 부여군 — find anchor with text 사이트맵
print('\n=== 부여군: extract sitemap link ===')
real, html = fetch('https://www.buyeo.go.kr/html/kr/')
soup = BeautifulSoup(html, 'html.parser')
for a in soup.find_all('a', href=True):
txt = a.get_text(strip=True)
if '사이트맵' in txt:
full = urljoin(real, a['href'])
print(f' 사이트맵 → {full}')
# Also look in JS for sitemap URL
for s in re.findall(r"location\.(?:href|replace)\s*=\s*['\"]([^'\"]+)['\"]", html):
if 'sitemap' in s.lower():
print(f' JS sitemap → {s}')
# 아산시 — gnb-menu inside the page
print('\n=== 아산시: scan for inline gnb menu ===')
real, html = fetch('https://www.asan.go.kr/main/')
soup = BeautifulSoup(html, 'html.parser')
# Find any container with many anchors that looks menu-like
for el in soup.select('nav, [class*=gnb], [class*=menu]'):
a_count = len(el.find_all('a'))
if a_count >= 50:
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
print(f' {el.name}#{eid}.{cls} a={a_count}')
# 서산시 — same approach
print('\n=== 서산시: scan for inline menu ===')
real, html = fetch('https://www.seosan.go.kr/www/index.do')
soup = BeautifulSoup(html, 'html.parser')
for el in soup.select('nav, [class*=gnb], [class*=menu], [class*=allMenu]'):
a_count = len(el.find_all('a'))
if a_count >= 30:
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
print(f' {el.name}#{eid}.{cls} a={a_count}')

76
_스크립트/_probe7.py Normal file
View File

@ -0,0 +1,76 @@
"""Inspect specific selectors found in probe6 for 아산시, 서산시, 부여군."""
import warnings
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
def fetch(url):
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
return r.url, r.text
def outline(el, depth=0, max_lines=120, lines=None):
if lines is None:
lines = []
if len(lines) >= max_lines:
return lines
name = el.name
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
label = name
if eid:
label += f'#{eid}'
if cls:
label += '.' + cls.replace(' ', '.')
if name == 'a':
txt = el.get_text(strip=True)[:50]
href = el.get('href', '')[:80]
lines.append(' ' * depth + f'{label} "{txt}"{href}')
else:
lines.append(' ' * depth + label)
for c in el.find_all(recursive=False):
if c.name in ('script', 'style'):
continue
outline(c, depth + 1, max_lines, lines)
if len(lines) >= max_lines:
return lines
return lines
# 부여군: find .pc_sitemap
print('\n=== 부여군: .pc_sitemap 또는 inline sitemap ===')
real, html = fetch('https://www.buyeo.go.kr/html/kr/')
soup = BeautifulSoup(html, 'html.parser')
for sel in ['.pc_sitemap', '#pc_sitemap', '.sitemap_grep', '.allMenu', '.allmenu',
'div[class*=sitemap]', 'div[class*=gnb]']:
for el in soup.select(sel):
ac = len(el.find_all('a'))
if ac >= 30:
cls = ' '.join(el.get('class', []))
print(f' ★ sel={sel} {el.name}.{cls} id={el.get("id","")} a={ac}')
# 아산시: mobile-nav.krds-gnb-mobile
print('\n=== 아산시: mobile-nav ===')
real, html = fetch('https://www.asan.go.kr/main/')
soup = BeautifulSoup(html, 'html.parser')
el = soup.select_one('nav#mobile-nav')
if el:
for line in outline(el, max_lines=80):
print(' ', line)
# 서산시: top_menu
print('\n=== 서산시: ul#top_menu ===')
real, html = fetch('https://www.seosan.go.kr/www/index.do')
soup = BeautifulSoup(html, 'html.parser')
el = soup.select_one('ul#top_menu') or soup.select_one('div.menu_wrap')
if el:
for line in outline(el, max_lines=80):
print(' ', line)

85
_스크립트/_probe8.py Normal file
View File

@ -0,0 +1,85 @@
"""Find sitemap containers in 부여군 main page and 아산시 desktop GNB."""
import warnings
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
def fetch(url):
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
return r.url, r.text
def outline(el, depth=0, max_lines=120, lines=None):
if lines is None:
lines = []
if len(lines) >= max_lines:
return lines
name = el.name
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
label = name
if eid:
label += f'#{eid}'
if cls:
label += '.' + cls.replace(' ', '.')
if name == 'a':
txt = el.get_text(strip=True)[:50]
href = el.get('href', '')[:80]
lines.append(' ' * depth + f'{label} "{txt}"{href}')
else:
lines.append(' ' * depth + label)
for c in el.find_all(recursive=False):
if c.name in ('script', 'style'):
continue
outline(c, depth + 1, max_lines, lines)
if len(lines) >= max_lines:
return lines
return lines
# 부여군 — scan ALL container types for menu-like content
print('=== 부여군: scan all containers ===')
real, html = fetch('https://www.buyeo.go.kr/html/kr/')
soup = BeautifulSoup(html, 'html.parser')
# Find divs with most anchors
candidates = []
for d in soup.find_all(['div', 'nav', 'ul']):
a_count = len(d.find_all('a'))
if a_count >= 100:
cls = ' '.join(d.get('class', []))
eid = d.get('id', '')
candidates.append((a_count, d.name, eid, cls, d))
candidates.sort(reverse=True, key=lambda x: x[0])
for ac, name, eid, cls, _ in candidates[:10]:
print(f' {name}#{eid}.{cls[:60]} a={ac}')
# Outline top container
if candidates:
print('\n Top container outline:')
for line in outline(candidates[0][4], max_lines=60):
print(' ', line)
# 아산시 — desktop GNB
print('\n\n=== 아산시: desktop GNB ===')
real, html = fetch('https://www.asan.go.kr/main/')
soup = BeautifulSoup(html, 'html.parser')
# Look for desktop GNB sub-lists
for el in soup.select('[id^=mGnb-anchor], .gnb-sub-list, .submenu-wrap, .gnb-wrap'):
ac = len(el.find_all('a'))
if ac >= 30:
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
print(f' {el.name}#{eid}.{cls} a={ac}')
# Try the main desktop nav specifically
el = soup.select_one('nav.krds-gnb:not(#mobile-nav)')
if el:
print('\n Desktop nav outline:')
for line in outline(el, max_lines=80):
print(' ', line)

110
_스크립트/_probe9.py Normal file
View File

@ -0,0 +1,110 @@
"""Detailed structural look at 논산시, 아산시, 부여군 main sitemaps."""
import warnings
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
def fetch(url):
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
return r.url, r.text
def outline(el, depth=0, max_lines=120, lines=None):
if lines is None:
lines = []
if len(lines) >= max_lines:
return lines
name = el.name
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
label = name
if eid:
label += f'#{eid}'
if cls:
label += '.' + cls.replace(' ', '.')
if name == 'a':
txt = el.get_text(strip=True)[:50]
href = el.get('href', '')[:80]
lines.append(' ' * depth + f'{label} "{txt}"{href}')
elif name == 'button':
txt = el.get_text(strip=True)[:50]
lines.append(' ' * depth + f'{label} BTN "{txt}"')
else:
lines.append(' ' * depth + label)
for c in el.find_all(recursive=False):
if c.name in ('script', 'style'):
continue
outline(c, depth + 1, max_lines, lines)
if len(lines) >= max_lines:
return lines
return lines
# 논산시 - look at div.sitemap.type1
print('=== 논산시 div.sitemap.type1 ===')
real, html = fetch('https://nonsan.go.kr/kor/html/sub07/0701.html')
soup = BeautifulSoup(html, 'html.parser')
el = soup.select_one('div.sitemap.type1') or soup.select_one('div.sitemap')
if el:
cls = ' '.join(el.get('class', []))
print(f' Container: {el.name}.{cls} a={len(el.find_all("a"))}')
for line in outline(el, max_lines=70):
print(' ', line)
# 부여군 - try /html/kr/sitemap_pop.html or similar
print('\n\n=== 부여군 alternate sitemap probes ===')
for path in [
'/html/kr/sitemap_pop.html',
'/html/kr/sitemap.html',
'/html/kr/html/sub06/0601.html',
'/html/kr/html/sub05/0501.html',
'/html/kr/html/sub04/0401.html',
'/html/kr/html/sub09/0901.html',
'/html/kr/popup/sitemap.html',
'/html/kr/include/sitemap.html',
]:
url = f'https://www.buyeo.go.kr{path}'
try:
r = requests.get(url, headers=H, timeout=8, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
if r.status_code == 200:
s = BeautifulSoup(r.text, 'html.parser')
best_count = 0
best_sel = ''
for sel in ['.sitemap', '#sitemap', '.sitemap_grep', '.allMenu', '.pc_sitemap', 'div[class*=sitemap]']:
for el in s.select(sel):
ac = len(el.find_all('a'))
if ac > best_count:
best_count = ac
best_sel = sel
print(f' {path}: HTTP 200, best_sel={best_sel}, a={best_count}')
except Exception as e:
pass
# 부여군 main page - look for inline pc_sitemap or modal
print('\n\n=== 부여군 main page: look for inline allMenu / pc_sitemap modal ===')
real, html = fetch('https://www.buyeo.go.kr/html/kr/')
soup = BeautifulSoup(html, 'html.parser')
# Search for elements with id or class containing "sitemap"
for el in soup.select('[class*=sitemap], [id*=sitemap], [id*=allMenu], [class*=allMenu], [id*=allmenu], [class*=allmenu]'):
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
ac = len(el.find_all('a'))
print(f' {el.name}#{eid}.{cls} a={ac}')
# 아산시 - look at the .gnb-menu inline structure (use sectioned anchors mGnb-anchor1..6)
print('\n\n=== 아산시: all gnb-sub-list sections combined ===')
real, html = fetch('https://www.asan.go.kr/main/')
soup = BeautifulSoup(html, 'html.parser')
all_a = []
for sec in soup.select('div[id^=mGnb-anchor]'):
eid = sec.get('id', '')
section_a = sec.find_all('a', href=True)
print(f' {eid}: a={len(section_a)}')
for a in section_a[:3]:
print(f' "{a.get_text(strip=True)[:40]}"{a["href"][:80]}')

View File

@ -0,0 +1,35 @@
"""Look at raw HTML around h4.site01 in 보령시 to find category name."""
import re
import warnings
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
for name, url in [('보령시', 'https://www.brcn.go.kr/kor/sitemap_11.do'),
('서천군', 'https://www.seocheon.go.kr/kor/sitemap_11.do')]:
print(f'\n=== {name} === {url}')
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
html = r.text
# Find h4.siteNN occurrences
for m in re.finditer(r'<h4[^>]*site\d+[^>]*>([\s\S]*?)</h4>', html):
inner = re.sub(r'\s+', ' ', m.group(1)).strip()[:200]
full = re.sub(r'\s+', ' ', m.group(0)).strip()[:200]
print(f' h4: {full}')
# Find gnb anchors
soup = BeautifulSoup(html, 'html.parser')
gnb = soup.select_one('#gnb, #tm, nav#topmenu, nav.gnb, .gnb_wrap, .menu_wrap')
if gnb:
a_top = [a.get_text(strip=True) for a in gnb.select('> ul > li > a') or gnb.select('ul > li > a.th_1st, ul > li > a.depth1_ti, ul > li > a.first, ul > li > a')]
print(f' GNB top items: {a_top[:10]}')
# Find the top-level menu anchors (any way)
print(' Possible main category anchors:')
for cls_pattern in ['ov', 'th_1st', 'depth1_ti', 'first', 'gnb-main-trigger']:
for a in soup.find_all('a', class_=cls_pattern):
t = a.get_text(strip=True)
if t and len(t) < 30:
print(f' .{cls_pattern}: "{t}"')
break

View File

@ -0,0 +1,157 @@
"""충청북도 11개 시·군 사이트맵 URL/패턴 탐지."""
import re
import warnings
from urllib.parse import urljoin, urlparse
from concurrent.futures import ThreadPoolExecutor, as_completed
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
CITIES = [
('괴산군', 'https://www.goesan.go.kr/www/index.do'),
('단양군', 'https://www.danyang.go.kr/dy21/1'),
('보은군', 'https://www.boeun.go.kr/www/index.do'),
('영동군', 'https://www.yd21.go.kr/'),
('옥천군', 'https://www.oc.go.kr/www/'),
('음성군', 'https://www.eumseong.go.kr/www/index.do'),
('제천시', 'https://www.jecheon.go.kr/www/index.do'),
('증평군', 'https://www.jp.go.kr/kor.do'),
('진천군', 'https://www.jincheon.go.kr/home/intro.do'),
('청주시', 'https://www.cheongju.go.kr/www/index.do'),
('충주시', 'https://www.chungju.go.kr/www/index.do'),
]
def fetch(url, timeout=12):
try:
r = requests.get(url, headers=H, timeout=timeout, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
if r.status_code == 200:
return r.url, r.text
except Exception as e:
return None, f'ERR: {e}'
return None, f'HTTP {r.status_code}'
def find_sitemap_anchor(html, base):
soup = BeautifulSoup(html, 'html.parser')
found = []
for a in soup.find_all('a', href=True):
txt = a.get_text(strip=True)
href = a['href']
if not (txt and href):
continue
if '사이트맵' in txt or '전체메뉴' in txt or 'sitemap' in href.lower() or 'allMenu' in href:
full = urljoin(base, href)
found.append((txt, full))
seen = set()
uniq = []
for t, u in found:
if u not in seen and u != base:
seen.add(u)
uniq.append((t, u))
return uniq
def analyze_sitemap_page(html):
"""Score candidate containers by anchor count and class hints."""
soup = BeautifulSoup(html, 'html.parser')
candidates = []
# Common containers
for sel in [
'div.sitemap.type1', 'div.sitemap.type2', 'div.sitemap',
'div.sitemap_grep', 'div.amThum', 'ul.sitemap_list',
'#sitemap', '#contents ul.sitemap', 'ul.sitemap',
'ul.depth1_ul', 'ul.depth1-ul', 'ul.depth1',
'ul.top_menu', '#gnb', 'nav#gnb', 'nav.gnb',
'.allmenu', '.allMenu', '#allMenu',
'div.menu_all', 'div.totalMenu',
]:
for el in soup.select(sel):
ac = len(el.find_all('a'))
if ac < 30:
continue
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
candidates.append((ac, sel, f'{el.name}#{eid}.{cls}'))
# Fallback: search any container by id/class containing 'sitemap' or 'allMenu'
for el in soup.find_all(True, class_=True):
cls = ' '.join(el.get('class', []))
if re.search(r'\b(sitemap|allmenu|amthum|depth1)\b', cls, re.I):
ac = len(el.find_all('a'))
if 30 <= ac <= 2000:
candidates.append((ac, f'class~{cls[:30]}', f'{el.name}.{cls}'))
candidates.sort(reverse=True)
return candidates[:5]
def probe(name, base):
print(f'\n=== {name} === {base}')
real, html = fetch(base)
if not real:
print(f' 메인 실패: {html}')
return name, None
print(f' 메인 OK: {real}')
# 1) Find sitemap links from main page
links = find_sitemap_anchor(html, real)
print(f' 사이트맵 링크 후보: {len(links)}')
for t, u in links[:5]:
print(f' "{t}"{u}')
# 2) Common paths to try
parsed = urlparse(base)
origin = f'{parsed.scheme}://{parsed.netloc}'
common = [
'/www/sitemap.do', '/www/sub.do?key=121', '/www/contents.do?key=121',
'/kor/sitemap.do', '/kor/sitemap_11.do', '/kor/sitemap_1.do',
'/sitemap.do', '/sitemap.html',
'/www/sitemap/', '/main/sitemap.do',
'/home/sitemap.do', '/dy21/sitemap.do',
'/www/cms/sitemap.do',
]
candidates = [u for _, u in links] + [origin + p for p in common]
seen = set()
best = None
for c in candidates:
if c in seen:
continue
seen.add(c)
u2, h2 = fetch(c, timeout=10)
if not u2:
continue
info = analyze_sitemap_page(h2)
if info:
top = info[0]
if best is None or top[0] > best[1][0]:
best = (c, top, info)
if best:
url, top, info = best
print(f' ★ 사이트맵 URL: {url}')
for ac, sel, desc in info:
print(f' {sel}: a={ac} {desc[:80]}')
return name, {'url': url, 'best_sel': info[0][1], 'a_count': info[0][0]}
print(' 사이트맵 못 찾음')
return name, None
def main():
results = {}
with ThreadPoolExecutor(max_workers=6) as ex:
futs = {ex.submit(probe, n, b): n for n, b in CITIES}
for f in as_completed(futs):
n, info = f.result()
results[n] = info
print('\n=== 요약 ===')
for n, _ in CITIES:
info = results.get(n)
if info:
print(f' {n}: {info["url"]} [{info["best_sel"]}, a={info["a_count"]}]')
else:
print(f' {n}: NOT FOUND')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,135 @@
"""충청북도 추가 분석:
1) depth1 패턴 구조 outline (괴산·청주·충주 샘플)
2) 단양·보은·영동·진천 사이트맵 재탐색
"""
import re
import warnings
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
def fetch(url, timeout=15):
try:
r = requests.get(url, headers=H, timeout=timeout, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
return r.status_code, r.url, r.text
except Exception as e:
return 0, str(e), ''
def outline(el, d=0, max_lines=60, lines=None):
if lines is None: lines = []
if len(lines) >= max_lines: return lines
name = el.name
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
lbl = name
if eid: lbl += f'#{eid}'
if cls: lbl += '.' + cls.replace(' ', '.')
if name == 'a':
t = el.get_text(strip=True)[:40]
h = el.get('href', '')[:80]
lines.append(' ' * d + f'{lbl} "{t}"{h}')
else:
lines.append(' ' * d + lbl)
for c in el.find_all(recursive=False):
if c.name in ('script', 'style'): continue
outline(c, d+1, max_lines, lines)
if len(lines) >= max_lines: return lines
return lines
# (1) Depth1 패턴 outline
print('='*70)
print('Depth1 pattern — 괴산군')
print('='*70)
_, _, html = fetch('https://www.goesan.go.kr/www/sitemap.do?key=28')
soup = BeautifulSoup(html, 'html.parser')
for sel in ['div.depth1', 'div.depth.depth1', '#sitemap', '.sitemap']:
el = soup.select_one(sel)
if el:
print(f'\n>>> selector: {sel} (a={len(el.find_all("a"))})')
for line in outline(el, max_lines=50):
print(' ', line)
break
print('\n' + '='*70)
print('Depth1 pattern — 청주시')
print('='*70)
_, _, html = fetch('https://www.cheongju.go.kr/www/sitemap.do?key=589')
soup = BeautifulSoup(html, 'html.parser')
for sel in ['#sitemap div.sitemap', 'div#sitemap.sitemap', 'div.sitemap']:
el = soup.select_one(sel)
if el:
print(f'\n>>> selector: {sel} (a={len(el.find_all("a"))})')
for line in outline(el, max_lines=50):
print(' ', line)
break
print('\n' + '='*70)
print('Depth1 pattern — 충주시')
print('='*70)
_, _, html = fetch('https://www.chungju.go.kr/www/sub.do?key=692')
soup = BeautifulSoup(html, 'html.parser')
for sel in ['#sitemap']:
el = soup.select_one(sel)
if el:
print(f'\n>>> selector: {sel} (a={len(el.find_all("a"))})')
for line in outline(el, max_lines=50):
print(' ', line)
break
# (2) 실패한 사이트들 재탐색
print('\n\n' + '='*70)
print('실패 사이트 재탐색')
print('='*70)
for name, url in [
('단양군', 'https://www.danyang.go.kr/dy21/1'),
('보은군', 'https://www.boeun.go.kr/www/index.do'),
('영동군', 'https://www.yd21.go.kr/'),
('진천군', 'https://www.jincheon.go.kr/home/intro.do'),
]:
print(f'\n--- {name} {url} ---')
code, real, html = fetch(url)
if code != 200:
print(f' 메인 실패: {code} {real}')
continue
print(f' 메인 OK: {real}')
# Find sitemap link
soup = BeautifulSoup(html, 'html.parser')
cands = []
for a in soup.find_all('a', href=True):
txt = a.get_text(strip=True)
if '사이트맵' in txt or '전체메뉴' in txt or '누리집' in txt or 'sitemap' in (a.get('href','')+txt).lower():
cands.append((txt, a['href']))
# Deduplicate
seen = set()
uniq = []
from urllib.parse import urljoin
for t, h in cands:
full = urljoin(real, h)
if full not in seen and full != real:
seen.add(full)
uniq.append((t, full))
for t, u in uniq[:5]:
print(f' 사이트맵 후보: "{t}"{u}')
# Test first candidate
if uniq:
c, c_url = uniq[0]
c2, r2, h2 = fetch(c_url)
if c2 == 200:
s2 = BeautifulSoup(h2, 'html.parser')
best = (0, '', '')
for sel in ['div.depth1', 'div.depth.depth1', 'ul.depth1_ul', 'ul.depth1-ul',
'#sitemap', 'div.sitemap_grep', 'ul.sitemap', 'div.sitemap',
'div.sitemap_11', 'div.amThum']:
for el in s2.select(sel):
ac = len(el.find_all('a'))
if ac > best[0]:
best = (ac, sel, ' '.join(el.get('class', [])))
print(f' 사이트맵 분석: best_sel={best[1]} cls={best[2]} a={best[0]}')

View File

@ -0,0 +1,213 @@
"""실패한 4개 사이트 (단양·보은·영동·진천) 심층 탐색."""
import re
import ssl
import warnings
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.ssl_ import create_urllib3_context
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
class WeakSSLAdapter(HTTPAdapter):
"""레거시 SSL/TLS handshake를 허용하는 어댑터 (영동군 등 구형 SSL)."""
def init_poolmanager(self, *args, **kwargs):
ctx = create_urllib3_context()
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
ctx.options |= 0x4 # ssl.OP_LEGACY_SERVER_CONNECT (Python 3.12+)
ctx.check_hostname = False
ctx.verify_mode = ssl.CERT_NONE
kwargs['ssl_context'] = ctx
return super().init_poolmanager(*args, **kwargs)
def make_session(weak_ssl=False):
s = requests.Session()
s.headers.update(H)
if weak_ssl:
s.mount('https://', WeakSSLAdapter())
return s
def fetch(url, session=None, timeout=20):
s = session or requests.Session()
if not session:
s.headers.update(H)
try:
r = s.get(url, timeout=timeout, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
return r.status_code, r.url, r.text
except Exception as e:
return 0, str(e), ''
def outline(el, d=0, max_lines=40, lines=None):
if lines is None: lines = []
if len(lines) >= max_lines: return lines
name = el.name
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
lbl = name
if eid: lbl += f'#{eid}'
if cls: lbl += '.' + cls.replace(' ', '.')
if name == 'a':
t = el.get_text(strip=True)[:40]
h = el.get('href', '')[:80]
lines.append(' ' * d + f'{lbl} "{t}"{h}')
else:
lines.append(' ' * d + lbl)
for c in el.find_all(recursive=False):
if c.name in ('script', 'style'): continue
outline(c, d+1, max_lines, lines)
if len(lines) >= max_lines: return lines
return lines
def analyze(html):
soup = BeautifulSoup(html, 'html.parser')
best = (0, '', '', None)
for sel in [
'div.depth.depth1', 'div.depth1', '#sitemap div.site_map_col', '#sitemap',
'div.sitemap_grep', 'ul.sitemap_list', 'div.sitemap_box',
'ul.depth1_ul', 'ul.depth1-ul', 'ul.depth1',
'ul.sitemap', 'div.sitemap', 'div.amThum', '.allMenu', 'div.menu_all',
'nav#gnb', 'nav.gnb', '#gnb',
]:
for el in soup.select(sel):
ac = len(el.find_all('a'))
if ac > best[0]:
cls = ' '.join(el.get('class', []))
best = (ac, sel, cls, el)
return best
# 단양군 - try /dy21/98
print('='*70)
print('단양군 — /dy21/98')
print('='*70)
sess = make_session()
code, real, html = fetch('https://www.danyang.go.kr/dy21/98', sess)
print(f' HTTP {code}, len {len(html)}')
if code == 200:
soup = BeautifulSoup(html, 'html.parser')
# Search for the sitemap container
print(f' All elements w/ a>=50:')
for el in soup.find_all(['div', 'ul', 'nav']):
ac = len(el.find_all('a'))
if 50 <= ac:
cls = ' '.join(el.get('class', []))[:50]
eid = el.get('id', '')
print(f' {el.name}#{eid}.{cls} a={ac}')
best = analyze(html)
if best[3]:
print(f' Best: {best[1]} cls={best[2]} a={best[0]}')
for line in outline(best[3], max_lines=40):
print(' ', line)
# 진천군 - try /home/main.do
print('\n' + '='*70)
print('진천군 — /home/main.do')
print('='*70)
code, real, html = fetch('https://www.jincheon.go.kr/home/main.do', sess)
print(f' HTTP {code}, len {len(html)}')
if code == 200:
soup = BeautifulSoup(html, 'html.parser')
cands = []
for a in soup.find_all('a', href=True):
txt = a.get_text(strip=True)
if '사이트맵' in txt or '전체메뉴' in txt or '누리집' in txt:
from urllib.parse import urljoin
cands.append((txt, urljoin(real, a['href'])))
print(' 사이트맵 후보:')
for t, u in cands[:8]:
print(f' "{t}"{u}')
# Common paths
from urllib.parse import urlparse
parsed = urlparse(real)
origin = f'{parsed.scheme}://{parsed.netloc}'
test_urls = [u for _, u in cands] + [
origin + '/home/sitemap.do',
origin + '/home/contents.do?key=121',
origin + '/home/sub.do?key=121',
]
for url in test_urls:
c2, r2, h2 = fetch(url, sess)
if c2 != 200:
continue
best = analyze(h2)
if best[0] >= 50:
print(f'{url}: {best[1]} cls={best[2]} a={best[0]}')
# 보은군 - retry
print('\n' + '='*70)
print('보은군 — retry')
print('='*70)
code, real, html = fetch('https://www.boeun.go.kr/www/index.do', sess)
print(f' HTTP {code}, len {len(html)}')
if code == 200:
soup = BeautifulSoup(html, 'html.parser')
cands = []
for a in soup.find_all('a', href=True):
txt = a.get_text(strip=True)
if '사이트맵' in txt or '전체메뉴' in txt or '누리집 지도' in txt:
from urllib.parse import urljoin
cands.append((txt, urljoin(real, a['href'])))
print(' 사이트맵 후보:')
for t, u in cands[:8]:
print(f' "{t}"{u}')
# Try probable paths
from urllib.parse import urlparse
parsed = urlparse(real)
origin = f'{parsed.scheme}://{parsed.netloc}'
test_urls = [u for _, u in cands] + [
origin + '/www/sitemap.do',
origin + '/www/sub.do?key=121',
]
for url in set(test_urls):
c2, r2, h2 = fetch(url, sess)
if c2 != 200:
continue
best = analyze(h2)
if best[0] >= 50:
print(f'{url}: {best[1]} cls={best[2]} a={best[0]}')
# 영동군 - try with weak SSL adapter
print('\n' + '='*70)
print('영동군 — weak SSL adapter')
print('='*70)
sess_weak = make_session(weak_ssl=True)
code, real, html = fetch('https://www.yd21.go.kr/', sess_weak)
print(f' HTTP {code}, len {len(html)}')
if code == 200:
soup = BeautifulSoup(html, 'html.parser')
cands = []
for a in soup.find_all('a', href=True):
txt = a.get_text(strip=True)
if '사이트맵' in txt or '전체메뉴' in txt or '누리집' in txt:
from urllib.parse import urljoin
cands.append((txt, urljoin(real, a['href'])))
print(' 사이트맵 후보:')
for t, u in cands[:8]:
print(f' "{t}"{u}')
from urllib.parse import urlparse
parsed = urlparse(real)
origin = f'{parsed.scheme}://{parsed.netloc}'
test_urls = [u for _, u in cands] + [
origin + '/kor/sitemap.do',
origin + '/sitemap.html',
origin + '/sitemap.do',
origin + '/contents/contents.html?cid=2151',
]
for url in set(test_urls):
c2, r2, h2 = fetch(url, sess_weak)
if c2 != 200:
continue
best = analyze(h2)
if best[0] >= 50:
print(f'{url}: {best[1]} cls={best[2]} a={best[0]}')

View File

@ -0,0 +1,86 @@
"""단양·진천 구조 outline + 보은군 alternative."""
import warnings
import socket
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
def fetch(url, timeout=20):
try:
r = requests.get(url, headers=H, timeout=timeout, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
return r.status_code, r.url, r.text
except Exception as e:
return 0, str(e), ''
def outline(el, d=0, max_lines=80, lines=None):
if lines is None: lines = []
if len(lines) >= max_lines: return lines
name = el.name
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
lbl = name
if eid: lbl += f'#{eid}'
if cls: lbl += '.' + cls.replace(' ', '.')
if name == 'a':
t = el.get_text(strip=True)[:40]
h = el.get('href', '')[:80]
lines.append(' ' * d + f'{lbl} "{t}"{h}')
else:
lines.append(' ' * d + lbl)
for c in el.find_all(recursive=False):
if c.name in ('script', 'style'): continue
outline(c, d+1, max_lines, lines)
if len(lines) >= max_lines: return lines
return lines
# 단양군 - outline #menu_sitemap
print('='*70)
print('단양군 — #menu_sitemap outline')
print('='*70)
_, _, html = fetch('https://www.danyang.go.kr/dy21/98')
soup = BeautifulSoup(html, 'html.parser')
el = soup.select_one('#menu_sitemap') or soup.select_one('#contents_sitemap')
if el:
print(f' Container a={len(el.find_all("a"))}')
for line in outline(el, max_lines=60):
print(' ', line)
# 진천군 - outline nav#gnb on sub.do?menukey=445
print('\n' + '='*70)
print('진천군 — sub.do?menukey=445 nav#gnb outline')
print('='*70)
_, _, html = fetch('https://www.jincheon.go.kr/home/sub.do?menukey=445')
soup = BeautifulSoup(html, 'html.parser')
el = soup.select_one('nav#gnb') or soup.select_one('#gnb')
if el:
print(f' Container a={len(el.find_all("a"))}')
for line in outline(el, max_lines=60):
print(' ', line)
# 보은군 — try alternate hosts
print('\n' + '='*70)
print('보은군 — DNS/ alternate hosts')
print('='*70)
for host in ['www.boeun.go.kr', 'boeun.go.kr', 'boeun.chungbuk.go.kr']:
try:
ip = socket.gethostbyname(host)
print(f' {host}{ip}')
except Exception as e:
print(f' {host}: {e}')
# Try fetching via curl-style direct
for url in ['https://www.boeun.go.kr/www/index.do',
'http://www.boeun.go.kr/www/index.do',
'https://boeun.go.kr/www/index.do',
'https://www.boeun.go.kr/']:
code, real, _ = fetch(url, timeout=10)
print(f' {url} → HTTP {code} ({real[:60]})')

View File

@ -0,0 +1,61 @@
"""진천군: nav#gnb 더 깊이 들어가서 실제 메뉴 찾기. + 보은군 한번 더."""
import re
import warnings
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
def fetch(url, timeout=20):
try:
r = requests.get(url, headers=H, timeout=timeout, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
return r.status_code, r.url, r.text
except Exception as e:
return 0, str(e), ''
# 진천군 - find all elements with menu-like classes
print('='*70)
print('진천군 — find menu containers')
print('='*70)
_, _, html = fetch('https://www.jincheon.go.kr/home/sub.do?menukey=445')
soup = BeautifulSoup(html, 'html.parser')
# Look for elements with 'depth' / 'gnb' / 'menu' / 'sitemap' in classes
print('All elements with 100+ anchors and depth/menu/sitemap in class:')
for el in soup.find_all(['div', 'ul', 'nav', 'section']):
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
ac = len(el.find_all('a'))
if ac >= 100 and re.search(r'depth|menu|sitemap|allmenu|total|all-menu|gnb', cls + ' ' + eid, re.I):
print(f' {el.name}#{eid}.{cls[:60]} a={ac}')
# Find all .depth* classes
print('\n.depth* classes with anchors:')
for el in soup.select('[class*=depth]'):
cls = ' '.join(el.get('class', []))
ac = len(el.find_all('a'))
if ac >= 50:
print(f' {el.name}.{cls[:80]} a={ac}')
# 보은군 한번 더
print('\n' + '='*70)
print('보은군 retry')
print('='*70)
import time
time.sleep(1)
code, real, html = fetch('https://www.boeun.go.kr/www/index.do')
print(f' {code} {real[:100]}')
if code == 200:
print(' 성공! 사이트맵 후보 탐색')
from urllib.parse import urljoin
s2 = BeautifulSoup(html, 'html.parser')
for a in s2.find_all('a', href=True)[:200]:
txt = a.get_text(strip=True)
if '사이트맵' in txt or '전체메뉴' in txt or '누리집 지도' in txt:
print(f' "{txt}"{urljoin(real, a["href"])}')

View File

@ -0,0 +1,39 @@
"""진천군 div.sitemap outline."""
import warnings
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
r = requests.get('https://www.jincheon.go.kr/home/sub.do?menukey=445', headers=H, timeout=15, verify=False)
r.encoding = r.apparent_encoding
soup = BeautifulSoup(r.text, 'html.parser')
def outline(el, d=0, max_lines=70, lines=None):
if lines is None: lines = []
if len(lines) >= max_lines: return lines
name = el.name
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
lbl = name
if eid: lbl += f'#{eid}'
if cls: lbl += '.' + cls.replace(' ', '.')
if name == 'a':
t = el.get_text(strip=True)[:40]
h = el.get('href', '')[:80]
lines.append(' ' * d + f'{lbl} "{t}"{h}')
else:
lines.append(' ' * d + lbl)
for c in el.find_all(recursive=False):
if c.name in ('script', 'style'): continue
outline(c, d+1, max_lines, lines)
if len(lines) >= max_lines: return lines
return lines
el = soup.select_one('div.sitemap') or soup.select_one('ul.gnb-wrap')
if el:
print(f'Container: {el.name}.{" ".join(el.get("class",[]))} a={len(el.find_all("a"))}')
for line in outline(el, max_lines=60):
print(' ', line)

View File

@ -0,0 +1,74 @@
"""Inspect e-Gov sitemap_grep structure — count amThum sections and h2 text."""
import warnings
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
SITES = {
'당진시': 'https://www.dangjin.go.kr/kor/sitemap_11.do',
'보령시': 'https://www.brcn.go.kr/kor/sitemap_11.do',
'서천군': 'https://www.seocheon.go.kr/kor/sitemap_11.do',
'청양군': 'https://www.cheongyang.go.kr/kor/sitemap_11.do',
'태안군': 'https://www.taean.go.kr/kor/sitemap_11.do',
'서산시(top_menu)': 'https://www.seosan.go.kr/www/index.do',
'예산군': 'https://www.yesan.go.kr/kor/sitemap.do',
'천안시': 'https://www.cheonan.go.kr/kor/sitemap.do',
'홍성군': 'https://www.hongseong.go.kr/kor/sitemap.do',
}
for name, url in SITES.items():
print(f'\n=== {name} ===')
try:
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
except Exception as e:
print(f' ERR: {e}')
continue
soup = BeautifulSoup(r.text, 'html.parser')
# Find amThum sections (eGov sitemap_grep)
amthums = soup.select('div.amThum')
if amthums:
print(f' amThum 섹션: {len(amthums)}')
for at in amthums[:10]:
h2 = at.find('h2')
h2_text = h2.get_text(strip=True) if h2 else ''
grep = at.find('div', class_='sitemap_grep')
n_list = len(grep.find_all('ul', class_='sitemap_list')) if grep else 0
n_first = len(at.find_all('a', class_='first'))
print(f' "{h2_text}" — sitemap_list={n_list}, a.first={n_first}')
continue
# Holsung 패턴 — div.sitemap.type2.nN > dl > dt + dd
sm = soup.select_one('div.sitemap[class*=type2]')
if sm:
dls = sm.find_all('dl', recursive=False)
print(f' sitemap type2 dl: {len(dls)}')
for dl in dls[:10]:
dt = dl.find('dt')
dt_text = dt.get_text(strip=True) if dt else ''
dds = dl.find_all('dd', recursive=False)
print(f' "{dt_text}" — dd={len(dds)}')
continue
# Yesan 패턴 — ul.depth1_ul
dep1 = soup.select('ul.depth1_ul > li, ul.depth1-ul > li')
if dep1:
print(f' depth1 li: {len(dep1)}')
for li in dep1[:10]:
a = li.find(['a', 'button'], recursive=False) or li.find(['a', 'button'])
if a:
txt = a.get_text(strip=True)
print(f' "{txt}"')
continue
# Seosan top_menu pattern
tm = soup.select_one('ul.top_menu')
if tm:
deps = tm.find_all('li', class_='depth1', recursive=False)
print(f' top_menu li.depth1: {len(deps)}')
for li in deps[:10]:
a = li.find('a', class_='depth1_ti')
if a:
print(f' "{a.get_text(strip=True)}"')
continue
print(' Unknown pattern')

View File

@ -0,0 +1,98 @@
"""Targeted probes for 청양군, 아산시, 부여군 to fix parsers."""
import warnings
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
def fetch(url):
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
return r.text
# 청양군 — look at where the actual sitemap content lives
print('=== 청양군 ===')
html = fetch('https://www.cheongyang.go.kr/kor/sitemap_11.do')
soup = BeautifulSoup(html, 'html.parser')
# Look for #contents or main content
for sel in ['#contents', '.contents', 'main', '#txt', '#mainSection']:
el = soup.select_one(sel)
if el:
a_count = len(el.find_all('a'))
print(f' {sel}: a={a_count}')
# Look for the right sitemap container
print(' All elements with 200+ anchors:')
for el in soup.find_all(['div', 'ul', 'section']):
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
ac = len(el.find_all('a'))
if 200 <= ac <= 800 and (cls or eid):
print(f' {el.name}#{eid}.{cls[:60]} a={ac}')
print('\n=== 아산시 ===')
html = fetch('https://www.asan.go.kr/main/')
soup = BeautifulSoup(html, 'html.parser')
# Inspect mGnb-anchor1 structure deeply
sec = soup.find('div', id='mGnb-anchor1')
if sec:
# outline
def outline(el, d=0, max_lines=50, lines=None):
if lines is None: lines = []
if len(lines) >= max_lines: return lines
name = el.name
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
lbl = name
if eid: lbl += f'#{eid}'
if cls: lbl += '.' + cls.replace(' ', '.')
if name == 'a':
t = el.get_text(strip=True)[:30]
h = el.get('href', '')[:60]
lines.append(' ' * d + f'{lbl} "{t}"{h}')
elif name == 'button':
t = el.get_text(strip=True)[:30]
lines.append(' ' * d + f'{lbl} BTN "{t}"')
else:
lines.append(' ' * d + lbl)
for c in el.find_all(recursive=False):
if c.name in ('script', 'style'): continue
outline(c, d+1, max_lines, lines)
if len(lines) >= max_lines: return lines
return lines
for line in outline(sec, max_lines=70):
print(' ', line)
print('\n=== 부여군 ===')
html = fetch('https://www.buyeo.go.kr/html/kr/')
soup = BeautifulSoup(html, 'html.parser')
# Inspect nav#gnb structure
sec = soup.select_one('nav#gnb')
if sec:
def outline(el, d=0, max_lines=80, lines=None):
if lines is None: lines = []
if len(lines) >= max_lines: return lines
name = el.name
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
lbl = name
if eid: lbl += f'#{eid}'
if cls: lbl += '.' + cls.replace(' ', '.')
if name == 'a':
t = el.get_text(strip=True)[:30]
h = el.get('href', '')[:60]
lines.append(' ' * d + f'{lbl} "{t}"{h}')
else:
lines.append(' ' * d + lbl)
for c in el.find_all(recursive=False):
if c.name in ('script', 'style'): continue
outline(c, d+1, max_lines, lines)
if len(lines) >= max_lines: return lines
return lines
for line in outline(sec, max_lines=80):
print(' ', line)

View File

@ -0,0 +1,60 @@
"""제천시·증평군 구조 재확인."""
import warnings
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
def outline(el, d=0, max_lines=70, lines=None):
if lines is None: lines = []
if len(lines) >= max_lines: return lines
name = el.name
cls = ' '.join(el.get('class', []))
eid = el.get('id', '')
lbl = name
if eid: lbl += f'#{eid}'
if cls: lbl += '.' + cls.replace(' ', '.')
if name == 'a':
t = el.get_text(strip=True)[:40]
h = el.get('href', '')[:80]
lines.append(' ' * d + f'{lbl} "{t}"{h}')
else:
lines.append(' ' * d + lbl)
for c in el.find_all(recursive=False):
if c.name in ('script', 'style'): continue
outline(c, d+1, max_lines, lines)
if len(lines) >= max_lines: return lines
return lines
# 제천시
print('='*70)
print('제천시 — div.depth1 outline')
print('='*70)
r = requests.get('https://www.jecheon.go.kr/www/sitemap.do?key=553', headers=H, timeout=15, verify=False)
r.encoding = r.apparent_encoding
soup = BeautifulSoup(r.text, 'html.parser')
for sel in ['div.depth1', 'div.depth.depth1', '#sitemap']:
el = soup.select_one(sel)
if el:
print(f'\n>>> {sel} (a={len(el.find_all("a"))}, classes={el.get("class")})')
for line in outline(el, max_lines=50):
print(' ', line)
break
# 증평군 - ul.depth1_ul outline
print('\n' + '='*70)
print('증평군 — ul.depth1_ul outline')
print('='*70)
r = requests.get('https://www.jp.go.kr/kor/sitemap_11.do', headers=H, timeout=15, verify=False)
r.encoding = r.apparent_encoding
soup = BeautifulSoup(r.text, 'html.parser')
el = soup.select_one('ul.depth1_ul')
if el:
print(f' (a={len(el.find_all("a"))}, classes={el.get("class")})')
for line in outline(el, max_lines=60):
print(' ', line)

View File

@ -0,0 +1,181 @@
"""전북·제주 16개 시·군 사이트맵 URL 탐색 + 컨테이너 구조 덤프."""
import re
import ssl
import sys
import warnings
from urllib.parse import urljoin, urlparse
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.ssl_ import create_urllib3_context
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
class WeakSSLAdapter(HTTPAdapter):
def init_poolmanager(self, *args, **kwargs):
ctx = create_urllib3_context()
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
ctx.options |= 0x4
ctx.check_hostname = False
ctx.verify_mode = ssl.CERT_NONE
kwargs['ssl_context'] = ctx
return super().init_poolmanager(*args, **kwargs)
def make_session(weak=False):
s = requests.Session()
s.headers.update(H)
if weak:
s.mount('https://', WeakSSLAdapter())
return s
def fetch(session, url, timeout=20):
r = session.get(url, timeout=timeout, verify=False, allow_redirects=True)
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
if meta:
r.encoding = meta.group(1).decode('ascii', errors='ignore')
else:
r.encoding = r.apparent_encoding
return r
SITES = [
('고창군', 'https://www.gochang.go.kr/index.gochang?contentsSid=3136'),
('군산시', 'https://www.gunsan.go.kr/main'),
('김제시', 'https://www.gimje.go.kr/index.gimje'),
('남원시', 'https://www.namwon.go.kr/index.do?menuUid=ff8080818e3beff0018e40e8f63e02d2'),
('무주군', 'https://www.muju.go.kr/index.9is'),
('부안군', 'https://www.buan.go.kr/index.buan?contentsSid=1'),
('순창군', 'https://www.sunchang.go.kr/'),
('완주군', 'https://www.wanju.go.kr/index.9is'),
('익산시', 'https://www.iksan.go.kr/index.do?menuUid=ff8080819a39930e019a4de8c1ae0afd'),
('임실군', 'https://www.imsil.go.kr/index.imsil'),
('장수군', 'https://www.jangsu.go.kr/index.jangsu'),
('전주시', 'https://www.jeonju.go.kr/index.9is'),
('정읍시', 'https://www.jeongeup.go.kr/index.jeongeup'),
('진안군', 'https://www.jinan.go.kr/index.jinan?contentsSid=1379'),
('서귀포시', 'https://www.seogwipo.go.kr/index.htm'),
('제주시', 'https://www.jejusi.go.kr/index.ac'),
]
def find_sitemap_links(soup, base):
"""페이지에서 사이트맵으로 보이는 링크 후보 수집."""
cands = []
for a in soup.find_all('a', href=True):
txt = (a.get_text() or '').strip()
href = a['href']
onclick = a.get('onclick', '')
blob = f'{txt} {href} {onclick}'.lower()
if '사이트맵' in txt or 'sitemap' in blob:
url = href
if url.startswith('#') or url.lower().startswith('javascript:'):
# onclick에서 추출 시도
m = re.search(r"""['"]([^'"]*(?:sitemap|site_map)[^'"]*)['"]""", onclick, re.I)
if m:
url = m.group(1)
else:
continue
cands.append((txt, urljoin(base, url)))
# 중복 제거
seen = set(); out = []
for t, u in cands:
if u not in seen:
seen.add(u); out.append((t, u))
return out
def dump_structure(soup):
"""사이트맵 페이지 주요 컨테이너 후보 출력."""
# 흔한 사이트맵 컨테이너 셀렉터 후보
sels = [
'div.sitemap', 'div#sitemap', 'div.site_map', 'div#site_map',
'ul#menu_sitemap', 'div.depth.depth1', 'div.depth1',
'ul.depth1_ul', 'div.amThum', 'ul.sitemap', 'div.sitemap_wrap',
'div.contents_sitemap', 'div.sitemapWrap', 'div.site-map',
'div.allmenu', 'div#allmenu', 'div.gnb_all', 'div.total_menu',
]
found = []
for sel in sels:
els = soup.select(sel)
if els:
found.append((sel, len(els)))
print(f' 매칭 셀렉터: {found}')
# 사이트맵스러운 컨테이너 한 개 잡아서 자식 구조 덤프
target = None
for sel, _ in found:
target = soup.select_one(sel)
if target:
print(f' >>> {sel} 내부 구조:')
break
if not target:
# body에서 class에 sitemap/menu/depth 포함 div 찾기
for div in soup.find_all(['div', 'ul'], class_=True):
cls = ' '.join(div.get('class', []))
if re.search(r'sitemap|site_map|allmenu|depth1|total_menu', cls, re.I):
target = div
print(f' >>> <{div.name} class="{cls}"> 내부 구조:')
break
if not target:
print(' !! 사이트맵 컨테이너 미발견')
return
# 자식 1~3레벨 태그/클래스 요약
def summarize(el, depth=0, maxdepth=4):
if depth > maxdepth:
return
for child in el.find_all(recursive=False):
cls = '.'.join(child.get('class', []))
idv = child.get('id', '')
tag = child.name
label = tag + (f'#{idv}' if idv else '') + (f'.{cls}' if cls else '')
a = child.find('a', recursive=False)
atxt = (a.get_text().strip()[:20] if a else '')
print(' ' * (depth + 1) + f'{label}' + (f' a="{atxt}"' if atxt else ''))
if depth < 2:
summarize(child, depth + 1, maxdepth)
summarize(target)
def main():
targets = sys.argv[1:]
for name, url in SITES:
if targets and name not in targets:
continue
print(f'\n{"="*70}\n[{name}] {url}')
for weak in (False, True):
try:
sess = make_session(weak=weak)
r = fetch(sess, url)
base = f'{urlparse(r.url).scheme}://{urlparse(r.url).netloc}'
soup = BeautifulSoup(r.text, 'html.parser')
print(f' status={r.status_code} final={r.url} weak_ssl={weak}')
links = find_sitemap_links(soup, base)
print(f' 사이트맵 링크 후보: {links[:6]}')
# 가장 그럴듯한 후보 따라가기
if links:
smurl = links[0][1]
try:
r2 = fetch(sess, smurl)
soup2 = BeautifulSoup(r2.text, 'html.parser')
print(f' 사이트맵 페이지: {r2.url} (status {r2.status_code})')
dump_structure(soup2)
except Exception as e:
print(f' 사이트맵 페이지 fetch 실패: {e}')
else:
print(' >>> 인덱스 자체 구조 확인:')
dump_structure(soup)
break
except Exception as e:
if weak:
print(f' !! 실패(weak 포함): {type(e).__name__}: {e}')
else:
print(f' (일반 SSL 실패 → weak 재시도): {type(e).__name__}')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,84 @@
"""전북 미해결 사이트 2차 정밀 탐색."""
import re, ssl, sys, warnings
from urllib.parse import urljoin, urlparse
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.ssl_ import create_urllib3_context
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
def sess():
s = requests.Session(); s.headers.update({'User-Agent': UA}); return s
def fetch(s, url, t=20):
r = s.get(url, timeout=t, verify=False, allow_redirects=True)
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
r.encoding = meta.group(1).decode('ascii','ignore') if meta else r.apparent_encoding
return r
# 사이트맵 링크를 못 찾은 사이트: 인덱스 HTML에서 sitemap/menuCd/전체메뉴 힌트 검색
NOLINK = {
'군산시': 'https://www.gunsan.go.kr/main',
'남원시': 'https://www.namwon.go.kr/index.do?menuUid=ff8080818e3beff0018e40e8f63e02d2',
'부안군': 'https://www.buan.go.kr/index.buan?contentsSid=1',
'완주군': 'https://www.wanju.go.kr/index.9is',
'임실군': 'https://www.imsil.go.kr/index.imsil',
'정읍시': 'https://www.jeongeup.go.kr/index.jeongeup',
}
def probe_links(name, url):
print(f'\n{"="*70}\n[{name}] {url}')
s = sess()
try:
r = fetch(s, url)
except Exception as e:
print(' fetch fail', e); return
base = f'{urlparse(r.url).scheme}://{urlparse(r.url).netloc}'
soup = BeautifulSoup(r.text, 'html.parser')
# sitemap/전체메뉴 텍스트나 href를 가진 a 전부
hits = []
for a in soup.find_all('a'):
txt = (a.get_text() or '').strip()
href = a.get('href','') or ''
oc = a.get('onclick','') or ''
blob = f'{txt}|{href}|{oc}'
if re.search(r'사이트맵|전체메뉴|sitemap|site_map|allmenu', blob, re.I):
hits.append((txt[:20], href[:90], oc[:90]))
for h in hits[:15]:
print(' a:', h)
# menuCd 패턴 가진 href 중 sitemap 후보 (DOM_...02000000 류)
cds = set(re.findall(r'menuCd=DOM_\d+', r.text))
print(' menuCd 샘플:', list(cds)[:10])
for n, u in NOLINK.items():
probe_links(n, u)
# ---- 구조 깊이 확인이 필요한 사이트들 ----
DEEP = {
'고창군': 'https://www.gochang.go.kr/index.gochang?menuCd=DOM_000000106006000000',
'익산시': 'https://www.iksan.go.kr/index.do?menuUid=ff80808199f0d11c019a041b8e35174a',
'전주시': 'https://www.jeonju.go.kr/index.9is?contentUid=ff8080818c7c2e8e018c7fd67b9a03b6',
'순창군': 'https://www.sunchang.go.kr/',
'진안군': 'https://www.jinan.go.kr/index.jinan?menuCd=DOM_000000110002000000',
}
def deep(name, url, sel):
print(f'\n{"#"*70}\n[{name}] DEEP {url} sel={sel}')
s = sess()
try:
r = fetch(s, url)
except Exception as e:
print(' fail', e); return
soup = BeautifulSoup(r.text, 'html.parser')
cont = soup.select_one(sel)
if not cont:
print(' 컨테이너 없음'); return
# 첫 블록 하나만 골라 li/a href까지 자세히
print(cont.prettify()[:2500])
deep('고창군 첫메뉴', 'https://www.gochang.go.kr/index.gochang?menuCd=DOM_000000106006000000', 'div.sitemap div.menu1')
deep('익산시 group수', 'https://www.iksan.go.kr/index.do?menuUid=ff80808199f0d11c019a041b8e35174a', 'div.sitemap_group')
deep('전주시', 'https://www.jeonju.go.kr/index.9is?contentUid=ff8080818c7c2e8e018c7fd67b9a03b6', 'div.sitemap_Warp')
deep('순창군 gnb', 'https://www.sunchang.go.kr/', 'div.sitemap_box ul.gnb_list')
deep('진안군 첫dl', 'https://www.jinan.go.kr/index.jinan?menuCd=DOM_000000110002000000', 'div.sitemap dl')

View File

@ -0,0 +1,99 @@
"""전북 미해결 사이트 3차: 군산/남원/부안/완주/임실/순창 사이트맵 URL 확정."""
import re, sys, warnings
from urllib.parse import urljoin, urlparse
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
def sess():
s = requests.Session(); s.headers.update({'User-Agent': UA}); return s
def fetch(s, url, t=20):
r = s.get(url, timeout=t, verify=False, allow_redirects=True)
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
r.encoding = meta.group(1).decode('ascii','ignore') if meta else r.apparent_encoding
return r
def show_container(soup, label):
for sel in ['div.sitemap','div#sitemap','div.sitemap_group','div.sitemap_Warp',
'div.sitemap_wrap','ul.siteMapList','div.allmenu','div#allmenu',
'div.contents','div.allmenubox','div.total_menu','div.site_map']:
els = soup.select(sel)
if els:
print(f' [{label}] sel={sel} x{len(els)}')
S = sess()
# 정읍 확정: 전체메뉴보기 menuCd
r = fetch(S, 'https://www.jeongeup.go.kr/index.jeongeup?menuCd=DOM_000000106002000000')
soup = BeautifulSoup(r.text, 'html.parser')
print('=== 정읍 사이트맵 ===', r.url)
show_container(soup, '정읍')
cont = soup.select_one('div.sitemap')
if cont:
blocks = cont.find_all('div', recursive=False)
print(' div.sitemap 직계 div:', len(blocks), [ '.'.join(b.get('class',[])) for b in blocks[:8]])
if blocks:
print(blocks[0].prettify()[:800])
# 부안/임실: index 페이지에서 "사이트맵/누리집지도/전체메뉴" 텍스트 가진 a 또는 그 주변 menuCd
for name, url in [('부안군','https://www.buan.go.kr/index.buan?contentsSid=1'),
('임실군','https://www.imsil.go.kr/index.imsil')]:
r = fetch(S, url)
soup = BeautifulSoup(r.text, 'html.parser')
print(f'\n=== {name} index — 사이트맵류 a ===')
for a in soup.find_all('a'):
t = (a.get_text() or '').strip()
if re.search(r'사이트맵|누리집지도|전체메뉴', t):
print(' a:', repr(t[:25]), '| href=', a.get('href'), '| onclick=', (a.get('onclick') or '')[:80])
# 부모/형제에 menuCd 있나
par = a.find_parent()
print(' parent menuCd:', re.findall(r'menuCd=DOM_\d+', str(par))[:3])
# 남원/완주: menuUid/contentUid 기반 — 사이트맵 링크 a 검색
for name, url, key in [('남원시','https://www.namwon.go.kr/index.do?menuUid=ff8080818e3beff0018e40e8f63e02d2','menuUid'),
('완주군','https://www.wanju.go.kr/index.9is','contentUid')]:
r = fetch(S, url)
soup = BeautifulSoup(r.text, 'html.parser')
print(f'\n=== {name} index — 사이트맵/전체메뉴 a ({key}) ===')
found = False
for a in soup.find_all('a'):
t = (a.get_text() or '').strip()
href = a.get('href') or ''
if re.search(r'사이트맵|누리집지도|전체메뉴|site_map|sitemap', t + href, re.I):
print(' a:', repr(t[:25]), '| href=', href[:100])
found = True
if not found:
print(' (없음) — 모든 a href에서 sitemap 토큰 검색:')
for a in soup.find_all('a', href=True):
if re.search(r'sitemap|site_map', a['href'], re.I):
print(' ', a['href'][:110], '|', (a.get_text() or '').strip()[:20])
# 군산: 사이트맵 페이지 추정 — /main 외 흔한 경로 시도
print('\n=== 군산 사이트맵 후보 ===')
for guess in ['https://www.gunsan.go.kr/sitemap','https://www.gunsan.go.kr/kor/sitemap.do',
'https://www.gunsan.go.kr/main?menuCd=','https://www.gunsan.go.kr/sitemap.do']:
try:
r = fetch(S, guess, t=10)
print(f' {guess} -> {r.status_code} {r.url}')
except Exception as e:
print(f' {guess} -> ERR {type(e).__name__}')
r = fetch(S, 'https://www.gunsan.go.kr/main')
soup = BeautifulSoup(r.text, 'html.parser')
print(' 군산 index 사이트맵류 a:')
for a in soup.find_all('a'):
t = (a.get_text() or '').strip()
href = a.get('href') or ''
if re.search(r'사이트맵|누리집지도|전체메뉴|site_map|sitemap', t + href, re.I):
print(' ', repr(t[:25]), '| href=', href[:100])
# 순창: 정적 사이트맵 페이지 탐색
print('\n=== 순창 사이트맵 후보 ===')
r = fetch(S, 'https://www.sunchang.go.kr/')
soup = BeautifulSoup(r.text, 'html.parser')
for a in soup.find_all('a'):
t = (a.get_text() or '').strip()
href = a.get('href') or ''
if re.search(r'사이트맵|누리집지도|전체메뉴|site_map|sitemap', t + href, re.I):
print(' a:', repr(t[:25]), '| href=', href[:110])

View File

@ -0,0 +1,85 @@
"""4차: 부안/완주/군산/순창 메뉴 소스 확정."""
import re, warnings
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
def sess():
s = requests.Session(); s.headers.update({'User-Agent': UA}); return s
def fetch(s, url, t=20):
r = s.get(url, timeout=t, verify=False, allow_redirects=True)
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
r.encoding = meta.group(1).decode('ascii','ignore') if meta else r.apparent_encoding
return r
S = sess()
def dump_menu_divs(soup, label):
print(f'\n--- {label}: class에 menu/gnb/sitemap/allmenu 포함 div·nav·ul (직계 a 텍스트) ---')
for el in soup.find_all(['div','nav','ul']):
cls = ' '.join(el.get('class', []))
idv = el.get('id','')
if re.search(r'menu|gnb|sitemap|site_map|allmenu|lnb|depth', cls + ' ' + idv, re.I):
n_a = len(el.find_all('a'))
n_li = len(el.find_all('li'))
if n_a >= 15: # 메가메뉴 후보
print(f' <{el.name} id="{idv}" class="{cls}"> a={n_a} li={n_li}')
# 부안: 사이트맵 menuCd 추정 — index의 모든 menuCd 중 *07*(정보공개류) 상위 노드 시도
print('=== 부안군 ===')
r = fetch(S, 'https://www.buan.go.kr/index.buan?contentsSid=1')
soup = BeautifulSoup(r.text, 'html.parser')
dump_menu_divs(soup, '부안 index')
# 누리집지도/사이트맵 후보 menuCd 시도
for cd in ['DOM_000000107003000000','DOM_000000108003000000','DOM_000000107002000000',
'DOM_000000106002000000','DOM_000000109002000000']:
try:
rr = fetch(S, f'https://www.buan.go.kr/index.buan?menuCd={cd}', t=10)
ss = BeautifulSoup(rr.text, 'html.parser')
has = bool(ss.select_one('div.sitemap'))
ttl = (ss.title.get_text().strip()[:30] if ss.title else '')
print(f' {cd}: div.sitemap={has} title={ttl!r}')
except Exception as e:
print(f' {cd}: ERR {type(e).__name__}')
# 완주: index 메가메뉴 + contentUid 사이트맵 후보
print('\n=== 완주군 ===')
r = fetch(S, 'https://www.wanju.go.kr/index.9is')
soup = BeautifulSoup(r.text, 'html.parser')
dump_menu_divs(soup, '완주 index')
for a in soup.find_all('a'):
t = (a.get_text() or '').strip()
if re.search(r'누리집|사이트맵|전체메뉴|지도', t):
print(' 완주 a:', repr(t[:25]), '| href=', (a.get('href') or '')[:100], '| onclick=', (a.get('onclick') or '')[:90])
# 군산: allmenubox 전체 구조(탭 포함)
print('\n=== 군산시 ===')
r = fetch(S, 'https://www.gunsan.go.kr/main')
soup = BeautifulSoup(r.text, 'html.parser')
dump_menu_divs(soup, '군산 index')
box = soup.select_one('div.allmenubox') or soup.select_one('div.allmw')
if box:
# 탭/카테고리 구조 — 직계 자식 요약
print(' allmenubox 하위 ul/div (a수):')
for el in box.find_all(['ul','div'], recursive=True):
cls = ' '.join(el.get('class', []))
n_a = len(el.find_all('a', recursive=False))
if n_a >= 5:
print(f' <{el.name} class="{cls}"> 직계a={n_a} 첫a={el.find("a").get_text().strip()[:15]!r}')
# 순창: HTML 내 sitemap/menu json/ajax 흔적
print('\n=== 순창군 ===')
r = fetch(S, 'https://www.sunchang.go.kr/')
html = r.text
soup = BeautifulSoup(html, 'html.parser')
dump_menu_divs(soup, '순창 index')
print(' html 내 sitemap/menu ajax url 흔적:')
for m in set(re.findall(r'["\']([^"\']*(?:sitemap|menu)[^"\']*\.(?:do|json|html|jsp))["\']', html, re.I)):
print(' ', m[:100])
# 순창 gnb_list가 ajax라면, 흔한 e-Gov 전체메뉴 경로 시도
for guess in ['/site/main/menu/sitemap','/kr/sitemap','/sitemap.do','/main/sitemap']:
try:
rr = fetch(S, urljoin('https://www.sunchang.go.kr', guess), t=10)
print(f' {guess} -> {rr.status_code}')
except Exception as e:
print(f' {guess} -> ERR')

View File

@ -0,0 +1,58 @@
"""5차: 부안/완주/군산/순창 인라인 메가메뉴 중첩 구조 정밀 덤프."""
import re, warnings
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
def fetch(url, t=20):
r = requests.get(url, timeout=t, verify=False, headers={'User-Agent': UA})
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
r.encoding = meta.group(1).decode('ascii','ignore') if meta else r.apparent_encoding
return r
def tree(el, depth=0, maxd=4, maxchild=6):
if el is None or depth > maxd: return
kids = el.find_all(recursive=False)
for i, c in enumerate(kids):
if i >= maxchild and depth >= 1:
print(' '*(depth+1) + '...'); break
cls = '.'.join(c.get('class', []))
a = c.find('a', recursive=False)
atxt = a.get_text().strip()[:18] if a else ''
ah = (a.get('href') or '')[:55] if a else ''
print(' '*(depth+1) + f'{c.name}.{cls}' + (f' a={atxt!r} {ah}' if atxt else ''))
tree(c, depth+1, maxd, maxchild)
print('############ 부안 nav#onmenu ############')
soup = BeautifulSoup(fetch('https://www.buan.go.kr/index.buan?contentsSid=1').text, 'html.parser')
nav = soup.select_one('nav#onmenu')
# 첫 1~2개 top li만
if nav:
top = nav.find('ul')
print('nav>ul 첫 li 2개:')
for li in (top.find_all('li', recursive=False)[:2] if top else []):
tree(li, 0, 4, 5)
print(' ----')
print('\n############ 완주 div.top_menu_wrap ############')
soup = BeautifulSoup(fetch('https://www.wanju.go.kr/index.9is').text, 'html.parser')
w = soup.select_one('div.top_menu_wrap')
tree(w, 0, 3, 4)
print('\n############ 군산 첫 allmenubox ############')
soup = BeautifulSoup(fetch('https://www.gunsan.go.kr/main').text, 'html.parser')
pc = soup.select_one('div#all_pcmenu')
if pc:
boxes = pc.select('div.allmenubox')
print(f'allmenubox 수: {len(boxes)}')
b = boxes[0]
print('Bmenu(대분류):', repr((b.find('a', class_='Bmenu') or b.find('a')).get_text().strip()[:20]))
tree(b, 0, 4, 5)
print('\n############ 순창 ul.gnb ############')
soup = BeautifulSoup(fetch('https://www.sunchang.go.kr/').text, 'html.parser')
g = soup.select_one('ul.gnb')
if g:
print(f'ul.gnb 직계 li: {len(g.find_all("li", recursive=False))}')
for li in g.find_all('li', recursive=False)[:1]:
tree(li, 0, 4, 5)

View File

@ -0,0 +1,46 @@
"""6차: 완주 메뉴/사이트맵 확정 + 부안 대분류 라벨 확인."""
import re, warnings
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
def fetch(url, t=20):
r = requests.get(url, timeout=t, verify=False, headers={'User-Agent': UA})
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
r.encoding = meta.group(1).decode('ascii','ignore') if meta else r.apparent_encoding
return r
# 부안 대분류 라벨: 각 depth_boxcon의 strong/p 텍스트
print('=== 부안 대분류(D) 라벨 ===')
soup = BeautifulSoup(fetch('https://www.buan.go.kr/index.buan?contentsSid=1').text, 'html.parser')
nav = soup.select_one('nav#onmenu')
top = nav.find('ul')
for li in top.find_all('li', recursive=False):
a0 = li.find('a', recursive=False)
box = li.find('div', class_='depth_boxcon')
strong = box.find('strong') if box else None
p = box.find('p') if box else None
print(' topA=', repr((a0.get_text().strip()[:20]) if a0 else ''),
'| strong=', repr(strong.get_text().strip()[:20] if strong else ''),
'| p=', repr(p.get_text(' ',strip=True)[:25] if p else ''))
# 완주: 전체 a 중 contentUid 사이트맵/메가메뉴 흔적. gnb mega menu 클래스 탐색
print('\n=== 완주 메뉴 구조 탐색 ===')
soup = BeautifulSoup(fetch('https://www.wanju.go.kr/index.9is').text, 'html.parser')
# 모든 div/ul 중 a>=30 이며 menuUid/contentUid href 다수인 것
for el in soup.find_all(['div','ul','nav']):
cls = ' '.join(el.get('class', [])); idv = el.get('id','')
n_a = len(el.find_all('a'))
if n_a >= 30:
sample = el.find('a', href=re.compile(r'contentUid|menuUid'))
print(f' <{el.name} id={idv!r} class={cls!r}> a={n_a}',
('| 샘플=' + (sample.get('href')[:60] if sample else 'none')))
# 사이트맵 페이지 후보: 전주처럼 index.9is?contentUid= 의 sitemap_Warp
# 완주 모든 contentUid 수집 후 'sitemap_Warp' 또는 'sitemap' 포함 페이지 찾기엔 비용 큼.
# 대신 footer 영역 a 전부 출력
foot = soup.find('footer') or soup.select_one('div.footer, #footer')
if foot:
print(' footer a:')
for a in foot.find_all('a')[:40]:
t=(a.get_text() or '').strip()
if t: print(' ', repr(t[:18]), (a.get('href') or '')[:60])

View File

@ -0,0 +1,172 @@
"""Probe sitemap pages directly using candidate URLs found in probe pass 1."""
import re
import warnings
from urllib.parse import urljoin, urlparse
from concurrent.futures import ThreadPoolExecutor, as_completed
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
# Round 1 results — best candidate sitemap URL per site
CANDIDATES = {
'논산시': 'https://nonsan.go.kr/kor/html/sub07/0701.html',
'당진시': 'https://www.dangjin.go.kr/kor/sitemap_11.do',
'보령시': 'https://www.brcn.go.kr/kor/sitemap_11.do',
'부여군': 'https://www.buyeo.go.kr/html/kr/sitemap.do',
'서산시': 'https://www.seosan.go.kr/www/sitemap.do',
'서천군': 'https://www.seocheon.go.kr/kor/sitemap_11.do',
'아산시': 'https://www.asan.go.kr/main/sitemap.do',
'예산군': 'https://www.yesan.go.kr/kor/sitemap.do',
'천안시': 'https://www.cheonan.go.kr/kor/sitemap.do',
'청양군': 'https://www.cheongyang.go.kr/kor/sitemap_11.do',
'태안군': 'https://www.taean.go.kr/kor/sitemap_11.do',
'홍성군': 'https://www.hongseong.go.kr/kor/sitemap.do',
}
# Alternative URLs to try if main candidate fails
ALTS = {
'부여군': ['https://www.buyeo.go.kr/html/kr/html/sub07/0701.html',
'https://www.buyeo.go.kr/html/kr/sitemap.html',
'https://www.buyeo.go.kr/html/kr/sitemap.do'],
'서산시': ['https://www.seosan.go.kr/www/sitemap.do',
'https://www.seosan.go.kr/www/contents.do?key=151'],
'아산시': ['https://www.asan.go.kr/main/sitemap.do',
'https://www.asan.go.kr/main/sub01_01.do'],
'예산군': ['https://www.yesan.go.kr/kor/sitemap.do',
'https://www.yesan.go.kr/kor/sitemap01.do'],
}
def fetch(url, timeout=15):
try:
r = requests.get(url, headers=H, timeout=timeout, verify=False, allow_redirects=True)
r.encoding = r.apparent_encoding
if r.status_code == 200:
return r.url, r.text
except Exception as e:
return None, f'ERR: {e}'
return None, f'HTTP {r.status_code}'
def deep_analyze(html):
"""Deeply analyze HTML to find sitemap container regardless of class name."""
soup = BeautifulSoup(html, 'html.parser')
# Try several common selectors
selectors = [
'.sitemap', '#sitemap', '.site_map', '.sitemap_wrap',
'.allMenu', '.allmenu', '#allMenu', '.all_menu', '.allMenuWrap',
'.contents .menu', '#content .sitemap', '.menu_all',
'.gnb_all', '.totalMenu', '.totMenu',
'div[class*=sitemap]', 'div[class*=allMenu]',
]
candidates = []
for sel in selectors:
for el in soup.select(sel):
# Count nested anchors as a measure of usefulness
a_count = len(el.find_all('a'))
if a_count >= 20:
candidates.append((a_count, sel, el))
if not candidates:
# Fallback — find any container with most anchors (excluding header/footer)
all_divs = soup.find_all(['div', 'section', 'main', 'article'])
for d in all_divs:
cls = ' '.join(d.get('class', []))
if 'header' in cls.lower() or 'footer' in cls.lower() or 'gnb' in cls.lower() and 'all' not in cls.lower():
continue
a_count = len(d.find_all('a'))
if a_count >= 50:
candidates.append((a_count, f'div.{cls}', d))
candidates.sort(reverse=True)
return candidates[:3]
def describe(el):
"""Describe DOM structure of an element."""
info = {
'tag': el.name,
'class': ' '.join(el.get('class', [])),
'id': el.get('id', ''),
'a_count': len(el.find_all('a')),
'dl_count': len(el.find_all('dl')),
'ul_count': len(el.find_all('ul')),
'li_count': len(el.find_all('li')),
}
# Identify pattern
dls_direct = el.find_all('dl', recursive=False)
uls_direct = el.find_all('ul', recursive=False)
info['dl_direct'] = len(dls_direct)
info['ul_direct'] = len(uls_direct)
if dls_direct and dls_direct[0].find('dt') and dls_direct[0].find('dd'):
info['pattern'] = 'A (dl>dt|dd>ul)'
elif uls_direct:
info['pattern'] = 'B (ul nested)'
else:
# Search one level deeper
nested_dl = []
nested_ul = []
for child in el.find_all(['div', 'section'], recursive=False):
nested_dl.extend(child.find_all('dl', recursive=False))
nested_ul.extend(child.find_all('ul', recursive=False))
if nested_dl:
info['pattern'] = f'A nested 1 deep (dl={len(nested_dl)})'
elif nested_ul:
info['pattern'] = f'B nested 1 deep (ul={len(nested_ul)})'
else:
info['pattern'] = 'UNKNOWN'
return info
def probe(name, url):
print(f'\n=== {name} === {url}')
u, html = fetch(url)
if not u:
print(f' 실패: {html}')
# Try alternatives
for alt in ALTS.get(name, []):
u, html = fetch(alt)
if u:
print(f' 대체 URL: {alt}')
break
else:
return name, None
print(f' 최종 URL: {u} ({len(html)} bytes)')
cands = deep_analyze(html)
if not cands:
print(' 사이트맵 컨테이너 못 찾음')
# Dump some snippets to help diagnose
soup = BeautifulSoup(html, 'html.parser')
for tag in ['title', 'h1', 'h2']:
for t in soup.find_all(tag)[:3]:
print(f' {tag}: {t.get_text(strip=True)[:80]}')
return name, None
for a_count, sel, el in cands:
info = describe(el)
print(f' [{a_count} anchors] selector="{sel}"{info}')
# Return best candidate info
best_count, best_sel, best_el = cands[0]
return name, {'url': u, 'selector': best_sel, **describe(best_el)}
def main():
results = {}
with ThreadPoolExecutor(max_workers=4) as ex:
futs = {ex.submit(probe, n, u): n for n, u in CANDIDATES.items()}
for f in as_completed(futs):
n, info = f.result()
results[n] = info
print('\n=== 최종 요약 ===')
for n in CANDIDATES:
info = results.get(n)
if info:
print(f' {n}: {info["url"]} | sel={info["selector"]} | a={info["a_count"]} dl={info["dl_count"]} ul={info["ul_count"]} | {info["pattern"]}')
else:
print(f' {n}: NOT FOUND')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,221 @@
# -*- coding: utf-8 -*-
"""계룡시 방식 일괄 적용 (수정판) — 충청남도(2~15) + 충청북도(1~11).
원본 phase234 검출 로직을 그대로 재사용한다(import):
- get_body(body_sel) '본문 영역' 한정해 검출 헤더/푸터 외부링크 노이즈 제거
- extract_detail_urls(JS fn_detail 폴백 포함) 게시판 상세를 원본과 동일하게 추적
- KOGL_IMG_PAT = img_opentype(\\d{2}).png (2자리 png), KOGL_LINK_PAT = licenseType(\\d)
부착 행을 재크롤링하여
1) O열(15) = 이미지(img_opentype) 유형만으로 재판정 (링크 숫자는 합치지 않음)
2) 링크가 '존재'하면서 이미지링크면 S열(19) 비고에 '링크주소 오기' (링크 없음은 비움)
파일별 *_backup_kogl재판정전.xlsx 백업 저장(기존 백업 있으면 보존).
사용: python -X utf8 _recheck_kogl_all.py [기관명 ...]
"""
import sys, os, re, shutil, warnings, openpyxl
from concurrent.futures import ThreadPoolExecutor
from urllib.parse import urlparse
sys.path.insert(0, r'D:\01.프로젝트\DB수집\_스크립트')
sys.stdout.reconfigure(encoding='utf-8')
warnings.filterwarnings('ignore')
import _chungnam_phase234_all as CN
import _chungbuk_phase234_all as CB
O_COL, K_COL, S_COL = 15, 11, 19
WORKERS = 6
# CN/CB 모듈 SITES에 없는 기관(개별 스크립트만 존재) 보완 — body_sel은 계룡시와 동일
EXTRA_CFG = {
'공주시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx',
'body_sel': ['#txt', '#contents', 'main']},
'금산군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\충청남도_금산군.xlsx',
'body_sel': ['#txt', '#contents', 'main']},
}
def get_cfg(name, region):
mod = CN if region == 'cn' else CB
return mod.SITES.get(name) or EXTRA_CFG[name]
# (기관명, region, weak_ssl) — xlsx/body_sel 은 각 모듈 SITES 에서 가져옴
TARGETS = [
('공주시', 'cn', False), ('금산군', 'cn', False), ('논산시', 'cn', False),
('당진시', 'cn', False), ('보령시', 'cn', False), ('부여군', 'cn', False),
('서산시', 'cn', False), ('서천군', 'cn', False), ('아산시', 'cn', False),
('예산군', 'cn', False), ('천안시', 'cn', False), ('청양군', 'cn', False),
('태안군', 'cn', False), ('홍성군', 'cn', False),
('괴산군', 'cb', False), ('단양군', 'cb', False), ('영동군', 'cb', True),
('옥천군', 'cb', False), ('음성군', 'cb', False), ('제천시', 'cb', False),
('증평군', 'cb', False), ('진천군', 'cb', False), ('청주시', 'cb', False),
('충주시', 'cb', False),
]
# 확장 이미지 패턴: img_opentype / img_opencode, 1~2자리, png/jpg/jpeg/gif
BROAD_IMG_PAT = re.compile(r'(?:new_)?img_open(?:type|code)(\d{1,2})\.(?:png|jpe?g|gif)', re.I)
def _valid(n):
return 1 <= n <= 4 # 공공누리 유형은 1~4만 유효
def detect_split(body, LINK_PAT):
"""이미지명(확장패턴) 유형과 링크(licenseType) 유형을 분리 검출. 유형 1~4만."""
img_t, link_t = set(), set()
for a in body.find_all('a', href=True):
m = LINK_PAT.search(a['href'])
if m and _valid(int(m.group(1))):
link_t.add(int(m.group(1)))
# 이미지: src 속성 + style 배경이미지 + raw HTML(누락 방지)
blob = ' '.join(filter(None, (img.get('src', '') for img in body.find_all('img'))))
blob += ' ' + ' '.join(el.get('style', '') for el in body.find_all(style=True))
blob += ' ' + str(body)
for m in BROAD_IMG_PAT.finditer(blob):
n = int(m.group(1))
if _valid(n):
img_t.add(n)
return img_t, link_t
def decide_O(img_t, link_t):
"""O열 값과 '링크주소 오기' 비고 여부 결정.
- 이미지 있으면 이미지 우선(권위). 링크 존재 & 이미지링크면 비고.
- 이미지 없고 링크만: 1·2·3·4 전부면 범례(설명)페이지 미부착. 아니면 링크 유형 인정.
- 없으면 미부착."""
if img_t:
newO = ','.join(f'{n}유형' for n in sorted(img_t))
return newO, bool(link_t) and (link_t != img_t)
if link_t:
if {1, 2, 3, 4}.issubset(link_t):
return '미부착', False
return ','.join(f'{n}유형' for n in sorted(link_t)), False
return '미부착', False
def make_fetch(region, weak_ssl):
if region == 'cn':
return CN.fetch # fetch(url, timeout)
sess = CB.make_session(weak_ssl)
return lambda url, timeout=12: CB.fetch(sess, url, timeout)
def _domain(host):
"""go.kr/or.kr 등 2단계 TLD 고려해 등록 도메인(끝 3라벨) 반환."""
labels = (host or '').split('.')
return '.'.join(labels[-3:]) if len(labels) >= 3 else host
def same_site(base_url, target_url):
"""상세 링크가 같은 기관 도메인일 때만 True (외부 사이트 KOGL 오탐 방지)."""
return _domain(urlparse(base_url).hostname) == _domain(urlparse(target_url).hostname)
def detect_row(url, body_sel, mod, fetch_fn):
soup = fetch_fn(url)
if soup is None:
return None, None, 'fetch_fail'
body = mod.get_body(soup, body_sel)
form, _ = mod.detect_form(body)
img_t, link_t = detect_split(body, mod.KOGL_LINK_PAT)
if form == '게시판':
for du in mod.extract_detail_urls(body, url, limit=5):
if not same_site(url, du): # 외부 도메인 상세링크 제외
continue
ds = fetch_fn(du, 10)
if ds is None:
continue
db = mod.get_body(ds, body_sel)
di, dl = detect_split(db, mod.KOGL_LINK_PAT)
img_t |= di
link_t |= dl
return img_t, link_t, 'ok'
def process_site(name, region, weak_ssl):
mod = CN if region == 'cn' else CB
cfg = get_cfg(name, region)
xlsx, body_sel = cfg['xlsx'], cfg['body_sel']
print(f'\n{"="*70}\n{name} [{region}] ({os.path.relpath(xlsx)})')
wb = openpyxl.load_workbook(xlsx)
ws = wb.active
targets = []
for r in range(3, ws.max_row + 1):
o = ws.cell(r, O_COL).value
if not o or str(o).strip() == '미부착':
continue
url = ws.cell(r, K_COL).value
if url:
targets.append((r, str(url).strip(), o, ws.cell(r, S_COL).value))
if not targets:
print(' 부착 행 없음 → 변경 없음')
return dict(name=name, total=0, o_changed=0, to_none=0, mismatch=0, fetch_fail=0)
print(f' 부착 행 {len(targets)}건 재크롤링 (본문영역+상세추적, workers={WORKERS}, weak_ssl={weak_ssl})')
fetch_fn = make_fetch(region, weak_ssl)
def work(t):
r, url, oldO, oldS = t
img_t, link_t, status = detect_row(url, body_sel, mod, fetch_fn)
return (r, url, oldO, oldS, img_t, link_t, status)
results = []
with ThreadPoolExecutor(max_workers=WORKERS) as ex:
for res in ex.map(work, targets):
results.append(res)
o_changed = to_none = mismatch = fetch_fail = 0
for r, url, oldO, oldS, img_t, link_t, status in sorted(results):
if status != 'ok':
fetch_fail += 1
print(f' row{r:>3} | FETCH_FAIL | {url[:70]}')
continue
newO, do_flag = decide_O(img_t, link_t)
if newO != str(oldO).strip():
o_changed += 1
if newO == '미부착':
to_none += 1
ws.cell(r, O_COL).value = newO
print(f' row{r:>3} | O: {str(oldO):>16} -> {newO:<16} <== 변경 | img={sorted(img_t)} link={sorted(link_t)}')
if do_flag:
mismatch += 1
cur = (oldS or '').strip()
if '링크주소 오기' not in cur:
ws.cell(r, S_COL).value = '링크주소 오기' if not cur else cur + ' / 링크주소 오기'
print(f' row{r:>3} | 비고+= 링크주소 오기 | img={sorted(img_t)} != link={sorted(link_t)}')
backup = os.path.join(os.path.dirname(xlsx), f'{name}_backup_kogl재판정전.xlsx')
if not os.path.exists(backup):
shutil.copy2(xlsx, backup)
wb.save(xlsx)
print(f'{name} 완료: O변경 {o_changed} (미부착化 {to_none}) | 불일치비고 {mismatch} | fetch실패 {fetch_fail} | 저장')
return dict(name=name, total=len(targets), o_changed=o_changed, to_none=to_none,
mismatch=mismatch, fetch_fail=fetch_fail)
def main():
sel = [a for a in sys.argv[1:] if not a.startswith('-')]
sites = [t for t in TARGETS if (not sel or t[0] in sel)]
print(f'대상 {len(sites)}개: {[s[0] for s in sites]}')
summary = []
for name, region, weak in sites:
try:
summary.append(process_site(name, region, weak))
except Exception as e:
import traceback; traceback.print_exc()
print(f' !! {name} 오류: {e}')
summary.append(dict(name=name, total=-1, o_changed=0, to_none=0, mismatch=0, fetch_fail=0))
print('\n' + '=' * 70 + '\n[전체 요약]')
print(f'{"기관":<8}{"부착":>6}{"O변경":>7}{"미부착化":>8}{"불일치":>7}{"실패":>6}')
for s in summary:
print(f'{s["name"]:<8}{s["total"]:>6}{s["o_changed"]:>7}{s["to_none"]:>8}{s["mismatch"]:>7}{s["fetch_fail"]:>6}')
tot = lambda k: sum(x[k] for x in summary if x[k] >= 0)
print(f'\n합계: 부착 {tot("total")} | O변경 {tot("o_changed")} | 미부착化 {tot("to_none")} | 불일치비고 {tot("mismatch")} | fetch실패 {tot("fetch_fail")}')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,178 @@
# -*- coding: utf-8 -*-
"""전북(14)+제주(2) 공공누리(O열) 권위 재판정 — 충청도 _recheck_kogl_all.py 방식.
_jeonbuk_phase234_all 검출 로직(get_body/detect_form/extract_detail_urls/KOGL_LINK_PAT)
그대로 재사용. 부착 행을 재크롤링하여:
1) O열 = 이미지명(img_opentype/opencode) 유형 우선 권위 재판정
2) 이미지 존재 & 이미지링크 S열 비고 '링크주소 오기'
3) 링크만 있고 1·2·3·4 전부 범례페이지로 미부착
규칙: feedback_kogl_image_rule. 상세추적은 같은 기관 도메인만.
사용: python -X utf8 _recheck_kogl_jeonbuk.py [기관명 ...]
"""
import sys, os, re, shutil, warnings, openpyxl
from concurrent.futures import ThreadPoolExecutor
from urllib.parse import urlparse
sys.path.insert(0, r'D:\01.프로젝트\DB수집\_스크립트')
sys.stdout.reconfigure(encoding='utf-8')
warnings.filterwarnings('ignore')
import _jeonbuk_phase234_all as JB
O_COL, K_COL, S_COL = 15, 11, 19
WORKERS = 12
DETAIL_TIMEOUT = 5
DETAIL_LIMIT = 3
# (기관명, weak_ssl) — xlsx/body_sel 은 JB.SITES 에서 가져옴
TARGETS = [(name, JB.SITES[name].get('weak_ssl', False)) for name in JB.SITES]
BROAD_IMG_PAT = re.compile(r'(?:new_)?img_open(?:type|code)(\d{1,2})\.(?:png|jpe?g|gif)', re.I)
def _valid(n):
return 1 <= n <= 4
def detect_split(body, LINK_PAT):
img_t, link_t = set(), set()
for a in body.find_all('a', href=True):
m = LINK_PAT.search(a['href'])
if m and _valid(int(m.group(1))):
link_t.add(int(m.group(1)))
blob = ' '.join(filter(None, (img.get('src', '') for img in body.find_all('img'))))
blob += ' ' + ' '.join(el.get('style', '') for el in body.find_all(style=True))
blob += ' ' + str(body)
for m in BROAD_IMG_PAT.finditer(blob):
n = int(m.group(1))
if _valid(n):
img_t.add(n)
return img_t, link_t
def decide_O(img_t, link_t):
if img_t:
newO = ','.join(f'{n}유형' for n in sorted(img_t))
return newO, bool(link_t) and (link_t != img_t)
if link_t:
if {1, 2, 3, 4}.issubset(link_t):
return '미부착', False
return ','.join(f'{n}유형' for n in sorted(link_t)), False
return '미부착', False
def make_fetch(weak_ssl):
sess = JB.make_session(weak_ssl)
return lambda url, timeout=DETAIL_TIMEOUT: JB.fetch(sess, url, timeout)
def _domain(host):
labels = (host or '').split('.')
return '.'.join(labels[-3:]) if len(labels) >= 3 else host
def same_site(base_url, target_url):
return _domain(urlparse(base_url).hostname) == _domain(urlparse(target_url).hostname)
def detect_row(url, body_sel, fetch_fn):
soup = fetch_fn(url)
if soup is None:
return None, None, 'fetch_fail'
body = JB.get_body(soup, body_sel)
form, _ = JB.detect_form(body)
img_t, link_t = detect_split(body, JB.KOGL_LINK_PAT)
if form == '게시판':
for du in JB.extract_detail_urls(body, url, limit=DETAIL_LIMIT):
if not same_site(url, du):
continue
ds = fetch_fn(du, DETAIL_TIMEOUT)
if ds is None:
continue
db = JB.get_body(ds, body_sel)
di, dl = detect_split(db, JB.KOGL_LINK_PAT)
img_t |= di
link_t |= dl
return img_t, link_t, 'ok'
def process_site(name, weak_ssl):
cfg = JB.SITES[name]
xlsx, body_sel = cfg['xlsx'], cfg['body_sel']
print(f'\n{"="*70}\n{name} ({os.path.relpath(xlsx)})')
wb = openpyxl.load_workbook(xlsx)
ws = wb.active
targets = []
for r in range(3, ws.max_row + 1):
o = ws.cell(r, O_COL).value
if not o or str(o).strip() == '미부착':
continue
url = ws.cell(r, K_COL).value
if url:
targets.append((r, str(url).strip(), o, ws.cell(r, S_COL).value))
if not targets:
print(' 부착 행 없음 → 변경 없음')
return dict(name=name, total=0, o_changed=0, to_none=0, mismatch=0, fetch_fail=0)
print(f' 부착 행 {len(targets)}건 재크롤링 (workers={WORKERS}, weak_ssl={weak_ssl})')
fetch_fn = make_fetch(weak_ssl)
def work(t):
r, url, oldO, oldS = t
img_t, link_t, status = detect_row(url, body_sel, fetch_fn)
return (r, url, oldO, oldS, img_t, link_t, status)
results = []
with ThreadPoolExecutor(max_workers=WORKERS) as ex:
for res in ex.map(work, targets):
results.append(res)
o_changed = to_none = mismatch = fetch_fail = 0
for r, url, oldO, oldS, img_t, link_t, status in sorted(results):
if status != 'ok':
fetch_fail += 1
continue
newO, do_flag = decide_O(img_t, link_t)
if newO != str(oldO).strip():
o_changed += 1
if newO == '미부착':
to_none += 1
ws.cell(r, O_COL).value = newO
if do_flag:
mismatch += 1
cur = (oldS or '').strip()
if '링크주소 오기' not in cur:
ws.cell(r, S_COL).value = '링크주소 오기' if not cur else cur + ' / 링크주소 오기'
backup = os.path.join(os.path.dirname(xlsx), f'{name}_backup_kogl재판정전.xlsx')
if not os.path.exists(backup):
shutil.copy2(xlsx, backup)
wb.save(xlsx)
print(f'{name} 완료: O변경 {o_changed} (미부착化 {to_none}) | 불일치비고 {mismatch} | fetch실패 {fetch_fail}')
return dict(name=name, total=len(targets), o_changed=o_changed, to_none=to_none,
mismatch=mismatch, fetch_fail=fetch_fail)
def main():
sel = [a for a in sys.argv[1:] if not a.startswith('-')]
sites = [t for t in TARGETS if (not sel or t[0] in sel)]
print(f'대상 {len(sites)}개: {[s[0] for s in sites]}')
summary = []
for name, weak in sites:
try:
summary.append(process_site(name, weak))
except Exception as e:
import traceback; traceback.print_exc()
summary.append(dict(name=name, total=-1, o_changed=0, to_none=0, mismatch=0, fetch_fail=0))
print('\n' + '=' * 70 + '\n[전체 요약]')
print(f'{"기관":<8}{"부착":>6}{"O변경":>7}{"미부착化":>8}{"불일치":>7}{"실패":>6}')
for s in summary:
print(f'{s["name"]:<8}{s["total"]:>6}{s["o_changed"]:>7}{s["to_none"]:>8}{s["mismatch"]:>7}{s["fetch_fail"]:>6}')
tot = lambda k: sum(x[k] for x in summary if x[k] >= 0)
print(f'\n합계: 부착 {tot("total")} | O변경 {tot("o_changed")} | 미부착化 {tot("to_none")} | 불일치비고 {tot("mismatch")} | fetch실패 {tot("fetch_fail")}')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,174 @@
시트 .do 페이지행: 407 — 재귀 크롤 시작...
크롤한 고유 페이지: 598
=== 수량 변경(증가/감소) 대상: 26행 ===
행306 주민복지 M 6 → 32 https://www.gongju.go.kr/kr/sub06_01_06_07_01.do
행398 교통약자 특별교통차량이용안내 M 1 → 18 https://www.gongju.go.kr/kr/sub06_09_03.do
행400 공주시 주정차 위반 단속 M 1 → 18 https://www.gongju.go.kr/kr/sub06_09_08.do
행401 주정차단속 문자알림 M 1 → 18 https://www.gongju.go.kr/kr/sub06_09_05.do
행402 공영주차장 현황 M 1 → 18 https://www.gongju.go.kr/kr/sub06_09_06.do
행403 자전거대여 M 1 → 18 https://www.gongju.go.kr/kr/sub06_09_07.do
행404 공주시 행복택시 M 1 → 18 https://www.gongju.go.kr/kr/sub06_09_09.do
행356 국민재난안전 포털 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_01.do
행357 충청남도 재난안전포털 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_05.do
행358 공주시 재난안전포털 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_09.do
행359 도민안전점검 청구제 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_02.do
행360 비닐하우스 피해경감 농가 행동요령 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_03.do
행361 민방위 정보 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_04.do
행363 시민안전보험 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_07.do
행364 어린이 안전보험 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_08.do
행365 풍수해보험 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_11.do
행367 반려동물을 위한 재난대처법 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_15.do
행110 제안하기 M 1 → 3 https://www.gongju.go.kr/kr/sub03_03_01.do
행111 나의제안 M 1 → 3 https://www.gongju.go.kr/kr/sub03_03_02.do
행112 공개제안 M 1 → 3 https://www.gongju.go.kr/kr/sub03_03_03.do
행334 공주시 행복누림 M 1 → 3 https://www.gongju.go.kr/kr/sub06_03_03.do
행335 공주시종합사회복지관 M 1 → 3 https://www.gongju.go.kr/kr/sub06_03_02.do
행336 충청남도 사이버교육 M 1 → 3 https://www.gongju.go.kr/kr/sub06_03_04.do
행167 미래전략실 M 1 → 2 https://www.gongju.go.kr/prog/deptPerson/kr/sub05_06_28_02/45003470000/list.do
행27 자동차등록안내 M 16 → 12 https://www.gongju.go.kr/kr/sub01_06_02_01.do
행396 버스정보 M 20 → 11 https://www.gongju.go.kr/kr/sub06_09_01_01.do
=== ⚠ 하위에 게시판 있어 합치지 않음(M 보류): 140행 ===
행19 납세자보호관 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub01_04_03_01.do
행20 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub01_04_03_02.do
행28 자동차검사 예약,신청조회 (현 M=3, 재귀=3) https://www.gongju.go.kr/kr/sub01_06_03_01.do
행33 정보공개제도안내 (현 M=6, 재귀=6) https://www.gongju.go.kr/kr/sub02_15_01_01.do
행35 정보공개목록 (현 M=1, 재귀=19) https://www.gongju.go.kr/kr/sub02_15_03.do
행36 정보공개목록(구) (현 M=1, 재귀=19) https://www.gongju.go.kr/kr/sub02_15_04.do
행38 정보공개청구 (현 M=1, 재귀=19) https://www.gongju.go.kr/kr/sub02_15_06.do
행68 공유재산 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub02_20_07_01.do
행70 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub02_20_07_03.do
행71 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub02_20_07_04.do
행75 공공데이터 개방목록 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub02_24_01.do
행77 축제 분석 인포그래픽 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub02_24_03.do
행81 기부제 안내 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub03_12_01.do
행82 답례품 안내 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub03_12_02.do
행84 홍보영상 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub03_12_04.do
행85 기부자 명예의 전당 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub03_12_05.do
행86 공모전 (현 M=1, 재귀=5) https://www.gongju.go.kr/prog/contest/kr/sub03_02_01/list.do
행90 공공언어 개선 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub03_02_11.do
행91 국민생각함 (현 M=1, 재귀=13) https://www.gongju.go.kr/prog/thinkBoxData/kr/sub03_03_04/getData.do
행92 공무원불친절 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_01.do
행93 민원부조리·부패신고 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_02.do
행94 공직자부조리신고센터 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_04.do
행95 예산낭비신고 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_05.do
행96 보조금부정수급신고 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_06.do
행98 안전신문고 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_09.do
행99 식품안전소비자신고 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_10.do
행100 부동산불법거래신고 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_11.do
행101 규제개혁 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_12.do
행102 직장 내 성희롱·성폭력·스토킹 고충상담 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_13.do
행103 공익신고 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_14.do
행109 법령유권해석 (현 M=2, 재귀=2) https://www.gongju.go.kr/kr/sub03_05_05_01.do
행116 여성인재DB 사업안내 (현 M=1, 재귀=1) https://www.gongju.go.kr/kr/sub03_10_01.do
행171 안전총괄과 (현 M=1, 재귀=2) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_06_05_02/45003590000/list.do
행172 (현 M=1, 재귀=2) https://www.gongju.go.kr/kr/sub05_06_05_01.do
행193 도시정책과 (현 M=1, 재귀=3) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_06_20_02/45003790000/list.do
행194 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub05_06_20_01.do
행196 허가건축과 (현 M=1, 재귀=3) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_06_21_02/45003800000/list.do
행197 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub05_06_21_01.do
행206 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_01_02.do
행207 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_01_03/yugu/dongList.do
행208 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_01_04/45000520000/list.do
행210 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_01_06.do
행212 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_02_02.do
행213 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_02_03/einmyun/dongList.do
행214 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_02_04/45000530000/list.do
행216 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_02_06.do
행218 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_03_02.do
행219 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_03_03/tancheon/dongList.do
행220 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_03_04/45000540000/list.do
행222 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_03_06.do
행224 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_04_02.do
행225 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_04_03/gyeryong/dongList.do
행226 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_04_04/45000550000/list.do
행228 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_04_06.do
행230 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_05_02.do
행231 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_05_03/banpo/dongList.do
행232 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_05_04/45000560000/list.do
행234 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_05_06.do
행236 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_06_02.do
행237 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_06_03/euidang/dongList.do
행238 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_06_04/45000580000/list.do
행240 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_06_06.do
행242 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_07_02.do
행243 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_07_03/jungan/dongList.do
행244 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_07_04/45000590000/list.do
행246 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_07_06.do
행248 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_08_02.do
행249 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_08_03/woosung/dongList.do
행250 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_08_04/45000600000/list.do
행252 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_08_06.do
행254 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_09_02.do
행255 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_09_03/sagok/dongList.do
행256 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_09_04/45000610000/list.do
행258 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_09_06.do
행260 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_10_02.do
행261 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_10_03/sinpoong/dongList.do
행262 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_10_04/45000620000/list.do
행264 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_10_06.do
행266 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_11_02.do
행267 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_11_03/junghak/dongList.do
행268 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_11_04/45000630000/list.do
행270 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_11_06.do
행272 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_12_02.do
행273 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_12_03/ungjin/dongList.do
행274 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_12_04/45000660000/list.do
행276 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_12_06.do
행278 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_13_02.do
행279 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_13_03/geumhak/dongList.do
행280 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_13_04/45000670000/list.do
행282 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_13_06.do
행284 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_14_02.do
행285 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_14_03/okryong/dongList.do
행286 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_14_04/45000680000/list.do
행288 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_14_06.do
행290 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_15_02.do
행291 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_15_03/sinkwan/dongList.do
행292 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_15_04/45000690000/list.do
행294 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_15_06.do
행296 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_16_02.do
행297 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_16_03/wolsong/dongList.do
행298 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_16_04/45002340000/list.do
행300 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_16_06.do
행307 노인 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_01_01_01.do
행308 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_01_01_02.do
행309 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_01_01_03.do
행310 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_01_01_08.do
행311 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_01_01_06.do
행313 장애인 (현 M=10, 재귀=81) https://www.gongju.go.kr/kr/sub06_01_02_01.do
행316 (현 M=1, 재귀=47) https://www.gongju.go.kr/kr/sub06_01_09_02_01.do
행317 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub06_01_09_03_01.do
행318 (현 M=1, 재귀=47) https://www.gongju.go.kr/kr/sub06_01_09_04_03.do
행319 (현 M=1, 재귀=46) https://www.gongju.go.kr/kr/sub06_01_09_05.do
행320 (현 M=1, 재귀=100) https://www.gongju.go.kr/kr/sub06_01_09_08.do
행328 공주시일자리센터 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_02_01.do
행330 임금체불 등 근로피해신고 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_02_04.do
행331 취업지원프로그램(고용24) (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_02_05.do
행332 충남인력개발원 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_02_06.do
행344 기업 (현 M=1, 재귀=1) https://www.gongju.go.kr/biz/index.do
행347 전통시장 (현 M=1, 재귀=28) https://www.gongju.go.kr/kr/sub06_06_01.do
행349 (현 M=1, 재귀=2) https://www.gongju.go.kr/kr/sub06_06_02_02.do
행352 유가정보서비스 (현 M=2, 재귀=2) https://www.gongju.go.kr/kr/sub06_06_05_01.do
행353 사회적·마을 기업 (현 M=4, 재귀=25) https://www.gongju.go.kr/kr/sub06_06_08.do
행368 지속가능발전협의회 (현 M=1, 재귀=328) https://www.gongju.go.kr/kr/sub06_08_01.do
행371 탄소중립포인트제 (현 M=1, 재귀=328) https://www.gongju.go.kr/kr/sub06_08_04.do
행375 야생동식물보호 (현 M=1, 재귀=328) https://www.gongju.go.kr/kr/sub06_08_08.do
행376 동물등록제 (현 M=1, 재귀=328) https://www.gongju.go.kr/kr/sub06_08_09.do
행377 상수도 (현 M=2, 재귀=5) https://www.gongju.go.kr/kr/sub06_08_12_01.do
행378 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_08_12_02.do
행379 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_08_12_03.do
행380 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_08_12_04.do
행382 하수도 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub06_08_13_01.do
행383 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub06_08_13_02_01.do
행385 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub06_08_13_04.do
행386 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub06_08_13_05.do
행388 수질검사 안내 (현 M=1, 재귀=4) https://www.gongju.go.kr/kr/sub06_08_14_01.do
행389 (현 M=1, 재귀=4) https://www.gongju.go.kr/kr/sub06_08_14_02.do
행391 (현 M=1, 재귀=4) https://www.gongju.go.kr/kr/sub06_08_14_04.do
행416 공주시 개인정보처리방침 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sitemap_03_01.do
행419 홈페이지 개인정보처리방침 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sitemap_03_03.do
행420 표준지방세정보시스템 개인정보 처리방침 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sitemap_03_04.do
(DRY — 적용하려면 --write)

View File

@ -0,0 +1,200 @@
# -*- coding: utf-8 -*-
"""수집완료 기관 N(저작물 유형) 일괄 재검토 (2026-06-01, 당진 규칙 일반화).
매뉴얼 3-1a/3-3/3-4b 반영:
- 이미지 = 별도 정보 주는 것만(사진·평면도·지도·QR·악보·소식지·상징물·도표).
- 어문 = 설명 인포그래픽(한눈에式)·로고·파트너로고·아이콘·배너·헤드라인텍스트·버튼·팝업·웹접근성마크.
- 오디오 = <audio>·mp3/wav/m4a 링크(다운로드 쿼리형 포함).
- 지도/PDF 임베드 = 이미지.
- 게시판(M=0) = 없음. 사이트(외부) = N 미변경(빈칸 유지).
N만 재계산하고 L/M/O/P/Q 다른 컬럼은 보존. 지역 phase234 모듈의 fetch/get_body/
extract_detail_urls/정규식/SITES 재사용. 검수완료(계룡·공주·금산·논산·보령)+당진 제외.
사용:
python -X utf8 _redo_N_all.py dry [기관...] # 변경 로그만(읽기전용)
python -X utf8 _redo_N_all.py run [기관...] # 백업 후 N 기입
(기관 미지정 3 지역 전체 기관)
"""
import sys, os, io, re, shutil, importlib.util, warnings
from concurrent.futures import ThreadPoolExecutor, as_completed
import openpyxl
warnings.filterwarnings('ignore')
HERE = os.path.dirname(os.path.abspath(__file__))
MODULES = ['_chungnam_phase234_all.py', '_chungbuk_phase234_all.py', '_jeonbuk_phase234_all.py']
EXCLUDE_INST = {'계룡시', '공주시', '금산군', '논산시', '보령시', '당진시'} # 검수완료 + 당진
# ── 이미지 분류 (벤더 공통) ───────────────────────────────
DECO = re.compile(r'/common/|move\.png|no[-_]?img|blank|spacer|/ico|/btn|bullet|arrow|/bg|icon|see_btn|/sample|mimetype|/file_|filedown|btn_dir', re.I)
EXCLUDE = re.compile(
r'한눈에|흐름도|절차도|처리절차|이용절차'
r'|로고(?!송)|logo(?!song)|아이콘|배너|banner'
r'|신문고|relation_item|tracer|headline'
r'|카피라이트|copyright|copy_logo|popup|/pup/|wa_mk|웹접근성|품질인증', re.I)
AUDIO = re.compile(r'\.(?:mp3|wav|m4a|ogg|flac)\b', re.I)
MAPPDF = re.compile(r'pdf|viewer\.html|/map|kakao|daum.*map', re.I)
def load_module(fname):
spec = importlib.util.spec_from_file_location(fname[:-3], os.path.join(HERE, fname))
m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m)
return m
def real_imgs(M, body):
out = []
for img in body.find_all('img'):
src = img.get('src') or ''
if not src or M.KOGL_IMG_PAT.search(src) or DECO.search(src):
continue
if EXCLUDE.search(src + ' ' + (img.get('alt') or '')):
continue
out.append(img)
return out
def media2(M, body):
has_text = len(body.get_text(strip=True)) > 30
img = len(real_imgs(M, body)) > 0
vid = False
for ifr in body.find_all('iframe'):
s = ifr.get('src') or ''
if M.YOUTUBE_PAT.search(s):
vid = True
elif MAPPDF.search(s):
img = True
if not vid and (body.find('video') or body.find('a', href=M.YOUTUBE_PAT) or M.VIDEO_EXT.search(str(body))):
vid = True
aud = bool(body.find('audio')) or bool(AUDIO.search(str(body)))
return img, vid, aud, has_text
def make_fetch(M, cfg):
"""모듈별 fetch 시그니처 차이 흡수: 충남=fetch(url), 충북/전북=fetch(session,url)."""
if hasattr(M, 'make_session'):
sess = M.make_session(weak_ssl=cfg.get('weak_ssl', False))
return lambda u: M.fetch(sess, u)
return lambda u: M.fetch(u)
def _fetch_body(M, url, body_sel, fetchfn, tries=3):
"""throttle 대비 재시도. 본문 텍스트>30 또는 미디어 잡히면 즉시 반환."""
import time
last = None
for k in range(tries):
soup = fetchfn(url)
if soup is not None:
body = M.get_body(soup, body_sel)
last = media2(M, body) + (body,)
if last[3] or last[0] or last[1] or last[2]: # txt/img/vid/aud 중 하나라도
return last
time.sleep(0.6 * (k + 1))
return last # 끝까지 비면 마지막(또는 None)
def n_of(M, url, L, body_sel, fetchfn):
res = _fetch_body(M, url, body_sel, fetchfn)
if res is None:
return None # 접근 실패 → 기존 N 유지
img, vid, aud, txt, body = res
if not (txt or img or vid or aud):
return None # 재시도해도 빈 본문 → 신뢰불가, 기존 N 유지(거짓 '없음' 차단)
if L == '게시판':
detail_urls = M.extract_detail_urls(body, url, limit=5)
data_rows = [tr for tr in body.select('table tbody tr, .board_list li, ul.bbs_list li') if tr.find('a')]
empty_msg = bool(re.search(r'게시물이?\s*없|등록된\s*(?:게시물|자료)\s*가?\s*없|자료가\s*없', body.get_text(' ', strip=True)))
if not detail_urls and not data_rows and (empty_msg or not txt):
return '없음' # 실제 글 0개 = 진짜 빈 게시판 (M값 무시, 내용기반)
for du in detail_urls:
ds = fetchfn(du)
if not ds:
continue
db = M.get_body(ds, body_sel)
i2, v2, a2, t2 = media2(M, db)
img = img or i2; vid = vid or v2; aud = aud or a2; txt = txt or t2
parts = []
if txt:
parts.append('어문')
if img:
parts.append('이미지')
if vid:
parts.append('영상')
if aud:
parts.append('오디오')
return ','.join(parts) if parts else '없음'
def do_inst(M, name, cfg, mode, log):
xlsx = cfg['xlsx']
if not os.path.exists(xlsx):
log.write(f'[{name}] 엑셀 없음\n'); return (name, 0, 0)
body_sel = cfg['body_sel']
fetchfn = make_fetch(M, cfg)
wb = openpyxl.load_workbook(xlsx); ws = wb.active
jobs = []
for r in range(3, ws.max_row + 1):
L = ws.cell(r, 12).value
K = ws.cell(r, 11).value
Mq = ws.cell(r, 13).value
if not (isinstance(K, str) and K.startswith('http')):
continue
if L == '사이트':
continue
if L not in ('페이지', '게시판'):
continue
jobs.append((r, K, L)) # M=0이어도 내용기반으로 n_of가 판정(거짓 M=0 대응)
results = {}
with ThreadPoolExecutor(max_workers=4) as ex:
futs = {}
for r, K, L in jobs:
if K == '__EMPTY__':
results[r] = '없음'
else:
futs[ex.submit(n_of, M, K, L, body_sel, fetchfn)] = r
for f in as_completed(futs):
r = futs[f]
try:
results[r] = f.result()
except Exception:
results[r] = None
changed = 0
for r, newN in sorted(results.items()):
if newN is None:
continue
oldN = ws.cell(r, 14).value
if newN != oldN:
changed += 1
log.write(f' [{name}] r{r} {oldN}{newN}\n')
if mode == 'run':
ws.cell(r, 14).value = newN
total = len([j for j in jobs])
if mode == 'run' and changed:
bak = xlsx.replace('.xlsx', '_backup_N재검토전.xlsx')
shutil.copy(xlsx, bak)
wb.save(xlsx)
return (name, total, changed)
def main():
mode = sys.argv[1] if len(sys.argv) > 1 else 'dry'
only = set(a for a in sys.argv[2:] if not a.startswith('--'))
log = io.open(os.path.join(HERE, '_redo_N_all_log.txt'), 'w', encoding='utf-8')
grand = []
for fname in MODULES:
M = load_module(fname)
for name, cfg in M.SITES.items():
if name in EXCLUDE_INST:
continue
if only and name not in only:
continue
res = do_inst(M, name, cfg, mode, log)
grand.append(res)
print(f'{name:7} 처리 {res[1]:4}행 | N변경 {res[2]:4} ({mode})', flush=True)
log.close()
print('=' * 50)
print(f'{len(grand)}개 기관 | 변경합계 {sum(r[2] for r in grand)}행 ({mode}) | 로그 _redo_N_all_log.txt')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,60 @@
[무주군] r3 어문,이미지→어문
[무주군] r4 어문,이미지→어문
[무주군] r6 어문,이미지→어문
[무주군] r12 어문,이미지→어문
[무주군] r13 어문,이미지→어문
[무주군] r14 어문,이미지→어문
[무주군] r15 어문,이미지→어문
[무주군] r17 어문,이미지→어문
[무주군] r26 어문,이미지→어문
[무주군] r27 어문,이미지→어문
[무주군] r30 어문,이미지→어문
[무주군] r89 어문,이미지→어문
[무주군] r90 어문,이미지→어문
[무주군] r91 어문,이미지→어문
[무주군] r92 어문,이미지→어문
[무주군] r99 어문,이미지→어문
[무주군] r103 어문,이미지→어문
[무주군] r116 어문,이미지→어문
[무주군] r120 어문,이미지→어문
[무주군] r121 어문,이미지→어문
[무주군] r126 어문,이미지→어문
[무주군] r134 어문,이미지→어문
[무주군] r155 어문,이미지→어문
[무주군] r156 어문,이미지→어문
[무주군] r157 어문,이미지→어문
[무주군] r158 어문,이미지→어문
[무주군] r159 어문,이미지→어문
[무주군] r163 어문,이미지→어문
[무주군] r164 어문,이미지→어문
[무주군] r165 어문,이미지→어문
[무주군] r185 어문,이미지→어문
[무주군] r186 어문,이미지→어문
[무주군] r187 어문,이미지→어문
[무주군] r188 어문,이미지→어문
[무주군] r189 어문,이미지→어문
[무주군] r190 어문,이미지→어문
[무주군] r192 어문,이미지→어문
[무주군] r193 어문,이미지→어문
[무주군] r194 어문,이미지→어문
[무주군] r195 어문,이미지→어문
[무주군] r196 어문,이미지→어문
[무주군] r197 어문,이미지→어문
[무주군] r198 어문,이미지→어문
[무주군] r199 어문,이미지→어문
[무주군] r200 어문,이미지→어문
[무주군] r207 어문,이미지→어문
[무주군] r239 어문,이미지,영상→어문,영상
[무주군] r247 어문,이미지→어문
[무주군] r262 어문,이미지→어문
[무주군] r263 어문,이미지→어문
[무주군] r314 어문,이미지→어문
[무주군] r323 어문,이미지→어문
[무주군] r326 어문,이미지→어문
[무주군] r349 어문,이미지→어문
[무주군] r350 어문,이미지→어문
[무주군] r360 어문,이미지→어문
[무주군] r362 어문,이미지→어문
[무주군] r368 어문,이미지→어문
[무주군] r370 어문,이미지→어문
[무주군] r379 어문,이미지→어문

View File

@ -0,0 +1,101 @@
# -*- coding: utf-8 -*-
"""작업파일명 앞에 광역(도) 접두 — {시}.xlsx → {도}_{시}.xlsx (2026-06-01).
1) 광역_사이트맵/{}/{N.}/{}.xlsx 메인 파일 rename (백업·부수파일 제외).
2) 모든 .py(_스크립트/ + 광역_사이트맵/**)에서 리터럴 '{시}.xlsx' '{도}_{시}.xlsx' 치환.
(충남·충북 phase234 SITES 하드코딩 경로, 기관폴더 스크립트 )
동적 빌더(f'{name}.xlsx' 5) 별도 수동 수정.
사용: python -X utf8 _rename_prefix.py dry|run
"""
import os, sys, io, re
ROOT = r'D:\01.프로젝트\DB수집'
MAP = os.path.join(ROOT, '작업파일', '광역_사이트맵')
PROVS = ['충청남도', '충청북도', '전북특별자치도', '제주특별자치도']
def build_map():
"""반환 [(도, 폴더명(N.시), 시, 기존xlsx경로, 새xlsx경로)]"""
out = []
for prov in PROVS:
pdir = os.path.join(MAP, prov)
if not os.path.isdir(pdir):
continue
for fold in sorted(os.listdir(pdir)):
fdir = os.path.join(pdir, fold)
if not os.path.isdir(fdir):
continue
si = fold.split('.', 1)[-1] # 'N.시' → 시
old = os.path.join(fdir, f'{si}.xlsx')
if os.path.exists(old):
new = os.path.join(fdir, f'{prov}_{si}.xlsx')
out.append((prov, fold, si, old, new))
return out
def py_files():
fs = []
sdir = os.path.join(ROOT, '_스크립트')
for f in os.listdir(sdir):
if f.endswith('.py'):
fs.append(os.path.join(sdir, f))
for dp, _, names in os.walk(MAP):
for n in names:
if n.endswith('.py'):
fs.append(os.path.join(dp, n))
return fs
def main():
mode = sys.argv[1] if len(sys.argv) > 1 else 'dry'
m = build_map()
# 시→도 (리터럴 치환용). 시명 길이 내림차순(부분일치 방지)
s2p = {si: prov for prov, fold, si, o, n in m}
print(f'== rename 대상 {len(m)}개 ==')
for prov, fold, si, o, n in m:
print(f' {prov}/{fold}/{si}.xlsx → {prov}_{si}.xlsx')
if mode == 'run':
if os.path.exists(n):
print(' !! 이미 존재, 스킵'); continue
os.rename(o, n)
# 리터럴 치환
print('\n== .py 리터럴 치환 ==')
keys = sorted(s2p.keys(), key=len, reverse=True)
total_files = 0; total_hits = 0
for fp in py_files():
try:
txt = io.open(fp, encoding='utf-8').read()
except Exception:
continue
orig = txt; hits = 0
for si in keys:
prov = s2p[si]
pat = si + '.xlsx'
rep = f'{prov}_{si}.xlsx'
# 이미 접두된 경우 중복 방지: {prov}_{si}.xlsx 는 건드리지 않음
# 음수 룩비하인드로 바로 앞이 '_'(도접두 직후)나 한글이면 스킵
def _sub(mo):
start = mo.start()
before = txt[max(0, start - 1):start]
# 직전이 '_' 또는 한글(다른 시명 꼬리)이면 치환 안 함
if before and (before == '_' or '' <= before <= ''):
return mo.group(0)
return rep
new = re.sub(re.escape(pat), _sub, txt)
if new != txt:
hits += new.count(rep) - orig.count(rep)
txt = new
if txt != orig:
total_files += 1
n_changed = sum(txt.count(f'{s2p[si]}_{si}.xlsx') for si in keys) - sum(orig.count(f'{s2p[si]}_{si}.xlsx') for si in keys)
print(f' {os.path.relpath(fp, ROOT)} (+{n_changed})')
total_hits += n_changed
if mode == 'run':
io.open(fp, 'w', encoding='utf-8').write(txt)
print(f'\n치환 파일 {total_files}개, 경로 {total_hits}건 ({mode})')
print('※ 동적 빌더 수동수정 필요: _dedup_all.py xpath, _tab_batch.py xlsx_path, _jeonbuk_phase234_all.py site(), phase1 3개')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,59 @@
# -*- coding: utf-8 -*-
"""조회 전용(저장 안 함): 부착 행 중 이미지(img_opentype) 없이
링크(licenseType) 있는 '텍스트 부착' 행의 URL을 ·군별로 나열."""
import sys, os, openpyxl
from concurrent.futures import ThreadPoolExecutor
sys.path.insert(0, r'D:\01.프로젝트\DB수집\_스크립트')
sys.stdout.reconfigure(encoding='utf-8')
import _recheck_kogl_all as RC
O_COL, K_COL = 15, 11
def scan(name, region, weak):
mod = RC.CN if region == 'cn' else RC.CB
cfg = RC.get_cfg(name, region)
xlsx, body_sel = cfg['xlsx'], cfg['body_sel']
wb = openpyxl.load_workbook(xlsx); ws = wb.active
targets = []
for r in range(3, ws.max_row + 1):
o = ws.cell(r, O_COL).value
if o and str(o).strip() != '미부착' and ws.cell(r, K_COL).value:
targets.append((r, str(ws.cell(r, K_COL).value).strip(), o))
if not targets:
return name, []
fetch_fn = RC.make_fetch(region, weak)
def work(t):
r, url, oldO = t
img_t, link_t, st = RC.detect_row(url, body_sel, mod, fetch_fn)
return (r, url, oldO, img_t, link_t, st)
out = []
with ThreadPoolExecutor(max_workers=6) as ex:
for r, url, oldO, img_t, link_t, st in ex.map(work, targets):
if st == 'ok' and not img_t and link_t:
out.append((r, url, str(oldO).strip(), sorted(link_t)))
return name, sorted(out)
def main():
sel = [a for a in sys.argv[1:] if not a.startswith('-')]
sites = [t for t in RC.TARGETS if (not sel or t[0] in sel)]
grand = 0
for name, region, weak in sites:
nm, rows = scan(name, region, weak)
if rows:
print(f'\n{nm} — 텍스트(링크)만 부착 {len(rows)}')
for r, url, oldO, lk in rows:
print(f' row{r:>3} | 기존O={oldO:<14} | 링크={lk} | {url}')
grand += len(rows)
else:
print(f'{nm} — 없음')
print(f'\n총 텍스트-only 부착 행: {grand}')
if __name__ == '__main__':
main()

File diff suppressed because it is too large Load Diff

174
_스크립트/_tab_batch.py Normal file
View File

@ -0,0 +1,174 @@
# -*- coding: utf-8 -*-
"""수집완료(✅) 기관 일괄 본문탭(1-5b) 확장 + 신규행 Phase2~4 수집 오케스트레이터.
매뉴얼 1-5b _tab_expand.py / _tab_phase234.py 기관 표대로 순차 호출한다.
- 검수완료(계룡시)·미처리(보은군)·이미 탭확장(공주시) 제외.
- base/domain 엑셀 K열 URL에서 자동 도출(가장 흔한 host 기준).
- 영동군만 가중 SSL(--weak-ssl).
모드:
python -X utf8 _tab_batch.py probe # 설정·행수·base/domain 검증만
python -X utf8 _tab_batch.py dry # 전 기관 확장계획(읽기전용)만 산출
python -X utf8 _tab_batch.py run [기관...] # 확장(--write)+Phase234 실제 수행
"""
import os
import re
import sys
import subprocess
from collections import Counter
from urllib.parse import urlsplit
import openpyxl
HERE = os.path.dirname(os.path.abspath(__file__))
ROOT = os.path.dirname(HERE)
MAP = os.path.join(ROOT, '작업파일', '광역_사이트맵')
CN_ALL = os.path.join(HERE, '_chungnam_phase234_all.py')
CB_ALL = os.path.join(HERE, '_chungbuk_phase234_all.py')
JB_ALL = os.path.join(HERE, '_jeonbuk_phase234_all.py')
GEUMSAN = os.path.join(MAP, '충청남도', '3.금산군', '_phase234.py')
JB_SEL = '#main-contents,#content,#contents,.contents,#txt,main,#container,#sub'
# (광역, idx, 기관, 폴더명, phase234모듈, body_sel(콤마/None), weak_ssl)
SITES = [
# 충청남도 (계룡=검수완료 제외, 공주=이미 탭확장 제외)
('충청남도', 3, '금산군', GEUMSAN, None, False),
('충청남도', 4, '논산시', CN_ALL, '#txt,#contents,main', False),
('충청남도', 5, '당진시', CN_ALL, '#txt,#contents,main', False),
('충청남도', 6, '보령시', CN_ALL, '#txt,#contents,main', False),
('충청남도', 7, '부여군', CN_ALL, '#txt,#contents,main', False),
('충청남도', 8, '서산시', CN_ALL, '#contents,#txt,main', False),
('충청남도', 9, '서천군', CN_ALL, '#txt,#contents,main', False),
('충청남도', 10, '아산시', CN_ALL, '#contents,main,#txt', False),
('충청남도', 11, '예산군', CN_ALL, '#txt,#contents,main', False),
('충청남도', 12, '천안시', CN_ALL, '#txt,#contents,main', False),
('충청남도', 13, '청양군', CN_ALL, '#txt,#contents,main', False),
('충청남도', 14, '태안군', CN_ALL, '#txt,#contents,main', False),
('충청남도', 15, '홍성군', CN_ALL, '#txt,#contents,main', False),
# 충청북도 (보은=미처리 제외)
('충청북도', 1, '괴산군', CB_ALL, '#contents,#txt,main', False),
('충청북도', 2, '단양군', CB_ALL, '#contents,#txt,main', False),
('충청북도', 4, '영동군', CB_ALL, '#txt,#contents,main', True),
('충청북도', 5, '옥천군', CB_ALL, '#contents,#txt,main', False),
('충청북도', 6, '음성군', CB_ALL, '#contents,#txt,main', False),
('충청북도', 7, '제천시', CB_ALL, '#contents,#txt,main', False),
('충청북도', 8, '증평군', CB_ALL, '#txt,#contents,main', False),
('충청북도', 9, '진천군', CB_ALL, '#contents,#txt,main', False),
('충청북도', 10, '청주시', CB_ALL, '#contents,#txt,main', False),
('충청북도', 11, '충주시', CB_ALL, '#contents,#txt,main', False),
# 전북특별자치도
('전북특별자치도', 1, '고창군', JB_ALL, JB_SEL, False),
('전북특별자치도', 2, '군산시', JB_ALL, JB_SEL, False),
('전북특별자치도', 3, '김제시', JB_ALL, JB_SEL, False),
('전북특별자치도', 4, '남원시', JB_ALL, JB_SEL, False),
('전북특별자치도', 5, '무주군', JB_ALL, JB_SEL, False),
('전북특별자치도', 6, '부안군', JB_ALL, JB_SEL, False),
('전북특별자치도', 7, '순창군', JB_ALL, JB_SEL, False),
('전북특별자치도', 8, '완주군', JB_ALL, JB_SEL, False),
('전북특별자치도', 9, '익산시', JB_ALL, JB_SEL, False),
('전북특별자치도', 10, '임실군', JB_ALL, JB_SEL, False),
('전북특별자치도', 11, '장수군', JB_ALL, JB_SEL, False),
('전북특별자치도', 12, '전주시', JB_ALL, JB_SEL, False),
('전북특별자치도', 13, '정읍시', JB_ALL, JB_SEL, False),
('전북특별자치도', 14, '진안군', JB_ALL, JB_SEL, False),
# 제주특별자치도
('제주특별자치도', 1, '서귀포시', JB_ALL, JB_SEL, False),
('제주특별자치도', 2, '제주시', JB_ALL, JB_SEL, False),
]
def xlsx_path(prov, idx, name):
return os.path.join(MAP, prov, f'{idx}.{name}', f'{prov}_{name}.xlsx')
def derive_base_domain(xlsx):
"""K열 URL에서 가장 흔한 host → base(scheme://host), domain(끝 3라벨)."""
wb = openpyxl.load_workbook(xlsx, read_only=True)
ws = wb.active
hosts = Counter()
scheme_by_host = {}
nrow = 0
for r in ws.iter_rows(min_row=3, min_col=11, max_col=11, values_only=True):
u = r[0]
if isinstance(u, str) and u.startswith('http'):
nrow += 1
sp = urlsplit(u)
if sp.hostname:
hosts[sp.hostname] += 1
scheme_by_host.setdefault(sp.hostname, sp.scheme)
# 데이터 행수(K 무관)
maxrow = ws.max_row
wb.close()
if not hosts:
return None, None, maxrow
host = hosts.most_common(1)[0][0]
base = f'{scheme_by_host[host]}://{host}'
labels = host.split('.')
domain = '.'.join(labels[-3:]) if len(labels) >= 3 else host
return base, domain, maxrow
def run(cmd):
print(' $', ' '.join(os.path.basename(c) if c.endswith('.py') else c for c in cmd[2:]))
p = subprocess.run(cmd, capture_output=True, text=True, encoding='utf-8', errors='replace')
out = (p.stdout or '') + (p.stderr or '')
return out
def main():
mode = sys.argv[1] if len(sys.argv) > 1 else 'probe'
only = set(sys.argv[2:]) if len(sys.argv) > 2 else None
py = [sys.executable, '-X', 'utf8']
sites = [s for s in SITES if not only or s[2] in only]
print(f'대상 기관: {len(sites)}개 (모드={mode})\n')
for prov, idx, name, mod, sel, weak in sites:
xp = xlsx_path(prov, idx, name)
if not os.path.exists(xp):
print(f'{prov} {name}: 엑셀 없음 → {xp}\n')
continue
base, domain, maxrow = derive_base_domain(xp)
tag = ' [weak-ssl]' if weak else ''
print(f'━━━ {prov} {name}{tag} (행~{maxrow}, base={base}, domain={domain})')
if mode == 'probe':
print(f' 모듈={os.path.basename(mod)} body_sel={sel}\n')
continue
# 1) 항상 dry 로 먼저 확장계획 산출(읽기전용)
ex = py + [os.path.join(HERE, '_tab_expand.py'), xp, base, domain]
if weak:
ex += ['--weak-ssl']
out = run(ex)
m = re.search(r'신규 추가 행:\s*(\d+)개', out)
newcnt = int(m.group(1)) if m else 0
grp = re.search(r'탭 확장 계획:\s*(\d+)개', out)
print(f' → 확장그룹 {grp.group(1) if grp else "?"}개 / 신규행 {newcnt}')
if mode == 'dry':
print()
continue
# 2) run 모드 & 신규행 있을 때만 --write 후 Phase234
if newcnt > 0:
exw = ex + ['--write']
run(exw)
ph = py + [os.path.join(HERE, '_tab_phase234.py'), xp, mod, domain]
if sel:
ph += [sel]
if weak:
ph += ['--weak-ssl']
out2 = run(ph)
for line in out2.splitlines():
if any(k in line for k in ('대상 신규행', 'L 분포', 'O 분포', '접근실패', '저장', '대체 저장')):
print(' ' + line)
else:
print(' (신규행 없음 → Phase234 생략)')
print()
print('\n=== 배치 종료 ===')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,81 @@
# -*- coding: utf-8 -*-
"""클래스명 무관 본문탭 탐지: 표본 페이지에서 '자기 URL을 포함하면서 2+실제링크 &
1+ 신규(엑셀에 없는) URL'을 가진 UL을 찾아 그 class 를 집계. 1-5b 조건2~4의 클래스 비의존 버전.
미지의 클래스명을 발견하기 위함.
사용: python -X utf8 _tab_classfind.py <엑셀> <base> <도메인> [표본=50] [--weak-ssl]
"""
import sys, warnings, ssl
from collections import Counter
from urllib.parse import urlsplit, urljoin
from concurrent.futures import ThreadPoolExecutor, as_completed
import openpyxl, requests
from requests.adapters import HTTPAdapter
from urllib3.util.ssl_ import create_urllib3_context
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
H={'User-Agent':'Mozilla/5.0 Chrome/120 Safari/537.36'}
class W(HTTPAdapter):
def init_poolmanager(self,*a,**k):
c=create_urllib3_context();c.set_ciphers('DEFAULT@SECLEVEL=0');c.options|=0x4
c.check_hostname=False;c.verify_mode=ssl.CERT_NONE;k['ssl_context']=c
return super().init_poolmanager(*a,**k)
S=requests.Session();S.headers.update(H)
if '--weak-ssl' in sys.argv: S.mount('https://',W())
def norm(u):
s=urlsplit(u);return urlsplit(urljoin('http://x/',s.path)).path.rstrip('/').lower()
def absu(h,base):
h=(h or '').strip()
if not h or h.startswith(('javascript:','#')): return ''
if h.startswith(('http://','https://')): return h
return urljoin(base.rstrip('/')+'/',h.lstrip('/'))
def main():
xlsx,base,domain=sys.argv[1],sys.argv[2],sys.argv[3]
nums=[a for a in sys.argv[4:] if a.isdigit()]
nsamp=int(nums[0]) if nums else 50
wb=openpyxl.load_workbook(xlsx,read_only=True);ws=wb.active
urls=[];existing=set()
for row in ws.iter_rows(min_row=3,min_col=11,max_col=11,values_only=True):
u=row[0]
if isinstance(u,str) and u.startswith('http'):
existing.add(norm(u))
if domain in u: urls.append(u)
wb.close()
step=max(1,len(urls)//nsamp);sample=urls[::step][:nsamp]
print(f'표본 {len(sample)}/{len(urls)} ({domain})')
cls_cnt=Counter();examples={}
def work(u):
try:
r=S.get(u,timeout=15,verify=False);return u,r.content
except Exception:return u,None
with ThreadPoolExecutor(max_workers=8) as ex:
for f in as_completed([ex.submit(work,u) for u in sample]):
u,html=f.result()
if not html:continue
cur=norm(u);soup=BeautifulSoup(html,'html.parser')
for ul in soup.find_all('ul'):
links=[]
ok=True
for a in ul.find_all('a'):
t=a.get_text(strip=True)
au=absu(a.get('href'),base)
if not t:continue
if not au: ok=False;break
links.append(au)
if not ok or len(links)<2:continue
norms=[norm(x) for x in links]
if cur not in norms:continue # 조건3: 자기 탭그룹
if sum(1 for n in norms if n not in existing)<1:continue # 조건4: 신규1+
cls=' '.join(ul.get('class') or []) or '(no-class)'
cls_cnt[cls]+=1
examples.setdefault(cls,(u,[ (a.get_text(strip=True), absu(a.get('href'),base)) for a in ul.find_all('a') if a.get_text(strip=True)][:6]))
print('\n=== 본문탭(자기URL포함+신규有) UL class 빈도 ===')
if not cls_cnt: print(' (없음 — 본문탭 진짜 없음 가능성)')
for cls,c in cls_cnt.most_common(15):
print(f' {c:3d}회 [{cls}]')
eu,el=examples[cls];print(f' 예: {eu}')
for t,h in el: print(f' - {t} -> {h}')
if __name__=='__main__':main()

View File

@ -0,0 +1,91 @@
# -*- coding: utf-8 -*-
"""미지 탭 클래스 발견: 표본 페이지 본문에서 '같은 본문에 2+ 실제링크를 가진 UL'
class 빈도순 집계한다. 사이트 전역 nav(거의 모든 페이지에 동일 링크집합) 제외 목적.
사용: python -X utf8 _tab_discover.py <엑셀> <도메인> [표본수=40]
"""
import sys, warnings, ssl
from collections import Counter
from urllib.parse import urlsplit
from concurrent.futures import ThreadPoolExecutor, as_completed
import openpyxl, requests
from requests.adapters import HTTPAdapter
from urllib3.util.ssl_ import create_urllib3_context
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
H = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/120 Safari/537.36'}
class WeakSSLAdapter(HTTPAdapter):
def init_poolmanager(self, *a, **k):
ctx = create_urllib3_context(); ctx.set_ciphers('DEFAULT@SECLEVEL=0')
ctx.options |= 0x4; ctx.check_hostname = False; ctx.verify_mode = ssl.CERT_NONE
k['ssl_context'] = ctx; return super().init_poolmanager(*a, **k)
S = requests.Session(); S.headers.update(H)
if '--weak-ssl' in sys.argv:
S.mount('https://', WeakSSLAdapter())
def norm(u):
s = urlsplit(u); return (s.path.rstrip('/')).lower()
def fetch(r, u):
try:
x = S.get(u, timeout=15, verify=False)
return r, u, x.content
except Exception:
return r, u, None
def main():
xlsx, domain = sys.argv[1], sys.argv[2]
nsamp = int([a for a in sys.argv[3:] if a.isdigit()][0]) if any(a.isdigit() for a in sys.argv[3:]) else 40
wb = openpyxl.load_workbook(xlsx, read_only=True); ws = wb.active
urls = []
for row in ws.iter_rows(min_row=3, min_col=11, max_col=11, values_only=True):
u = row[0]
if isinstance(u, str) and domain in u:
urls.append(u)
wb.close()
step = max(1, len(urls) // nsamp)
sample = urls[::step][:nsamp]
print(f'표본 {len(sample)}/{len(urls)} ({domain})')
# class별 등장 페이지 수 / 링크집합 다양성
cls_pages = Counter()
cls_linksets = {}
with ThreadPoolExecutor(max_workers=8) as ex:
futs = [ex.submit(fetch, i, u) for i, u in enumerate(sample)]
for f in as_completed(futs):
r, u, html = f.result()
if not html:
continue
soup = BeautifulSoup(html, 'html.parser')
for ul in soup.find_all(['ul', 'ol']):
links = []
for a in ul.find_all('a'):
t = a.get_text(strip=True)
h = (a.get('href') or '').strip()
if t and h and not h.startswith(('#', 'javascript:')):
links.append(norm(h))
if len(links) < 2:
continue
cls = ' '.join(ul.get('class') or []).strip() or '(no-class)'
cls_pages[cls] += 1
cls_linksets.setdefault(cls, set()).add(tuple(links))
print('\n=== UL/OL class별 (2+링크) — 등장페이지수 / 서로다른 링크집합수 ===')
print('(전역 nav = 등장多+집합1 / 본문탭 = 집합 다양) \n')
for cls, cnt in cls_pages.most_common(40):
nset = len(cls_linksets[cls])
flag = '★탭후보' if nset >= 3 and cnt >= 3 else ''
print(f' {cnt:3d}p / {nset:3d}집합 [{cls}] {flag}')
if __name__ == '__main__':
main()

107
_스크립트/_tab_dry.log Normal file
View File

@ -0,0 +1,107 @@
대상 기관: 39개 (모드=dry)
━━━ 충청남도 금산군 (행~526, base=https://www.geumsan.go.kr, domain=geumsan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\금산군.xlsx https://www.geumsan.go.kr geumsan.go.kr
→ 확장그룹 3개 / 신규행 9개
━━━ 충청남도 논산시 (행~738, base=https://nonsan.go.kr, domain=nonsan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\4.논산시\논산시.xlsx https://nonsan.go.kr nonsan.go.kr
→ 확장그룹 7개 / 신규행 22개
━━━ 충청남도 당진시 (행~315, base=https://www.dangjin.go.kr, domain=dangjin.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\5.당진시\당진시.xlsx https://www.dangjin.go.kr dangjin.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 충청남도 보령시 (행~573, base=https://www.brcn.go.kr, domain=brcn.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\6.보령시\보령시.xlsx https://www.brcn.go.kr brcn.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 충청남도 부여군 (행~315, base=https://www.buyeo.go.kr, domain=buyeo.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\7.부여군\부여군.xlsx https://www.buyeo.go.kr buyeo.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 충청남도 서산시 (행~349, base=https://www.seosan.go.kr, domain=seosan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\8.서산시\서산시.xlsx https://www.seosan.go.kr seosan.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 충청남도 서천군 (행~330, base=https://www.seocheon.go.kr, domain=seocheon.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\9.서천군\서천군.xlsx https://www.seocheon.go.kr seocheon.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 충청남도 아산시 (행~315, base=https://www.asan.go.kr, domain=asan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\10.아산시\아산시.xlsx https://www.asan.go.kr asan.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 충청남도 예산군 (행~387, base=https://www.yesan.go.kr, domain=yesan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\11.예산군\예산군.xlsx https://www.yesan.go.kr yesan.go.kr
→ 확장그룹 91개 / 신규행 233개
━━━ 충청남도 천안시 (행~371, base=https://www.cheonan.go.kr, domain=cheonan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\12.천안시\천안시.xlsx https://www.cheonan.go.kr cheonan.go.kr
→ 확장그룹 107개 / 신규행 297개
━━━ 충청남도 청양군 (행~315, base=http://www.cheongyang.go.kr, domain=cheongyang.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\13.청양군\청양군.xlsx http://www.cheongyang.go.kr cheongyang.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 충청남도 태안군 (행~315, base=https://www.taean.go.kr, domain=taean.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\14.태안군\태안군.xlsx https://www.taean.go.kr taean.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 충청남도 홍성군 (행~315, base=https://www.hongseong.go.kr, domain=hongseong.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\15.홍성군\홍성군.xlsx https://www.hongseong.go.kr hongseong.go.kr
→ 확장그룹 67개 / 신규행 215개
━━━ 충청북도 괴산군 (행~315, base=https://www.goesan.go.kr, domain=goesan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\1.괴산군\괴산군.xlsx https://www.goesan.go.kr goesan.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 충청북도 단양군 (행~486, base=https://www.danyang.go.kr, domain=danyang.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\2.단양군\단양군.xlsx https://www.danyang.go.kr danyang.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 충청북도 영동군 [weak-ssl] (행~674, base=https://www.yd21.go.kr, domain=yd21.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\4.영동군\영동군.xlsx https://www.yd21.go.kr yd21.go.kr --weak-ssl
→ 확장그룹 20개 / 신규행 143개
━━━ 충청북도 옥천군 (행~564, base=https://www.oc.go.kr, domain=oc.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\5.옥천군\옥천군.xlsx https://www.oc.go.kr oc.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 충청북도 음성군 (행~652, base=https://www.eumseong.go.kr, domain=eumseong.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\6.음성군\음성군.xlsx https://www.eumseong.go.kr eumseong.go.kr
→ 확장그룹 9개 / 신규행 18개
━━━ 충청북도 제천시 (행~644, base=https://www.jecheon.go.kr, domain=jecheon.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\7.제천시\제천시.xlsx https://www.jecheon.go.kr jecheon.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 충청북도 증평군 (행~315, base=https://www.jp.go.kr, domain=jp.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\8.증평군\증평군.xlsx https://www.jp.go.kr jp.go.kr
→ 확장그룹 42개 / 신규행 175개
━━━ 충청북도 진천군 (행~334, base=https://www.jincheon.go.kr, domain=jincheon.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\9.진천군\진천군.xlsx https://www.jincheon.go.kr jincheon.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 충청북도 청주시 (행~342, base=https://www.cheongju.go.kr, domain=cheongju.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\10.청주시\청주시.xlsx https://www.cheongju.go.kr cheongju.go.kr
→ 확장그룹 5개 / 신규행 97개
━━━ 충청북도 충주시 (행~403, base=https://www.chungju.go.kr, domain=chungju.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\11.충주시\충주시.xlsx https://www.chungju.go.kr chungju.go.kr
→ 확장그룹 3개 / 신규행 33개
━━━ 전북특별자치도 고창군 (행~381, base=https://www.gochang.go.kr, domain=gochang.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\1.고창군\고창군.xlsx https://www.gochang.go.kr gochang.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 전북특별자치도 군산시 (행~639, base=https://www.gunsan.go.kr, domain=gunsan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\2.군산시\군산시.xlsx https://www.gunsan.go.kr gunsan.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 전북특별자치도 김제시 (행~454, base=https://www.gimje.go.kr, domain=gimje.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\3.김제시\김제시.xlsx https://www.gimje.go.kr gimje.go.kr
→ 확장그룹 0개 / 신규행 0개
━━━ 전북특별자치도 남원시 (행~333, base=https://www.namwon.go.kr, domain=namwon.go.kr)

View File

@ -0,0 +1,433 @@
"""본문 탭(서브내비) → 카테고리 하위 확장.
규칙(공주시 기준, 일반화):
본문에서 '탭 UL' 찾아, 아래 조건을 모두 만족하는 그룹만 하위 카테고리로 확장한다.
1) UL(또는 직계 div) class 패턴(tab-ul ) 있다.
2) 탭이 2 이상이고, 모든 href '실제 페이지 링크'.
(#anchor 인페이지 탭, javascript:, 빈 href, 쿼리스트링만 다른 필터 탭은 제외)
3) 현재 페이지 URL URL 집합에 포함된다(= 자기 자신의 그룹).
4) '엑셀에 아직 없는 URL' 1 이상 있다(이미 사이트맵에 있으면 상위 nav 제외).
확장 방식(=사용자 지시: 탭들은 단계 아래 컬럼으로):
- 원래 행의 leaf 컬럼(E~J 가장 깊은 ) 부모(카테고리) 두고,
탭들을 leaf+1 컬럼에 순서대로 채운다.
- 현재 페이지와 같은 URL의 = 원래 행을 재사용(L~T 기존 데이터 보존, G만 탭명 기입).
- 나머지 = 신규 (L~T 비움, Phase 2~4 별도 수행 대상).
사용법:
python -X utf8 _tab_expand.py <엑셀경로> <BASE_URL> <도메인키워드> [--write]
) python -X utf8 _tab_expand.py 충청남도/2.공주시/충청남도_공주시.xlsx https://www.gongju.go.kr gongju.go.kr
(--write 없으면 계획만 출력 / 있으면 백업 실제 기입)
"""
import re
import sys
import shutil
import warnings
from copy import copy
from concurrent.futures import ThreadPoolExecutor, as_completed
from urllib.parse import urljoin, urlsplit, urlunsplit
import ssl
import openpyxl
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.ssl_ import create_urllib3_context
from bs4 import BeautifulSoup
from openpyxl.styles import Alignment, Font
from openpyxl.utils import get_column_letter
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
class WeakSSLAdapter(HTTPAdapter):
def init_poolmanager(self, *args, **kwargs):
ctx = create_urllib3_context()
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
ctx.options |= 0x4
ctx.check_hostname = False
ctx.verify_mode = ssl.CERT_NONE
kwargs['ssl_context'] = ctx
return super().init_poolmanager(*args, **kwargs)
SESSION = requests.Session()
SESSION.headers.update(H)
TAB_CLASS_PATS = ['tab-ul', 'tab_ul', 'tabmenu', 'tab-menu', 'tab_menu',
'tablist', 'tab-list', 'tab_list', 'subtab', 'sub-tab', 'sub_tab']
# 토큰 단위 탭 클래스 정규식: 'basic_tab', 'tab_wrap', 'tab-ul', 단독 'tab' 등 매칭.
# (UL 자신뿐 아니라 직계 부모 div 클래스도 검사 → div.basic_tab > ul 구조 대응)
TAB_CLASS_RE = re.compile(
r'(?:^|[-_ ])tab(?:[-_ ]|$)|tabmenu|tablist|tab[-_]?(?:ul|wrap|list|menu)|basic[-_]tab')
MAXCOL = 27 # AA
CAT_COLS = [5, 6, 7, 8, 9, 10] # E F G H I J
HEADER_MERGES = {'B1:R1', 'S1:W1', 'Y1:AA1'}
def norm_url(u):
"""비교용 키: 프래그먼트 제거 + 끝 슬래시 정리 + 쿼리 정렬 보존.
(쿼리만 다른 게시판 분류 ?code=A vs ?code=B 서로 다른 URL로 구분
is_cur·existing 중복판정 오류 방지)"""
if not u:
return ''
s = urlsplit(u)
path = s.path.rstrip('/')
q = '&'.join(sorted(s.query.split('&'))) if s.query else ''
return urlunsplit((s.scheme, s.netloc, path, q, '')).lower()
def path_key(u):
"""쿼리 제거한 경로 키 (필터 탭 판별용)."""
if not u:
return ''
s = urlsplit(u)
return urlunsplit((s.scheme, s.netloc, s.path.rstrip('/'), '', '')).lower()
def abs_url(href, base):
if not href:
return ''
href = href.strip()
if href.startswith(('javascript:', '#')):
return ''
if href.startswith(('http://', 'https://')):
return href
return urljoin(base.rstrip('/') + '/', href.lstrip('/'))
def has_tab_class(el):
if el is None or not getattr(el, 'get', None):
return False, ''
cls = ' '.join(el.get('class') or []).lower()
ok = any(p in cls for p in TAB_CLASS_PATS) or bool(TAB_CLASS_RE.search(cls))
return ok, cls
def _has_on_li(ul):
"""ul 직계 li(또는 그 a)에 활성 탭 마커(on/active/current/selected)가 있나."""
marks = {'on', 'active', 'current', 'selected', 'sel'}
for li in ul.find_all('li', recursive=False):
if marks & set(c.lower() for c in (li.get('class') or [])):
return True
a = li.find('a')
if a and (marks & set(c.lower() for c in (a.get('class') or []))):
return True
return False
def find_tab_groups(soup, base):
"""조건 1·2 만족하는 탭 그룹 후보 반환: [(cls, [(label, abs_url),...], has_on), ...]"""
out = []
seen = set()
for ul in soup.find_all('ul'):
ok, cls = has_tab_class(ul)
if not ok:
# UL 무클래스라도 직계 부모 div 에 탭 클래스가 있으면 인정 (div.basic_tab > ul)
pok, pcls = has_tab_class(ul.parent)
if not pok:
continue # 쿼리필터·무클래스 ul 배제
cls = pcls
links = []
real = True
for a in ul.find_all('a'):
t = a.get_text(strip=True).replace('\xa0', '').strip()
raw = (a.get('href') or '').strip()
au = abs_url(raw, base)
if not t:
continue
if not au: # #anchor / javascript / 빈 href → 인페이지 탭
real = False
break
links.append((t, au))
if not real or len(links) < 2:
continue
key = tuple(nu for _, (nu) in [(t, u) for t, u in links])
if key in seen:
continue
seen.add(key)
out.append((cls, links, _has_on_li(ul)))
return out
def fetch(r, url):
try:
resp = SESSION.get(url, timeout=15, verify=False)
# 바이트로 받아 BeautifulSoup가 meta charset으로 인코딩 판별 (apparent_encoding 오판 방지)
return r, url, resp.content, None
except Exception as e:
return r, url, None, str(e)[:60]
def load_flat(ws):
"""병합값을 채워넣어 행별 완전 데이터로 평탄화. 반환: rows(list of dict), 스타일/하이퍼링크 캡처."""
# 1) 병합값 채우기 (헤더 제외, 데이터 컬럼)
for mr in list(ws.merged_cells.ranges):
s = str(mr)
if s in HEADER_MERGES:
continue
top = ws.cell(mr.min_row, mr.min_col).value
ws.unmerge_cells(s)
for rr in range(mr.min_row, mr.max_row + 1):
for cc in range(mr.min_col, mr.max_col + 1):
if ws.cell(rr, cc).value in (None, ''):
ws.cell(rr, cc).value = top
rows = []
for r in range(3, ws.max_row + 1):
# 빈 행 스킵
if all(ws.cell(r, c).value in (None, '') for c in range(2, MAXCOL + 1)):
continue
vals = {c: ws.cell(r, c).value for c in range(1, MAXCOL + 1)}
styles = {}
for c in range(1, MAXCOL + 1):
sc = ws.cell(r, c)
styles[c] = (copy(sc.font), copy(sc.fill), copy(sc.border),
copy(sc.alignment), sc.number_format, copy(sc.protection))
hl = ws.cell(r, 11).hyperlink
rows.append({'src': r, 'vals': vals, 'styles': styles,
'hyperlink': hl.target if hl else None})
return rows
def leaf_col(vals):
"""E~J 중 값이 있는 가장 깊은 컬럼 인덱스."""
deep = 5
for c in CAT_COLS:
if vals.get(c) not in (None, ''):
deep = c
return deep
def main():
xlsx, base, domain = sys.argv[1], sys.argv[2], sys.argv[3]
write = '--write' in sys.argv
# menuCd 벤더(고창·임실·정읍·진안 등 index.{name}?menuCd=…): 모든 페이지가 같은 경로,
# 쿼리(menuCd)만 다름 = 다른 페이지. → 쿼리-경로 가드(조건2) 스킵 + cur∈탭(조건3) 대신
# div.basic_tab 등에 활성탭(li.on) 마커가 있는 '자기 sub-nav'만 인정(랜딩 URL이 탭과 달라도).
menucd = '--menucd' in sys.argv
if '--weak-ssl' in sys.argv:
SESSION.mount('https://', WeakSSLAdapter())
wb = openpyxl.load_workbook(xlsx)
ws = wb.active
rows = load_flat(ws)
existing = set()
for row in rows:
u = row['vals'].get(11)
if isinstance(u, str) and u.startswith('http'):
existing.add(norm_url(u))
# 같은 도메인 행 fetch
targets = [(row['src'], row['vals'].get(11)) for row in rows
if isinstance(row['vals'].get(11), str) and domain in row['vals'].get(11)]
html_by_src = {}
with ThreadPoolExecutor(max_workers=8) as ex:
futs = [ex.submit(fetch, r, u) for r, u in targets]
for fut in as_completed(futs):
r, url, html, err = fut.result()
if html:
html_by_src[r] = (url, html)
# 이미 사이트맵에 자식행이 있는 '랜딩행' 집합 (다음 평탄행이 같은 상위카테고리 + 더 깊은 leaf).
# menuCd 벤더 랜딩은 첫 자식의 탭을 렌더하므로, 확장하면 그 자식행과 중복 → 제외.
landing_src = set()
for i in range(len(rows) - 1):
v, nv = rows[i]['vals'], rows[i + 1]['vals']
lc = leaf_col(v)
if lc < 10 and nv.get(lc + 1) not in (None, '') \
and all((v.get(c) or '') == (nv.get(c) or '') for c in CAT_COLS if c <= lc):
landing_src.add(rows[i]['src'])
# 행별 확장 계획
plan = {} # src_row -> ordered [(label, url, is_existing)]
planned_new = set() # 이미 어느 부모행이 추가한 신규 탭 URL(전역 중복 방지)
for row in rows:
src = row['src']
if src not in html_by_src:
continue
if menucd and src in landing_src:
continue # 자식 보유 랜딩 → 확장 금지(첫 자식 탭 중복 방지)
url, html = html_by_src[src]
cur = norm_url(url)
soup = BeautifulSoup(html, 'html.parser')
groups = find_tab_groups(soup, base)
chosen = None
for cls, links, has_on in groups:
if menucd:
# menuCd 벤더: 현재 페이지와 '같은 경로'(menuCd만 다른 형제)인 탭만 인정.
# → /gochang/toc/GC…(향토문화대전 백과 목록) 등 다른경로 링크 배제.
# 활성탭(li.on) 마커가 있는 자기 sub-nav만(랜딩 URL이 탭에 없어도 OK).
if not has_on:
continue
links = [(t, u) for t, u in links if path_key(u) == path_key(cur)]
if len(links) < 2:
continue
tab_norms = [norm_url(u) for _, u in links]
else:
tab_norms = [norm_url(u) for _, u in links]
if len({path_key(u) for _, u in links}) == 1:
continue # 조건2: 같은 경로(쿼리만 다른 필터 탭) → 제외
if cur not in tab_norms:
continue # 조건3: 자기 탭그룹만
new_cnt = sum(1 for n in tab_norms if n not in existing)
if new_cnt < 1:
continue # 조건4: 신규 0 → 상위nav, 제외
chosen = links
break
if chosen:
ch = []
for label, u in chosen:
is_cur = norm_url(u) == cur
nu = norm_url(u)
if not is_cur and nu in existing:
continue # 이미 사이트맵에 별도 행으로 존재 → 중복행 방지(skip)
if not is_cur and nu in planned_new:
continue # 다른 부모행이 이미 추가한 신규 URL → 전역 중복 방지
if not is_cur:
planned_new.add(nu)
is_ext = domain not in u # 같은 도메인 아님 = 외부 사이트 탭
ch.append((label, u, is_cur, is_ext))
# 신규(비-cur) 탭이 1개 이상일 때만 확장. 아니면 원본행 그대로 둠(삭제 방지).
if any(not c[2] for c in ch):
plan[src] = ch
# 계획 출력
print(f'=== 탭 확장 계획: {len(plan)}개 그룹 ===\n')
total_new = 0
for src in sorted(plan):
row = next(r for r in rows if r['src'] == src)
v = row['vals']
lc = leaf_col(v)
childc = lc + 1
ev = v.get(5) or ''
fv = v.get(6) or ''
print(f'[행{src}] E={ev} F={fv} '
f'(leaf={get_column_letter(lc)} → 자식={get_column_letter(childc)})')
for label, u, is_cur, is_ext in plan[src]:
if is_cur:
tag = '재사용'
elif is_ext:
tag = '신규+(외부=사이트)'
total_new += 1
else:
tag = '신규+'
total_new += 1
print(f' [{tag}] {label} -> {u}')
print()
print(f'신규 추가 행: {total_new}개 / 기존 {len(rows)}행 → 최종 {len(rows)+total_new}')
if not write:
print('\n(계획만 출력. 실제 기입하려면 --write)')
return
# ===== 실제 기입 =====
backup = xlsx.replace('.xlsx', '_backup_tab전.xlsx')
shutil.copy(xlsx, backup)
print(f'\n백업: {backup}')
# 새 평탄 행 목록 구성
out_rows = [] # 각 원소 = {'vals':..., 'style_src':row_dict, 'url':..}
for row in rows:
src = row['src']
if src in plan:
v = row['vals']
lc = leaf_col(v)
childc = lc + 1
for label, u, is_cur, is_ext in plan[src]:
nv = dict(v)
# leaf 보다 깊은 컬럼 초기화 후 자식 컬럼에 탭명
for c in CAT_COLS:
if c > lc:
nv[c] = None
nv[childc] = label
if is_cur:
nv[11] = v.get(11) # 기존 URL/데이터 유지
out_rows.append({'vals': nv, 'style': row, 'url': nv.get(11)})
else:
for c in range(12, MAXCOL + 1):
nv[c] = None # L~T 비움 (신규)
nv[11] = u
nv[19] = None
if is_ext: # 외부 사이트 탭 → 즉시 L=사이트/M=1(Phase234 대상 아님)
nv[12] = '사이트'
nv[13] = 1
out_rows.append({'vals': nv, 'style': row, 'url': u})
else:
out_rows.append({'vals': dict(row['vals']), 'style': row, 'url': row['vals'].get(11)})
# 데이터 영역 클리어
for r in range(3, ws.max_row + 1):
for c in range(1, MAXCOL + 1):
ws.cell(r, c).value = None
ws.cell(r, c).hyperlink = None
START = 3
for i, orow in enumerate(out_rows):
r = START + i
sty = orow['style']['styles']
for c in range(1, MAXCOL + 1):
cell = ws.cell(r, c)
f, fl, bd, al, nf, pr = sty[c]
cell.font = copy(f); cell.fill = copy(fl); cell.border = copy(bd)
cell.alignment = copy(al); cell.number_format = nf; cell.protection = copy(pr)
v = orow['vals']
for c in range(1, MAXCOL + 1):
ws.cell(r, c).value = v.get(c)
ws.cell(r, 2).value = i + 1 # B 순번 재부여
END = START + len(out_rows) - 1
# D/E/F 재병합 (G는 leaf라 병합 안 함)
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
def merge_runs(col_letter, col_idx, group_cols=()):
cur_val = ws.cell(START, col_idx).value
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
run_start = START
runs = []
for r in range(START + 1, END + 1):
val = ws.cell(r, col_idx).value
grp = tuple(ws.cell(r, gg).value for gg in group_cols)
if val == cur_val and grp == cur_grp:
continue
if cur_val not in (None, '') and r - 1 > run_start:
runs.append((run_start, r - 1))
cur_val, cur_grp, run_start = val, grp, r
if cur_val not in (None, '') and END > run_start:
runs.append((run_start, END))
for s, e in runs:
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
ws.cell(s, col_idx).alignment = center
merge_runs('F', 6, group_cols=(4, 5))
merge_runs('E', 5, group_cols=(4,))
merge_runs('D', 4)
for r in range(START, END + 1):
for c in (4, 5, 6, 7):
if ws.cell(r, c).value is not None:
ws.cell(r, c).alignment = center
for r in range(1, END + 1):
ws.row_dimensions[r].height = 15
# K 하이퍼링크 재설정
for r in range(START, END + 1):
cell = ws.cell(r, 11)
u = cell.value
if isinstance(u, str) and u.startswith(('http://', 'https://')):
cell.hyperlink = u
old = cell.font
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
bold=old.bold, italic=old.italic,
color='0000FF', underline='single')
cell.alignment = left
wb.save(xlsx)
print(f'저장 완료: {xlsx} (총 {len(out_rows)}행)')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,350 @@
"""본문 탭(서브내비) → 카테고리 하위 확장.
규칙(공주시 기준, 일반화):
본문에서 '탭 UL' 찾아, 아래 조건을 모두 만족하는 그룹만 하위 카테고리로 확장한다.
1) UL(또는 직계 div) class 패턴(tab-ul ) 있다.
2) 탭이 2 이상이고, 모든 href '실제 페이지 링크'.
(#anchor 인페이지 탭, javascript:, 빈 href, 쿼리스트링만 다른 필터 탭은 제외)
3) 현재 페이지 URL URL 집합에 포함된다(= 자기 자신의 그룹).
4) '엑셀에 아직 없는 URL' 1 이상 있다(이미 사이트맵에 있으면 상위 nav 제외).
확장 방식(=사용자 지시: 탭들은 단계 아래 컬럼으로):
- 원래 행의 leaf 컬럼(E~J 가장 깊은 ) 부모(카테고리) 두고,
탭들을 leaf+1 컬럼에 순서대로 채운다.
- 현재 페이지와 같은 URL의 = 원래 행을 재사용(L~T 기존 데이터 보존, G만 탭명 기입).
- 나머지 = 신규 (L~T 비움, Phase 2~4 별도 수행 대상).
사용법:
python -X utf8 _tab_expand.py <엑셀경로> <BASE_URL> <도메인키워드> [--write]
) python -X utf8 _tab_expand.py 충청남도/2.공주시/충청남도_공주시.xlsx https://www.gongju.go.kr gongju.go.kr
(--write 없으면 계획만 출력 / 있으면 백업 실제 기입)
"""
import re
import sys
import shutil
import warnings
from copy import copy
from concurrent.futures import ThreadPoolExecutor, as_completed
from urllib.parse import urljoin, urlsplit, urlunsplit
import ssl
import openpyxl
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.ssl_ import create_urllib3_context
from bs4 import BeautifulSoup
from openpyxl.styles import Alignment, Font
from openpyxl.utils import get_column_letter
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
class WeakSSLAdapter(HTTPAdapter):
def init_poolmanager(self, *args, **kwargs):
ctx = create_urllib3_context()
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
ctx.options |= 0x4
ctx.check_hostname = False
ctx.verify_mode = ssl.CERT_NONE
kwargs['ssl_context'] = ctx
return super().init_poolmanager(*args, **kwargs)
SESSION = requests.Session()
SESSION.headers.update(H)
TAB_CLASS_PATS = ['tab-ul', 'tab_ul', 'tabmenu', 'tab-menu', 'tab_menu',
'tablist', 'tab-list', 'tab_list', 'subtab', 'sub-tab', 'sub_tab']
MAXCOL = 27 # AA
CAT_COLS = [5, 6, 7, 8, 9, 10] # E F G H I J
HEADER_MERGES = {'B1:R1', 'S1:W1', 'Y1:AA1'}
def norm_url(u):
"""쿼리/프래그먼트 제거 + 끝 슬래시 정리한 비교용 키."""
if not u:
return ''
s = urlsplit(u)
path = s.path.rstrip('/')
return urlunsplit((s.scheme, s.netloc, path, '', '')).lower()
def abs_url(href, base):
if not href:
return ''
href = href.strip()
if href.startswith(('javascript:', '#')):
return ''
if href.startswith(('http://', 'https://')):
return href
return urljoin(base.rstrip('/') + '/', href.lstrip('/'))
def has_tab_class(el):
cls = ' '.join(el.get('class') or []).lower()
return any(p in cls for p in TAB_CLASS_PATS), cls
def find_tab_groups(soup, base):
"""조건 1·2 만족하는 탭 그룹 후보 반환: [(cls, [(label, abs_url),...]), ...]"""
out = []
seen = set()
for ul in soup.find_all('ul'):
ok, cls = has_tab_class(ul)
if not ok:
continue # UL 자신에 탭 클래스 필요 (쿼리필터·무클래스 ul 배제)
links = []
real = True
for a in ul.find_all('a'):
t = a.get_text(strip=True).replace('\xa0', '').strip()
raw = (a.get('href') or '').strip()
au = abs_url(raw, base)
if not t:
continue
if not au: # #anchor / javascript / 빈 href → 인페이지 탭
real = False
break
links.append((t, au))
if not real or len(links) < 2:
continue
key = tuple(nu for _, (nu) in [(t, u) for t, u in links])
if key in seen:
continue
seen.add(key)
out.append((cls, links))
return out
def fetch(r, url):
try:
resp = SESSION.get(url, timeout=15, verify=False)
# 바이트로 받아 BeautifulSoup가 meta charset으로 인코딩 판별 (apparent_encoding 오판 방지)
return r, url, resp.content, None
except Exception as e:
return r, url, None, str(e)[:60]
def load_flat(ws):
"""병합값을 채워넣어 행별 완전 데이터로 평탄화. 반환: rows(list of dict), 스타일/하이퍼링크 캡처."""
# 1) 병합값 채우기 (헤더 제외, 데이터 컬럼)
for mr in list(ws.merged_cells.ranges):
s = str(mr)
if s in HEADER_MERGES:
continue
top = ws.cell(mr.min_row, mr.min_col).value
ws.unmerge_cells(s)
for rr in range(mr.min_row, mr.max_row + 1):
for cc in range(mr.min_col, mr.max_col + 1):
if ws.cell(rr, cc).value in (None, ''):
ws.cell(rr, cc).value = top
rows = []
for r in range(3, ws.max_row + 1):
# 빈 행 스킵
if all(ws.cell(r, c).value in (None, '') for c in range(2, MAXCOL + 1)):
continue
vals = {c: ws.cell(r, c).value for c in range(1, MAXCOL + 1)}
styles = {}
for c in range(1, MAXCOL + 1):
sc = ws.cell(r, c)
styles[c] = (copy(sc.font), copy(sc.fill), copy(sc.border),
copy(sc.alignment), sc.number_format, copy(sc.protection))
hl = ws.cell(r, 11).hyperlink
rows.append({'src': r, 'vals': vals, 'styles': styles,
'hyperlink': hl.target if hl else None})
return rows
def leaf_col(vals):
"""E~J 중 값이 있는 가장 깊은 컬럼 인덱스."""
deep = 5
for c in CAT_COLS:
if vals.get(c) not in (None, ''):
deep = c
return deep
def main():
xlsx, base, domain = sys.argv[1], sys.argv[2], sys.argv[3]
write = '--write' in sys.argv
if '--weak-ssl' in sys.argv:
SESSION.mount('https://', WeakSSLAdapter())
wb = openpyxl.load_workbook(xlsx)
ws = wb.active
rows = load_flat(ws)
existing = set()
for row in rows:
u = row['vals'].get(11)
if isinstance(u, str) and u.startswith('http'):
existing.add(norm_url(u))
# 같은 도메인 행 fetch
targets = [(row['src'], row['vals'].get(11)) for row in rows
if isinstance(row['vals'].get(11), str) and domain in row['vals'].get(11)]
html_by_src = {}
with ThreadPoolExecutor(max_workers=8) as ex:
futs = [ex.submit(fetch, r, u) for r, u in targets]
for fut in as_completed(futs):
r, url, html, err = fut.result()
if html:
html_by_src[r] = (url, html)
# 행별 확장 계획
plan = {} # src_row -> ordered [(label, url, is_existing)]
for row in rows:
src = row['src']
if src not in html_by_src:
continue
url, html = html_by_src[src]
cur = norm_url(url)
soup = BeautifulSoup(html, 'html.parser')
groups = find_tab_groups(soup, base)
chosen = None
for cls, links in groups:
tab_norms = [norm_url(u) for _, u in links]
if cur not in tab_norms:
continue # 조건3: 자기 탭그룹만
new_cnt = sum(1 for n in tab_norms if n not in existing)
if new_cnt < 1:
continue # 조건4: 신규 0 → 상위nav, 제외
chosen = links
break
if chosen:
ch = []
for label, u in chosen:
ch.append((label, u, norm_url(u) == cur))
plan[src] = ch
# 계획 출력
print(f'=== 탭 확장 계획: {len(plan)}개 그룹 ===\n')
total_new = 0
for src in sorted(plan):
row = next(r for r in rows if r['src'] == src)
v = row['vals']
lc = leaf_col(v)
childc = lc + 1
ev = v.get(5) or ''
fv = v.get(6) or ''
print(f'[행{src}] E={ev} F={fv} '
f'(leaf={get_column_letter(lc)} → 자식={get_column_letter(childc)})')
for label, u, exist in plan[src]:
tag = '재사용' if exist else '신규+'
if not exist:
total_new += 1
print(f' [{tag}] {label} -> {u}')
print()
print(f'신규 추가 행: {total_new}개 / 기존 {len(rows)}행 → 최종 {len(rows)+total_new}')
if not write:
print('\n(계획만 출력. 실제 기입하려면 --write)')
return
# ===== 실제 기입 =====
backup = xlsx.replace('.xlsx', '_backup_tab전.xlsx')
shutil.copy(xlsx, backup)
print(f'\n백업: {backup}')
# 새 평탄 행 목록 구성
out_rows = [] # 각 원소 = {'vals':..., 'style_src':row_dict, 'url':..}
for row in rows:
src = row['src']
if src in plan:
v = row['vals']
lc = leaf_col(v)
childc = lc + 1
for label, u, exist in plan[src]:
nv = dict(v)
# leaf 보다 깊은 컬럼 초기화 후 자식 컬럼에 탭명
for c in CAT_COLS:
if c > lc:
nv[c] = None
nv[childc] = label
if exist:
nv[11] = v.get(11) # 기존 URL/데이터 유지
out_rows.append({'vals': nv, 'style': row, 'url': nv.get(11)})
else:
for c in range(12, MAXCOL + 1):
nv[c] = None # L~T 비움 (신규)
nv[11] = u
nv[19] = None
out_rows.append({'vals': nv, 'style': row, 'url': u})
else:
out_rows.append({'vals': dict(row['vals']), 'style': row, 'url': row['vals'].get(11)})
# 데이터 영역 클리어
for r in range(3, ws.max_row + 1):
for c in range(1, MAXCOL + 1):
ws.cell(r, c).value = None
ws.cell(r, c).hyperlink = None
START = 3
for i, orow in enumerate(out_rows):
r = START + i
sty = orow['style']['styles']
for c in range(1, MAXCOL + 1):
cell = ws.cell(r, c)
f, fl, bd, al, nf, pr = sty[c]
cell.font = copy(f); cell.fill = copy(fl); cell.border = copy(bd)
cell.alignment = copy(al); cell.number_format = nf; cell.protection = copy(pr)
v = orow['vals']
for c in range(1, MAXCOL + 1):
ws.cell(r, c).value = v.get(c)
ws.cell(r, 2).value = i + 1 # B 순번 재부여
END = START + len(out_rows) - 1
# D/E/F 재병합 (G는 leaf라 병합 안 함)
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
def merge_runs(col_letter, col_idx, group_cols=()):
cur_val = ws.cell(START, col_idx).value
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
run_start = START
runs = []
for r in range(START + 1, END + 1):
val = ws.cell(r, col_idx).value
grp = tuple(ws.cell(r, gg).value for gg in group_cols)
if val == cur_val and grp == cur_grp:
continue
if cur_val not in (None, '') and r - 1 > run_start:
runs.append((run_start, r - 1))
cur_val, cur_grp, run_start = val, grp, r
if cur_val not in (None, '') and END > run_start:
runs.append((run_start, END))
for s, e in runs:
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
ws.cell(s, col_idx).alignment = center
merge_runs('F', 6, group_cols=(4, 5))
merge_runs('E', 5, group_cols=(4,))
merge_runs('D', 4)
for r in range(START, END + 1):
for c in (4, 5, 6, 7):
if ws.cell(r, c).value is not None:
ws.cell(r, c).alignment = center
for r in range(1, END + 1):
ws.row_dimensions[r].height = 15
# K 하이퍼링크 재설정
for r in range(START, END + 1):
cell = ws.cell(r, 11)
u = cell.value
if isinstance(u, str) and u.startswith(('http://', 'https://')):
cell.hyperlink = u
old = cell.font
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
bold=old.bold, italic=old.italic,
color='0000FF', underline='single')
cell.alignment = left
wb.save(xlsx)
print(f'저장 완료: {xlsx} (총 {len(out_rows)}행)')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,439 @@
"""본문 탭(서브내비) → 카테고리 하위 확장.
규칙(공주시 기준, 일반화):
본문에서 '탭 UL' 찾아, 아래 조건을 모두 만족하는 그룹만 하위 카테고리로 확장한다.
1) UL(또는 직계 div) class 패턴(tab-ul ) 있다.
2) 탭이 2 이상이고, 모든 href '실제 페이지 링크'.
(#anchor 인페이지 탭, javascript:, 빈 href, 쿼리스트링만 다른 필터 탭은 제외)
3) 현재 페이지 URL URL 집합에 포함된다(= 자기 자신의 그룹).
4) '엑셀에 아직 없는 URL' 1 이상 있다(이미 사이트맵에 있으면 상위 nav 제외).
확장 방식(=사용자 지시: 탭들은 단계 아래 컬럼으로):
- 원래 행의 leaf 컬럼(E~J 가장 깊은 ) 부모(카테고리) 두고,
탭들을 leaf+1 컬럼에 순서대로 채운다.
- 현재 페이지와 같은 URL의 = 원래 행을 재사용(L~T 기존 데이터 보존, G만 탭명 기입).
- 나머지 = 신규 (L~T 비움, Phase 2~4 별도 수행 대상).
사용법:
python -X utf8 _tab_expand.py <엑셀경로> <BASE_URL> <도메인키워드> [--write]
) python -X utf8 _tab_expand.py 충청남도/2.공주시/충청남도_공주시.xlsx https://www.gongju.go.kr gongju.go.kr
(--write 없으면 계획만 출력 / 있으면 백업 실제 기입)
"""
import re
import sys
import shutil
import warnings
from copy import copy
from concurrent.futures import ThreadPoolExecutor, as_completed
from urllib.parse import urljoin, urlsplit, urlunsplit
import ssl
import openpyxl
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.ssl_ import create_urllib3_context
from bs4 import BeautifulSoup
from openpyxl.styles import Alignment, Font
from openpyxl.utils import get_column_letter
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
class WeakSSLAdapter(HTTPAdapter):
def init_poolmanager(self, *args, **kwargs):
ctx = create_urllib3_context()
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
ctx.options |= 0x4
ctx.check_hostname = False
ctx.verify_mode = ssl.CERT_NONE
kwargs['ssl_context'] = ctx
return super().init_poolmanager(*args, **kwargs)
SESSION = requests.Session()
SESSION.headers.update(H)
TAB_CLASS_PATS = ['tab-ul', 'tab_ul', 'tabmenu', 'tab-menu', 'tab_menu',
'tablist', 'tab-list', 'tab_list', 'subtab', 'sub-tab', 'sub_tab']
# 토큰 단위 탭 클래스 정규식: 'basic_tab', 'tab_wrap', 'tab-ul', 단독 'tab' 등 매칭.
# (UL 자신뿐 아니라 직계 부모 div 클래스도 검사 → div.basic_tab > ul 구조 대응)
TAB_CLASS_RE = re.compile(
r'(?:^|[-_ ])tab(?:[-_ ]|$)|tabmenu|tablist|tab[-_]?(?:ul|wrap|list|menu)|basic[-_]tab')
MAXCOL = 27 # AA
CAT_COLS = [5, 6, 7, 8, 9, 10] # E F G H I J
HEADER_MERGES = {'B1:R1', 'S1:W1', 'Y1:AA1'}
def norm_url(u):
"""비교용 키: 프래그먼트 제거 + 끝 슬래시 정리 + 쿼리 정렬 보존.
(쿼리만 다른 게시판 분류 ?code=A vs ?code=B 서로 다른 URL로 구분
is_cur·existing 중복판정 오류 방지)"""
if not u:
return ''
s = urlsplit(u)
path = s.path.rstrip('/')
q = '&'.join(sorted(s.query.split('&'))) if s.query else ''
return urlunsplit((s.scheme, s.netloc, path, q, '')).lower()
def path_key(u):
"""쿼리 제거한 경로 키 (필터 탭 판별용)."""
if not u:
return ''
s = urlsplit(u)
return urlunsplit((s.scheme, s.netloc, s.path.rstrip('/'), '', '')).lower()
def abs_url(href, base):
if not href:
return ''
href = href.strip()
if href.startswith(('javascript:', '#')):
return ''
if href.startswith(('http://', 'https://')):
return href
return urljoin(base.rstrip('/') + '/', href.lstrip('/'))
def has_tab_class(el):
if el is None or not getattr(el, 'get', None):
return False, ''
cls = ' '.join(el.get('class') or []).lower()
ok = any(p in cls for p in TAB_CLASS_PATS) or bool(TAB_CLASS_RE.search(cls))
return ok, cls
def _has_on_li(ul):
"""ul 직계 li(또는 그 a)에 활성 탭 마커(on/active/current/selected)가 있나."""
marks = {'on', 'active', 'current', 'selected', 'sel'}
for li in ul.find_all('li', recursive=False):
if marks & set(c.lower() for c in (li.get('class') or [])):
return True
a = li.find('a')
if a and (marks & set(c.lower() for c in (a.get('class') or []))):
return True
return False
def find_tab_groups(soup, base):
"""조건 1·2 만족하는 탭 그룹 후보 반환: [(cls, [(label, abs_url),...], has_on), ...]"""
out = []
seen = set()
for ul in soup.find_all('ul'):
ok, cls = has_tab_class(ul)
if not ok:
# UL 무클래스라도 직계 부모 div 에 탭 클래스가 있으면 인정 (div.basic_tab > ul)
pok, pcls = has_tab_class(ul.parent)
if not pok:
continue # 쿼리필터·무클래스 ul 배제
cls = pcls
links = []
real = True
for a in ul.find_all('a'):
t = a.get_text(strip=True).replace('\xa0', '').strip()
raw = (a.get('href') or '').strip()
au = abs_url(raw, base)
if not t:
continue
if not au: # #anchor / javascript / 빈 href → 인페이지 탭
real = False
break
links.append((t, au))
if not real or len(links) < 2:
continue
key = tuple(nu for _, (nu) in [(t, u) for t, u in links])
if key in seen:
continue
seen.add(key)
out.append((cls, links, _has_on_li(ul)))
return out
def fetch(r, url):
try:
resp = SESSION.get(url, timeout=15, verify=False)
# 바이트로 받아 BeautifulSoup가 meta charset으로 인코딩 판별 (apparent_encoding 오판 방지)
return r, url, resp.content, None
except Exception as e:
return r, url, None, str(e)[:60]
def load_flat(ws):
"""병합값을 채워넣어 행별 완전 데이터로 평탄화. 반환: rows(list of dict), 스타일/하이퍼링크 캡처."""
# 1) 병합값 채우기 (헤더 제외, 데이터 컬럼)
for mr in list(ws.merged_cells.ranges):
s = str(mr)
if s in HEADER_MERGES:
continue
top = ws.cell(mr.min_row, mr.min_col).value
ws.unmerge_cells(s)
for rr in range(mr.min_row, mr.max_row + 1):
for cc in range(mr.min_col, mr.max_col + 1):
if ws.cell(rr, cc).value in (None, ''):
ws.cell(rr, cc).value = top
rows = []
for r in range(3, ws.max_row + 1):
# 빈 행 스킵
if all(ws.cell(r, c).value in (None, '') for c in range(2, MAXCOL + 1)):
continue
vals = {c: ws.cell(r, c).value for c in range(1, MAXCOL + 1)}
styles = {}
for c in range(1, MAXCOL + 1):
sc = ws.cell(r, c)
styles[c] = (copy(sc.font), copy(sc.fill), copy(sc.border),
copy(sc.alignment), sc.number_format, copy(sc.protection))
hl = ws.cell(r, 11).hyperlink
rows.append({'src': r, 'vals': vals, 'styles': styles,
'hyperlink': hl.target if hl else None})
return rows
def leaf_col(vals):
"""E~J 중 값이 있는 가장 깊은 컬럼 인덱스."""
deep = 5
for c in CAT_COLS:
if vals.get(c) not in (None, ''):
deep = c
return deep
def main():
xlsx, base, domain = sys.argv[1], sys.argv[2], sys.argv[3]
write = '--write' in sys.argv
# menuCd 벤더(고창·임실·정읍·진안 등 index.{name}?menuCd=…): 모든 페이지가 같은 경로,
# 쿼리(menuCd)만 다름 = 다른 페이지. → 쿼리-경로 가드(조건2) 스킵 + cur∈탭(조건3) 대신
# div.basic_tab 등에 활성탭(li.on) 마커가 있는 '자기 sub-nav'만 인정(랜딩 URL이 탭과 달라도).
menucd = '--menucd' in sys.argv
# 서산 등 쿼리기반(contents.do?key=, selectBbsNttList.do?bbsNo=) 벤더:
# 모든 탭이 같은 경로·쿼리만 다름(=다른 페이지)이고 한 탭바에 페이지+게시판 혼재.
# → 쿼리가드(같은경로=필터 제외) 끄고, div.tab_menu>ul.tab_button 그룹만 인정.
noqg = '--noqg' in sys.argv
if '--weak-ssl' in sys.argv:
SESSION.mount('https://', WeakSSLAdapter())
wb = openpyxl.load_workbook(xlsx)
ws = wb.active
rows = load_flat(ws)
existing = set()
for row in rows:
u = row['vals'].get(11)
if isinstance(u, str) and u.startswith('http'):
existing.add(norm_url(u))
# 같은 도메인 행 fetch
targets = [(row['src'], row['vals'].get(11)) for row in rows
if isinstance(row['vals'].get(11), str) and domain in row['vals'].get(11)]
html_by_src = {}
with ThreadPoolExecutor(max_workers=8) as ex:
futs = [ex.submit(fetch, r, u) for r, u in targets]
for fut in as_completed(futs):
r, url, html, err = fut.result()
if html:
html_by_src[r] = (url, html)
# 이미 사이트맵에 자식행이 있는 '랜딩행' 집합 (다음 평탄행이 같은 상위카테고리 + 더 깊은 leaf).
# menuCd 벤더 랜딩은 첫 자식의 탭을 렌더하므로, 확장하면 그 자식행과 중복 → 제외.
landing_src = set()
for i in range(len(rows) - 1):
v, nv = rows[i]['vals'], rows[i + 1]['vals']
lc = leaf_col(v)
if lc < 10 and nv.get(lc + 1) not in (None, '') \
and all((v.get(c) or '') == (nv.get(c) or '') for c in CAT_COLS if c <= lc):
landing_src.add(rows[i]['src'])
# 행별 확장 계획
plan = {} # src_row -> ordered [(label, url, is_existing)]
planned_new = set() # 이미 어느 부모행이 추가한 신규 탭 URL(전역 중복 방지)
for row in rows:
src = row['src']
if src not in html_by_src:
continue
if menucd and src in landing_src:
continue # 자식 보유 랜딩 → 확장 금지(첫 자식 탭 중복 방지)
url, html = html_by_src[src]
cur = norm_url(url)
soup = BeautifulSoup(html, 'html.parser')
groups = find_tab_groups(soup, base)
chosen = None
for cls, links, has_on in groups:
if menucd:
# menuCd 벤더: 현재 페이지와 '같은 경로'(menuCd만 다른 형제)인 탭만 인정.
# → /gochang/toc/GC…(향토문화대전 백과 목록) 등 다른경로 링크 배제.
# 활성탭(li.on) 마커가 있는 자기 sub-nav만(랜딩 URL이 탭에 없어도 OK).
if not has_on:
continue
links = [(t, u) for t, u in links if path_key(u) == path_key(cur)]
if len(links) < 2:
continue
tab_norms = [norm_url(u) for _, u in links]
else:
if noqg and not ('tab_button' in cls or 'tab_menu' in cls):
continue # 서산: 본문 탭바(tab_menu/tab_button)만 — GNB·푸터 nav 배제
tab_norms = [norm_url(u) for _, u in links]
if not noqg and len({path_key(u) for _, u in links}) == 1:
continue # 조건2: 같은 경로(쿼리만 다른 필터 탭) → 제외 (noqg면 쿼리=다른페이지라 허용)
if cur not in tab_norms:
continue # 조건3: 자기 탭그룹만
new_cnt = sum(1 for n in tab_norms if n not in existing)
if new_cnt < 1:
continue # 조건4: 신규 0 → 상위nav, 제외
chosen = links
break
if chosen:
ch = []
for label, u in chosen:
is_cur = norm_url(u) == cur
nu = norm_url(u)
if not is_cur and nu in existing:
continue # 이미 사이트맵에 별도 행으로 존재 → 중복행 방지(skip)
if not is_cur and nu in planned_new:
continue # 다른 부모행이 이미 추가한 신규 URL → 전역 중복 방지
if not is_cur:
planned_new.add(nu)
is_ext = domain not in u # 같은 도메인 아님 = 외부 사이트 탭
ch.append((label, u, is_cur, is_ext))
# 신규(비-cur) 탭이 1개 이상일 때만 확장. 아니면 원본행 그대로 둠(삭제 방지).
if any(not c[2] for c in ch):
plan[src] = ch
# 계획 출력
print(f'=== 탭 확장 계획: {len(plan)}개 그룹 ===\n')
total_new = 0
for src in sorted(plan):
row = next(r for r in rows if r['src'] == src)
v = row['vals']
lc = leaf_col(v)
childc = lc + 1
ev = v.get(5) or ''
fv = v.get(6) or ''
print(f'[행{src}] E={ev} F={fv} '
f'(leaf={get_column_letter(lc)} → 자식={get_column_letter(childc)})')
for label, u, is_cur, is_ext in plan[src]:
if is_cur:
tag = '재사용'
elif is_ext:
tag = '신규+(외부=사이트)'
total_new += 1
else:
tag = '신규+'
total_new += 1
print(f' [{tag}] {label} -> {u}')
print()
print(f'신규 추가 행: {total_new}개 / 기존 {len(rows)}행 → 최종 {len(rows)+total_new}')
if not write:
print('\n(계획만 출력. 실제 기입하려면 --write)')
return
# ===== 실제 기입 =====
backup = xlsx.replace('.xlsx', '_backup_tab전.xlsx')
shutil.copy(xlsx, backup)
print(f'\n백업: {backup}')
# 새 평탄 행 목록 구성
out_rows = [] # 각 원소 = {'vals':..., 'style_src':row_dict, 'url':..}
for row in rows:
src = row['src']
if src in plan:
v = row['vals']
lc = leaf_col(v)
childc = lc + 1
for label, u, is_cur, is_ext in plan[src]:
nv = dict(v)
# leaf 보다 깊은 컬럼 초기화 후 자식 컬럼에 탭명
for c in CAT_COLS:
if c > lc:
nv[c] = None
nv[childc] = label
if is_cur:
nv[11] = v.get(11) # 기존 URL/데이터 유지
out_rows.append({'vals': nv, 'style': row, 'url': nv.get(11)})
else:
for c in range(12, MAXCOL + 1):
nv[c] = None # L~T 비움 (신규)
nv[11] = u
nv[19] = None
if is_ext: # 외부 사이트 탭 → 즉시 L=사이트/M=1(Phase234 대상 아님)
nv[12] = '사이트'
nv[13] = 1
out_rows.append({'vals': nv, 'style': row, 'url': u})
else:
out_rows.append({'vals': dict(row['vals']), 'style': row, 'url': row['vals'].get(11)})
# 데이터 영역 클리어
for r in range(3, ws.max_row + 1):
for c in range(1, MAXCOL + 1):
ws.cell(r, c).value = None
ws.cell(r, c).hyperlink = None
START = 3
for i, orow in enumerate(out_rows):
r = START + i
sty = orow['style']['styles']
for c in range(1, MAXCOL + 1):
cell = ws.cell(r, c)
f, fl, bd, al, nf, pr = sty[c]
cell.font = copy(f); cell.fill = copy(fl); cell.border = copy(bd)
cell.alignment = copy(al); cell.number_format = nf; cell.protection = copy(pr)
v = orow['vals']
for c in range(1, MAXCOL + 1):
ws.cell(r, c).value = v.get(c)
ws.cell(r, 2).value = i + 1 # B 순번 재부여
END = START + len(out_rows) - 1
# D/E/F 재병합 (G는 leaf라 병합 안 함)
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
def merge_runs(col_letter, col_idx, group_cols=()):
cur_val = ws.cell(START, col_idx).value
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
run_start = START
runs = []
for r in range(START + 1, END + 1):
val = ws.cell(r, col_idx).value
grp = tuple(ws.cell(r, gg).value for gg in group_cols)
if val == cur_val and grp == cur_grp:
continue
if cur_val not in (None, '') and r - 1 > run_start:
runs.append((run_start, r - 1))
cur_val, cur_grp, run_start = val, grp, r
if cur_val not in (None, '') and END > run_start:
runs.append((run_start, END))
for s, e in runs:
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
ws.cell(s, col_idx).alignment = center
merge_runs('F', 6, group_cols=(4, 5))
merge_runs('E', 5, group_cols=(4,))
merge_runs('D', 4)
for r in range(START, END + 1):
for c in (4, 5, 6, 7):
if ws.cell(r, c).value is not None:
ws.cell(r, c).alignment = center
for r in range(1, END + 1):
ws.row_dimensions[r].height = 15
# K 하이퍼링크 재설정
for r in range(START, END + 1):
cell = ws.cell(r, 11)
u = cell.value
if isinstance(u, str) and u.startswith(('http://', 'https://')):
cell.hyperlink = u
old = cell.font
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
bold=old.bold, italic=old.italic,
color='0000FF', underline='single')
cell.alignment = left
wb.save(xlsx)
print(f'저장 완료: {xlsx} (총 {len(out_rows)}행)')
if __name__ == '__main__':
main()

View File

@ -0,0 +1,220 @@
# -*- coding: utf-8 -*-
"""탭 확장으로 새로 추가된 행만 골라 Phase 2~4(L/M/N/O/P) 수집.
- 대상: L(게시판형태) 비어있고 K가 같은 도메인 http URL = 신규 .
- L/M/N: 해당 기관 _phase234.py get_body/detect_form/detect_media/extract_detail_urls 재사용.
- O/P : 확정 KOGL 규칙(_recheck_kogl_all BROAD_IMG_PAT + detect_split + decide_O).
이미지명 우선, 게시판은 같은도메인 상세 5 추적, 이미지링크면 S열 '링크주소 오기'.
- 인코딩: 바이트로 받아 BeautifulSoup 자동판별(읍면동 오판 방지).
- 기존 (L 이미 채워짐) 절대 건드리지 않음.
사용: python -X utf8 _tab_phase234.py <엑셀경로> <Phase234 모듈경로> <도메인키워드> [body_sel(콤마)]
- 개별형(공주시 _phase234.py: get_body(soup)) body_sel 생략
- 일괄형(_chungnam_phase234_all.py: get_body(soup, sel)) body_sel 지정(미지정 #txt,#contents,main)
) python -X utf8 _tab_phase234.py 충청남도/2.공주시/공주시_탭확장.xlsx 충청남도/2.공주시/_phase234.py gongju.go.kr
python -X utf8 _tab_phase234.py 충청남도/4.논산시/충청남도_논산시.xlsx _chungnam_phase234_all.py nonsan.go.kr "#txt,#contents,main"
"""
import re
import sys
import time
import inspect
import importlib.util
import warnings
from urllib.parse import urlparse
from concurrent.futures import ThreadPoolExecutor, as_completed
import ssl
import openpyxl
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.ssl_ import create_urllib3_context
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
class WeakSSLAdapter(HTTPAdapter):
def init_poolmanager(self, *args, **kwargs):
ctx = create_urllib3_context()
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
ctx.options |= 0x4
ctx.check_hostname = False
ctx.verify_mode = ssl.CERT_NONE
kwargs['ssl_context'] = ctx
return super().init_poolmanager(*args, **kwargs)
SESSION = requests.Session()
SESSION.headers.update(H)
BROAD_IMG_PAT = re.compile(r'(?:new_)?img_open(?:type|code)(\d{1,2})\.(?:png|jpe?g|gif)', re.I)
def load_module(path):
spec = importlib.util.spec_from_file_location('city_p234', path)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod
def fetch_soup(url, timeout=14):
try:
r = SESSION.get(url, timeout=timeout, verify=False)
if r.status_code == 200:
return BeautifulSoup(r.content, 'html.parser') # 바이트 → 자동 인코딩
except Exception:
pass
return None
def valid(n):
return 1 <= n <= 4
def detect_split(body, LINK_PAT):
img_t, link_t = set(), set()
for a in body.find_all('a', href=True):
m = LINK_PAT.search(a['href'])
if m and valid(int(m.group(1))):
link_t.add(int(m.group(1)))
blob = ' '.join(filter(None, (img.get('src', '') for img in body.find_all('img'))))
blob += ' ' + ' '.join(el.get('style', '') for el in body.find_all(style=True))
blob += ' ' + str(body)
for m in BROAD_IMG_PAT.finditer(blob):
n = int(m.group(1))
if valid(n):
img_t.add(n)
return img_t, link_t
def decide_O(img_t, link_t):
if img_t:
return ','.join(f'{n}유형' for n in sorted(img_t)), (bool(link_t) and link_t != img_t)
if link_t:
if {1, 2, 3, 4}.issubset(link_t):
return '미부착', False
return ','.join(f'{n}유형' for n in sorted(link_t)), False
return '미부착', False
def domain3(host):
labels = (host or '').split('.')
return '.'.join(labels[-3:]) if len(labels) >= 3 else host
def same_site(a, b):
return domain3(urlparse(a).hostname) == domain3(urlparse(b).hostname)
def main():
xlsx, p234_path, domain = sys.argv[1], sys.argv[2], sys.argv[3]
if '--weak-ssl' in sys.argv:
SESSION.mount('https://', WeakSSLAdapter())
body_sel = None
if len(sys.argv) > 4 and sys.argv[4].strip() and not sys.argv[4].startswith('--'):
body_sel = [s.strip() for s in sys.argv[4].split(',') if s.strip()]
mod = load_module(p234_path)
LINK_PAT = mod.KOGL_LINK_PAT
# get_body 시그니처 자동 대응: 일괄형은 (soup, selectors), 개별형은 (soup)
needs_sel = len(inspect.signature(mod.get_body).parameters) >= 2
sel = body_sel or ['#txt', '#contents', 'main']
get_body = (lambda s: mod.get_body(s, sel)) if needs_sel else mod.get_body
print(f'get_body 인자 {2 if needs_sel else 1}개 | body_sel={sel if needs_sel else "(미사용)"}')
wb = openpyxl.load_workbook(xlsx)
ws = wb.active
targets = []
for r in range(3, ws.max_row + 1):
L = ws.cell(r, 12).value
url = ws.cell(r, 11).value
if L not in (None, '') :
continue # 기존 행 보존
if not (isinstance(url, str) and domain in url):
continue
targets.append((r, url))
print(f'대상 신규행: {len(targets)}')
def work(t):
r, url = t
soup = fetch_soup(url)
if soup is None:
return r, {'note': '접근 실패'}
body = get_body(soup)
form, count = mod.detect_form(body)
has_img, has_vid, has_txt = mod.detect_media(body)
img_t, link_t = detect_split(body, LINK_PAT)
P_loc = '게시판' if img_t or link_t else ''
if form == '게시판':
for du in mod.extract_detail_urls(body, url, limit=5):
if not same_site(url, du):
continue
ds = fetch_soup(du, 10)
if ds is None:
continue
db = get_body(ds)
di, dv, dt = mod.detect_media(db)
has_img = has_img or di; has_vid = has_vid or dv; has_txt = has_txt or dt
dimg, dlink = detect_split(db, LINK_PAT)
if (dimg or dlink) and not P_loc:
P_loc = '게시물'
img_t |= dimg; link_t |= dlink
N = mod.n_string(has_txt, has_img, has_vid)
O, mismatch = decide_O(img_t, link_t)
P = '' if O == '미부착' else (P_loc or '게시판')
return r, {'L': form, 'M': count if form == '게시판' else 1,
'N': N, 'O': O, 'P': P, 'mismatch': mismatch}
t0 = time.time()
results = {}
done = 0
with ThreadPoolExecutor(max_workers=8) as ex:
futs = [ex.submit(work, t) for t in targets]
for fut in as_completed(futs):
r, res = fut.result()
results[r] = res
done += 1
if done % 40 == 0:
print(f' 진행 {done}/{len(targets)} ({time.time()-t0:.0f}s)')
print(f'크롤링 완료 ({time.time()-t0:.0f}s)')
fail = 0
for r, res in results.items():
if res.get('note'):
fail += 1
if not ws.cell(r, 19).value:
ws.cell(r, 19).value = res['note']
continue
ws.cell(r, 12).value = res['L']
ws.cell(r, 13).value = res['M']
ws.cell(r, 14).value = res['N']
ws.cell(r, 15).value = res['O']
if res['P']:
ws.cell(r, 16).value = res['P']
if res['mismatch']:
cur = (ws.cell(r, 19).value or '').strip()
ws.cell(r, 19).value = '링크주소 오기' if not cur else cur + ' / 링크주소 오기'
try:
wb.save(xlsx)
saved = xlsx
except PermissionError:
saved = xlsx.replace('.xlsx', '_LP.xlsx')
wb.save(saved)
print(f'!! 원본 잠김(Excel 열림). 대체 저장: {saved}')
from collections import Counter
Lc = Counter(res.get('L') for res in results.values())
Oc = Counter(res.get('O') for res in results.values())
print(f'저장: {xlsx}')
print('L 분포:', dict(Lc))
print('O 분포:', dict(Oc))
print(f'접근실패: {fail}')
if __name__ == '__main__':
main()

265
_스크립트/_tab_run.log Normal file
View File

@ -0,0 +1,265 @@
대상 기관: 39개 (모드=run)
━━━ 충청남도 금산군 (행~526, base=https://www.geumsan.go.kr, domain=geumsan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\금산군.xlsx https://www.geumsan.go.kr geumsan.go.kr
→ 확장그룹 3개 / 신규행 9개
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\금산군.xlsx https://www.geumsan.go.kr geumsan.go.kr --write
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\금산군.xlsx _phase234.py geumsan.go.kr
대상 신규행: 13개
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\금산군.xlsx
L 분포: {None: 4, '페이지': 4, '게시판': 5}
O 분포: {None: 4, '미부착': 7, '4유형': 2}
접근실패: 4
━━━ 충청남도 논산시 (행~738, base=https://nonsan.go.kr, domain=nonsan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\4.논산시\논산시.xlsx https://nonsan.go.kr nonsan.go.kr
→ 확장그룹 7개 / 신규행 22개
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\4.논산시\논산시.xlsx https://nonsan.go.kr nonsan.go.kr --write
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\4.논산시\논산시.xlsx _chungnam_phase234_all.py nonsan.go.kr #txt,#contents,main
대상 신규행: 22개
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\4.논산시\논산시.xlsx
L 분포: {'페이지': 19, '게시판': 3}
O 분포: {'미부착': 22}
접근실패: 0
━━━ 충청남도 당진시 (행~315, base=https://www.dangjin.go.kr, domain=dangjin.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\5.당진시\당진시.xlsx https://www.dangjin.go.kr dangjin.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 충청남도 보령시 (행~573, base=https://www.brcn.go.kr, domain=brcn.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\6.보령시\보령시.xlsx https://www.brcn.go.kr brcn.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 충청남도 부여군 (행~315, base=https://www.buyeo.go.kr, domain=buyeo.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\7.부여군\부여군.xlsx https://www.buyeo.go.kr buyeo.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 충청남도 서산시 (행~349, base=https://www.seosan.go.kr, domain=seosan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\8.서산시\서산시.xlsx https://www.seosan.go.kr seosan.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 충청남도 서천군 (행~330, base=https://www.seocheon.go.kr, domain=seocheon.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\9.서천군\서천군.xlsx https://www.seocheon.go.kr seocheon.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 충청남도 아산시 (행~315, base=https://www.asan.go.kr, domain=asan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\10.아산시\아산시.xlsx https://www.asan.go.kr asan.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 충청남도 예산군 (행~387, base=https://www.yesan.go.kr, domain=yesan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\11.예산군\예산군.xlsx https://www.yesan.go.kr yesan.go.kr
→ 확장그룹 91개 / 신규행 233개
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\11.예산군\예산군.xlsx https://www.yesan.go.kr yesan.go.kr --write
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\11.예산군\예산군.xlsx _chungnam_phase234_all.py yesan.go.kr #txt,#contents,main
대상 신규행: 237개
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\11.예산군\예산군.xlsx
L 분포: {'페이지': 167, '게시판': 70}
O 분포: {'미부착': 237}
접근실패: 0
━━━ 충청남도 천안시 (행~371, base=https://www.cheonan.go.kr, domain=cheonan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\12.천안시\천안시.xlsx https://www.cheonan.go.kr cheonan.go.kr
→ 확장그룹 107개 / 신규행 297개
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\12.천안시\천안시.xlsx https://www.cheonan.go.kr cheonan.go.kr --write
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\12.천안시\천안시.xlsx _chungnam_phase234_all.py cheonan.go.kr #txt,#contents,main
대상 신규행: 283개
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\12.천안시\천안시.xlsx
L 분포: {'페이지': 225, '게시판': 58}
O 분포: {'미부착': 281, '1유형': 2}
접근실패: 0
━━━ 충청남도 청양군 (행~315, base=http://www.cheongyang.go.kr, domain=cheongyang.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\13.청양군\청양군.xlsx http://www.cheongyang.go.kr cheongyang.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 충청남도 태안군 (행~315, base=https://www.taean.go.kr, domain=taean.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\14.태안군\태안군.xlsx https://www.taean.go.kr taean.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 충청남도 홍성군 (행~315, base=https://www.hongseong.go.kr, domain=hongseong.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\15.홍성군\홍성군.xlsx https://www.hongseong.go.kr hongseong.go.kr
→ 확장그룹 67개 / 신규행 215개
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\15.홍성군\홍성군.xlsx https://www.hongseong.go.kr hongseong.go.kr --write
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\15.홍성군\홍성군.xlsx _chungnam_phase234_all.py hongseong.go.kr #txt,#contents,main
대상 신규행: 213개
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\15.홍성군\홍성군.xlsx
L 분포: {'게시판': 41, '페이지': 172}
O 분포: {'미부착': 212, '4유형': 1}
접근실패: 0
━━━ 충청북도 괴산군 (행~315, base=https://www.goesan.go.kr, domain=goesan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\1.괴산군\괴산군.xlsx https://www.goesan.go.kr goesan.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 충청북도 단양군 (행~486, base=https://www.danyang.go.kr, domain=danyang.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\2.단양군\단양군.xlsx https://www.danyang.go.kr danyang.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 충청북도 영동군 [weak-ssl] (행~674, base=https://www.yd21.go.kr, domain=yd21.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\4.영동군\영동군.xlsx https://www.yd21.go.kr yd21.go.kr --weak-ssl
→ 확장그룹 20개 / 신규행 143개
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\4.영동군\영동군.xlsx https://www.yd21.go.kr yd21.go.kr --weak-ssl --write
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\4.영동군\영동군.xlsx _chungbuk_phase234_all.py yd21.go.kr #txt,#contents,main --weak-ssl
대상 신규행: 146개
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\4.영동군\영동군.xlsx
L 분포: {None: 2, '페이지': 138, '게시판': 6}
O 분포: {None: 2, '미부착': 144}
접근실패: 2
━━━ 충청북도 옥천군 (행~564, base=https://www.oc.go.kr, domain=oc.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\5.옥천군\옥천군.xlsx https://www.oc.go.kr oc.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 충청북도 음성군 (행~652, base=https://www.eumseong.go.kr, domain=eumseong.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\6.음성군\음성군.xlsx https://www.eumseong.go.kr eumseong.go.kr
→ 확장그룹 9개 / 신규행 18개
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\6.음성군\음성군.xlsx https://www.eumseong.go.kr eumseong.go.kr --write
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\6.음성군\음성군.xlsx _chungbuk_phase234_all.py eumseong.go.kr #contents,#txt,main
대상 신규행: 18개
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\6.음성군\음성군.xlsx
L 분포: {'페이지': 9, '게시판': 9}
O 분포: {'미부착': 18}
접근실패: 0
━━━ 충청북도 제천시 (행~644, base=https://www.jecheon.go.kr, domain=jecheon.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\7.제천시\제천시.xlsx https://www.jecheon.go.kr jecheon.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 충청북도 증평군 (행~315, base=https://www.jp.go.kr, domain=jp.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\8.증평군\증평군.xlsx https://www.jp.go.kr jp.go.kr
→ 확장그룹 42개 / 신규행 175개
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\8.증평군\증평군.xlsx https://www.jp.go.kr jp.go.kr --write
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\8.증평군\증평군.xlsx _chungbuk_phase234_all.py jp.go.kr #txt,#contents,main
대상 신규행: 160개
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\8.증평군\증평군.xlsx
L 분포: {'게시판': 158, None: 1, '페이지': 1}
O 분포: {'미부착': 159, None: 1}
접근실패: 1
━━━ 충청북도 진천군 (행~334, base=https://www.jincheon.go.kr, domain=jincheon.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\9.진천군\진천군.xlsx https://www.jincheon.go.kr jincheon.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 충청북도 청주시 (행~342, base=https://www.cheongju.go.kr, domain=cheongju.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\10.청주시\청주시.xlsx https://www.cheongju.go.kr cheongju.go.kr
→ 확장그룹 5개 / 신규행 97개
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\10.청주시\청주시.xlsx https://www.cheongju.go.kr cheongju.go.kr --write
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\10.청주시\청주시.xlsx _chungbuk_phase234_all.py cheongju.go.kr #contents,#txt,main
대상 신규행: 90개
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\10.청주시\청주시.xlsx
L 분포: {None: 1, '페이지': 76, '게시판': 13}
O 분포: {None: 1, '미부착': 89}
접근실패: 1
━━━ 충청북도 충주시 (행~403, base=https://www.chungju.go.kr, domain=chungju.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\11.충주시\충주시.xlsx https://www.chungju.go.kr chungju.go.kr
→ 확장그룹 3개 / 신규행 33개
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\11.충주시\충주시.xlsx https://www.chungju.go.kr chungju.go.kr --write
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\11.충주시\충주시.xlsx _chungbuk_phase234_all.py chungju.go.kr #contents,#txt,main
대상 신규행: 36개
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\11.충주시\충주시.xlsx
L 분포: {None: 2, '게시판': 30, '페이지': 4}
O 분포: {None: 2, '미부착': 34}
접근실패: 2
━━━ 전북특별자치도 고창군 (행~381, base=https://www.gochang.go.kr, domain=gochang.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\1.고창군\고창군.xlsx https://www.gochang.go.kr gochang.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 전북특별자치도 군산시 (행~639, base=https://www.gunsan.go.kr, domain=gunsan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\2.군산시\군산시.xlsx https://www.gunsan.go.kr gunsan.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 전북특별자치도 김제시 (행~454, base=https://www.gimje.go.kr, domain=gimje.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\3.김제시\김제시.xlsx https://www.gimje.go.kr gimje.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 전북특별자치도 남원시 (행~333, base=https://www.namwon.go.kr, domain=namwon.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\4.남원시\남원시.xlsx https://www.namwon.go.kr namwon.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 전북특별자치도 무주군 (행~315, base=https://www.muju.go.kr, domain=muju.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\5.무주군\무주군.xlsx https://www.muju.go.kr muju.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 전북특별자치도 부안군 (행~315, base=https://www.buan.go.kr, domain=buan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\6.부안군\부안군.xlsx https://www.buan.go.kr buan.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 전북특별자치도 순창군 (행~454, base=https://www.sunchang.go.kr, domain=sunchang.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\7.순창군\순창군.xlsx https://www.sunchang.go.kr sunchang.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 전북특별자치도 완주군 (행~331, base=https://www.wanju.go.kr, domain=wanju.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\8.완주군\완주군.xlsx https://www.wanju.go.kr wanju.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 전북특별자치도 익산시 (행~537, base=https://www.iksan.go.kr, domain=iksan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\9.익산시\익산시.xlsx https://www.iksan.go.kr iksan.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 전북특별자치도 임실군 (행~315, base=https://www.imsil.go.kr, domain=imsil.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\10.임실군\임실군.xlsx https://www.imsil.go.kr imsil.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 전북특별자치도 장수군 (행~315, base=https://www.jangsu.go.kr, domain=jangsu.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\11.장수군\장수군.xlsx https://www.jangsu.go.kr jangsu.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 전북특별자치도 전주시 (행~315, base=https://www.jeonju.go.kr, domain=jeonju.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\12.전주시\전주시.xlsx https://www.jeonju.go.kr jeonju.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 전북특별자치도 정읍시 (행~315, base=https://www.jeongeup.go.kr, domain=jeongeup.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\13.정읍시\정읍시.xlsx https://www.jeongeup.go.kr jeongeup.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 전북특별자치도 진안군 (행~315, base=http://www.jinan.go.kr, domain=jinan.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\14.진안군\진안군.xlsx http://www.jinan.go.kr jinan.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
━━━ 제주특별자치도 서귀포시 (행~366, base=https://www.seogwipo.go.kr, domain=seogwipo.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\제주특별자치도\1.서귀포시\서귀포시.xlsx https://www.seogwipo.go.kr seogwipo.go.kr
→ 확장그룹 5개 / 신규행 24개
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\제주특별자치도\1.서귀포시\서귀포시.xlsx https://www.seogwipo.go.kr seogwipo.go.kr --write
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\제주특별자치도\1.서귀포시\서귀포시.xlsx _jeonbuk_phase234_all.py seogwipo.go.kr #main-contents,#content,#contents,.contents,#txt,main,#container,#sub
대상 신규행: 22개
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\제주특별자치도\1.서귀포시\서귀포시.xlsx
L 분포: {'페이지': 10, '게시판': 11, None: 1}
O 분포: {'미부착': 21, None: 1}
접근실패: 1
━━━ 제주특별자치도 제주시 (행~315, base=https://www.jejusi.go.kr, domain=jejusi.go.kr)
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\제주특별자치도\2.제주시\제주시.xlsx https://www.jejusi.go.kr jejusi.go.kr
→ 확장그룹 0개 / 신규행 0개
(신규행 없음 → Phase234 생략)
=== 배치 종료 ===

116
_스크립트/_tab_scan.py Normal file
View File

@ -0,0 +1,116 @@
"""사이트 본문 탭(서브내비) 탐지 스캐너.
사이트 페이지를 받아 본문 UL을 찾아, 탭이 2 이상인 페이지를 보고한다.
공주시: ul.tab-ul. 다른 사이트도 흔한 클래스 패턴을 함께 탐지.
사용법: python -X utf8 _tab_scan.py <엑셀경로> <도메인키워드>
: python -X utf8 _tab_scan.py 충청남도/2.공주시/충청남도_공주시.xlsx gongju.go.kr
"""
import re
import sys
import warnings
from concurrent.futures import ThreadPoolExecutor, as_completed
from urllib.parse import urljoin, urlparse
import openpyxl
import requests
from bs4 import BeautifulSoup
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
# 흔한 본문 탭(서브내비) 컨테이너 클래스 패턴 (소문자 비교)
TAB_CLASS_PATS = [
'tab-ul', # 공주시
'tab_ul',
'tabmenu', 'tab-menu', 'tab_menu',
'tablist', 'tab-list', 'tab_list',
'subtab', 'sub-tab', 'sub_tab',
'tab_wrap', 'tab-wrap',
]
def find_tab_uls(soup):
"""본문 탭으로 보이는 ul들을 반환. (ul, [(text,href),...]) 리스트."""
results = []
seen = set()
for ul in soup.find_all('ul'):
cls = ' '.join(ul.get('class') or []).lower()
if not cls:
# 부모 div의 클래스도 확인 (ul 자체엔 클래스 없고 div.tab > ul 구조)
parent = ul.find_parent(['div'])
pcls = ' '.join(parent.get('class') or []).lower() if parent else ''
if not any(p in pcls for p in TAB_CLASS_PATS):
continue
elif not any(p in cls for p in TAB_CLASS_PATS):
continue
links = []
for a in ul.find_all('a'):
t = a.get_text(strip=True).replace('\xa0', '').strip()
h = (a.get('href') or '').strip()
if t:
links.append((t, h))
if len(links) >= 2:
key = tuple(t for t, _ in links)
if key not in seen:
seen.add(key)
results.append((cls, links))
return results
def scan_one(r, url):
try:
resp = requests.get(url, headers=H, timeout=15, verify=False)
resp.encoding = resp.apparent_encoding
soup = BeautifulSoup(resp.text, 'html.parser')
tabs = find_tab_uls(soup)
return r, url, tabs, None
except Exception as e:
return r, url, [], str(e)[:60]
def main():
xlsx = sys.argv[1]
domain = sys.argv[2]
wb = openpyxl.load_workbook(xlsx)
ws = wb.active
targets = []
for r in range(3, ws.max_row + 1):
u = ws.cell(r, 11).value
if isinstance(u, str) and domain in u:
targets.append((r, u))
print(f'스캔 대상: {len(targets)}개 ({domain})')
found = []
errs = []
with ThreadPoolExecutor(max_workers=8) as ex:
futs = [ex.submit(scan_one, r, u) for r, u in targets]
for fut in as_completed(futs):
r, url, tabs, err = fut.result()
if err:
errs.append((r, url, err))
elif tabs:
found.append((r, url, tabs))
found.sort()
print(f'\n=== 탭 발견: {len(found)}개 페이지 ===')
for r, url, tabs in found:
e = ws.cell(r, 5).value or ''
f = ws.cell(r, 6).value or ''
g = ws.cell(r, 7).value or ''
print(f'[행{r}] E={e} F={f} G={g}')
print(f' {url}')
for cls, links in tabs:
print(f' <ul class={cls}> 탭 {len(links)}개:')
for t, h in links:
print(f' - {t} -> {h}')
if errs:
print(f'\n=== 에러 {len(errs)}개 ===')
for r, url, err in errs[:20]:
print(f'[행{r}] {url} : {err}')
if __name__ == '__main__':
main()

File diff suppressed because it is too large Load Diff

View File

@ -0,0 +1,76 @@
# -*- coding: utf-8 -*-
"""억제된 페이지→게시판 후보(아직 페이지인 행)를 빈게시판 신호로 재분석. 읽기전용."""
import sys, os, re, json, importlib.util, time
from concurrent.futures import ThreadPoolExecutor, as_completed
import openpyxl
HERE = os.path.dirname(os.path.abspath(__file__)); TEMP = os.path.join(HERE,'..','_temp')
MODULES = ['_chungnam_phase234_all.py','_chungbuk_phase234_all.py','_jeonbuk_phase234_all.py']
def load():
sites,mods={},{}
for f in MODULES:
sp=importlib.util.spec_from_file_location(f[:-3],os.path.join(HERE,f));m=importlib.util.module_from_spec(sp);sp.loader.exec_module(m)
for k,v in m.SITES.items(): sites[k]=v;mods[k]=m
return sites,mods
EMPTY=re.compile(r'게시물이?\s*없|등록된?\s*(?:게시물|자료|글|내용)\s*가?\s*없|자료가\s*없|검색된\s*(?:게시물|결과)\s*가?\s*없|데이터가\s*없')
BOARDDOM='.board_list,.bbs_list,table.board_list,.board_view,.board_wrap,.bbs,.tbl_list,.board,.list_wrap,.gallery_list,.photo_list,table.board'
WRITE=re.compile(r'글쓰기|글등록|등록하기|write|글작성')
def mkfetch(M,cfg):
if hasattr(M,'make_session'):
s=M.make_session(weak_ssl=cfg.get('weak_ssl',False)); return lambda u:M.fetch(s,u)
return lambda u:M.fetch(u)
def analyze(M,url,bsel,fetch):
soup=None
for k in range(3):
soup=fetch(url)
if soup is not None: break
time.sleep(0.4*(k+1))
if soup is None: return {'v':'fail'}
b=M.get_body(soup,bsel)
txt=b.get_text(' ',strip=True)
search=len([i for i in b.find_all('input') if (i.get('type') or 'text').lower() in ('text','search')])>0
rows=[el for el in b.select('table tbody tr, .board_list li, ul.bbs_list li, .bbs_list li') if el.find('a',href=True)]
empty=bool(EMPTY.search(txt))
boarddom=bool(b.select(BOARDDOM))
paging=bool(b.select('.pagination,.paging,nav.paging,.page_nav,.paginate'))
write=bool(WRITE.search(txt))
if empty or (boarddom and (paging or write)):
v='empty_board'
else:
v='page'
return {'v':v,'search':search,'rows':len(rows),'empty':empty,'boarddom':boarddom,'paging':paging,'write':write,
'snip':txt[:160]}
def main():
sites,mods=load(); names=[a for a in sys.argv[1:] if not a.startswith('--')]
for n in names:
cfg=sites[n];M=mods[n];fetch=mkfetch(M,cfg)
wb=openpyxl.load_workbook(cfg['xlsx'],read_only=True);ws=wb.active
d=json.load(open(os.path.join(TEMP,f'_audit_{n}.json'),encoding='utf-8'))
cands=[]
for x in d['L_diffs']:
if x['oldL']=='페이지' and x['newL']=='게시판':
if ws.cell(x['r'],12).value=='페이지': # 아직 페이지(억제됨)
cands.append(x)
wb.close()
res=[]
with ThreadPoolExecutor(max_workers=6) as ex:
futs={ex.submit(analyze,M,x['url'],cfg['body_sel'],fetch):x for x in cands}
for f in as_completed(futs):
x=futs[f]
try:r=f.result()
except:r={'v':'fail'}
res.append((x,r))
eb=[(x,r) for x,r in res if r['v']=='empty_board']
pg=[(x,r) for x,r in res if r['v']=='page']
nf=len([1 for _,r in res if r['v']=='fail'])
print('\n==== %s: 억제후보%d → 빈게시판의심%d 페이지%d 실패%d'%(n,len(cands),len(eb),len(pg),nf))
print('--- 빈게시판 의심(empty_board) 예시 ---')
for x,r in sorted(eb,key=lambda t:t[0]['r'])[:8]:
print(f" r{x['r']} {(x['label'] or '')[:20]} | empty={r['empty']} boarddom={r['boarddom']} paging={r['paging']} | {x['url'][:62]}")
print(f" · {r['snip'][:110]}")
print('--- 순수 페이지 예시 ---')
for x,r in sorted(pg,key=lambda t:t[0]['r'])[:5]:
print(f" r{x['r']} {(x['label'] or '')[:20]} | boarddom={r['boarddom']} | {x['url'][:62]}")
print(f" · {r['snip'][:110]}")
json.dump([{**x,**r} for x,r in res],open(os.path.join(TEMP,f'_emptyboard_{n}.json'),'w',encoding='utf-8'),ensure_ascii=False,indent=1)
if __name__=='__main__': main()

View File

@ -0,0 +1,37 @@
# -*- coding: utf-8 -*-
"""plan_{기관}.json 의 auto='어문'(콘텐츠이미지 없음) 행 → N에서 '이미지' 제거(강등).
candidate 행은 보존(사용자/추가 시각판정 영역). 하늘색 표시는 하지 않음(사용자 지시).
사용: python _공공기관_napply.py <기관명>
"""
import sys, os, json
import openpyxl
OUTDIR = r'D:\01.프로젝트\DB수집\작업파일\공공기관2'
TEMP = r'D:\01.프로젝트\DB수집\_temp'
def demote(N):
parts = [p for p in (N or '').split(',') if p and p != '이미지']
return ','.join(parts) if parts else '없음'
def run(name):
plan = json.load(open(os.path.join(TEMP, f'nshot_{name}', f'plan_{name}.json'), encoding='utf-8'))
xlsx = __import__('glob').glob(os.path.join(OUTDIR, f'*.{name}', f'{name}.xlsx'))[0]
wb = openpyxl.load_workbook(xlsx)
ws = wb.active
chg = 0
for it in plan:
if it['auto'] == '어문':
r = it['row']
cur = ws.cell(r, 14).value or ''
if '이미지' in cur:
ws.cell(r, 14).value = demote(cur)
chg += 1
wb.save(xlsx)
cand = sum(1 for it in plan if it['auto'] == 'candidate')
print(f'[{name}] 이미지강등 {chg}행 | 잔여이미지후보 {cand}행(시각판정 영역)')
if __name__ == '__main__':
run(sys.argv[1])

View File

@ -0,0 +1,34 @@
# -*- coding: utf-8 -*-
"""공공기관 N 이미지 몽타주 재판정 일괄: 각 기관 렌더+크기필터(nshot) → 이미지강등 적용(napply).
국토연구원은 이미 처리 스킵. 사용: python _공공기관_nbatch.py [기관명 ...]
"""
import sys, os, json, importlib.util
def load(mod):
spec = importlib.util.spec_from_file_location(mod, rf'D:\01.프로젝트\DB수집\_스크립트\{mod}.py')
m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m)
return m
nshot = load('_공공기관2_nshot')
napply = load('_공공기관2_napply')
probe = {r['name']: r for r in json.load(open(r'D:\01.프로젝트\DB수집\_스크립트\_공공기관2_probe.json', encoding='utf-8'))}
order = sorted(probe.values(), key=lambda x: -int(x['num']))
only = sys.argv[1:]
if only:
order = [p for p in order if p['name'] in only or str(p['num']) in only]
DONE = {'국토연구원'}
for p in order:
name = p['name']
if name in DONE:
print(f'[{name}] 이미 처리 — 스킵'); continue
try:
nshot.run(name)
napply.run(name)
except Exception as e:
import traceback
print(f'[{name}] 실패: {e}'); traceback.print_exc()
print('\n=== N이미지 몽타주 재판정 일괄 완료 ===')

View File

@ -0,0 +1,123 @@
# -*- coding: utf-8 -*-
"""공공기관 N 이미지 몽타주 재판정 — 1단계: 렌더+콘텐츠이미지 크기필터 + 풀페이지 스샷 + 몽타주.
대상: N에 '이미지' 포함 & L!=사이트 & URL 있는 .
출력: _temp\nshot_{기관}\ (스샷 png + montage_*.png) + plan_{기관}.json
plan: {row, url, L, N, n_imgs(필터통과), auto='어문'(이미지없음) | 'candidate'(몽타주판정)}
사용: python _공공기관_nshot.py <기관명>
"""
import sys, os, re, json, warnings
import openpyxl
from PIL import Image, ImageDraw, ImageFont
warnings.filterwarnings('ignore')
OUTDIR = r'D:\01.프로젝트\DB수집\작업파일\공공기관2'
TEMP = r'D:\01.프로젝트\DB수집\_temp'
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
NOISE = re.compile(r'(ico[_\-/]|/icon|logo|btn|bul[_\-]|bg[_\-]|banner|sns|blank|spacer|loading|arrow|/dot|line[_\.]|top_|foot|header|common|_icon|symbol|copyright|qr_|no_img|noimage|share|facebook|insta|twitter|naver|kakao|/skin/|/template/|/resources/|/images/common)', re.I)
JS_IMGS = """() => {
const out = [];
for (const i of document.querySelectorAll('img')) {
const r = i.getBoundingClientRect();
out.push({src: i.currentSrc||i.src||'', nw: i.naturalWidth, nh: i.naturalHeight, rw: Math.round(r.width), rh: Math.round(r.height)});
}
// 배경이미지도 일부
return out;
}"""
def content_imgs(items):
cnt = 0
for it in items:
src = it.get('src', '')
if not src or NOISE.search(src):
continue
nw, nh, rw, rh = it.get('nw', 0), it.get('nh', 0), it.get('rw', 0), it.get('rh', 0)
if nw >= 170 and nh >= 110 and rw >= 100 and rh >= 80:
cnt += 1
return cnt
def run(name):
from playwright.sync_api import sync_playwright
xlsx = __import__('glob').glob(os.path.join(OUTDIR, f'*.{name}', f'{name}.xlsx'))[0]
wb = openpyxl.load_workbook(xlsx)
ws = wb.active
targets = []
for r in range(3, ws.max_row + 1):
if ws.cell(r, 2).value is None:
break
L = ws.cell(r, 12).value
N = ws.cell(r, 14).value or ''
url = ws.cell(r, 11).value
if L != '사이트' and '이미지' in N and url and isinstance(url, str) and url.startswith('http'):
targets.append((r, url, L, N))
outdir = os.path.join(TEMP, f'nshot_{name}')
os.makedirs(outdir, exist_ok=True)
print(f'[{name}] 이미지후보 {len(targets)}행 렌더…')
plan = []
shots = [] # (row, N, path)
with sync_playwright() as p:
b = p.chromium.launch()
pg = b.new_page(user_agent=UA, viewport={'width': 1280, 'height': 1600})
for idx, (r, url, L, N) in enumerate(targets):
try:
pg.goto(url, timeout=25000, wait_until='domcontentloaded')
pg.wait_for_timeout(1800)
items = pg.evaluate(JS_IMGS)
nc = content_imgs(items)
if nc == 0:
plan.append({'row': r, 'url': url, 'L': L, 'N': N, 'n_imgs': 0, 'auto': '어문'})
else:
sp = os.path.join(outdir, f'r{r}.png')
try:
pg.screenshot(path=sp, full_page=False)
except Exception:
pass
plan.append({'row': r, 'url': url, 'L': L, 'N': N, 'n_imgs': nc, 'auto': 'candidate'})
if os.path.exists(sp):
shots.append((r, N, sp))
except Exception as e:
plan.append({'row': r, 'url': url, 'L': L, 'N': N, 'n_imgs': -1, 'auto': '어문', 'err': str(e)[:40]})
if (idx + 1) % 20 == 0:
print(f' {idx+1}/{len(targets)}')
b.close()
# 몽타주 그리드 (후보만)
cols, cw, ch = 4, 300, 360
try:
font = ImageFont.truetype('malgun.ttf', 16)
except Exception:
font = ImageFont.load_default()
mont_paths = []
per = cols * 5 # 20개/장
for gi in range(0, len(shots), per):
chunk = shots[gi:gi + per]
rows_n = (len(chunk) + cols - 1) // cols
canvas = Image.new('RGB', (cols * cw, rows_n * ch), 'white')
d = ImageDraw.Draw(canvas)
for j, (r, N, sp) in enumerate(chunk):
try:
im = Image.open(sp).convert('RGB')
im.thumbnail((cw - 8, ch - 26))
except Exception:
continue
cx, cy = (j % cols) * cw, (j // cols) * ch
canvas.paste(im, (cx + 4, cy + 22))
d.rectangle([cx, cy, cx + cw - 1, cy + ch - 1], outline='gray')
d.text((cx + 4, cy + 3), f'r{r}', fill='red', font=font)
mp = os.path.join(outdir, f'montage_{name}_{gi//per+1}.png')
canvas.save(mp)
mont_paths.append(mp)
auto_amun = sum(1 for x in plan if x['auto'] == '어문')
cand = sum(1 for x in plan if x['auto'] == 'candidate')
with open(os.path.join(outdir, f'plan_{name}.json'), 'w', encoding='utf-8') as f:
json.dump(plan, f, ensure_ascii=False, indent=1)
print(f'[{name}] 자동강등(이미지無) {auto_amun} | 몽타주후보 {cand} | 몽타주 {len(mont_paths)}')
for mp in mont_paths:
print(' MONT:', mp)
print(' PLAN:', os.path.join(outdir, f'plan_{name}.json'))
if __name__ == '__main__':
run(sys.argv[1])

View File

@ -0,0 +1,393 @@
# -*- coding: utf-8 -*-
"""공공기관 Phase 1: 사이트맵/메가메뉴 → 메뉴트리 D~K (Playwright 렌더 + 범용 중첩리스트 추출).
출력: D:\\01.프로젝트\\DB수집\\공공기관\\{기관}.xlsx (평면배치)
시트명: {번호}_{기관}
사용: python _공공기관_phase1.py [기관명 ...] (인자 없으면 OVERRIDES 정의된 전체)
"""
import sys, io, os, re, shutil, json, warnings
if __name__ == '__main__':
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8')
from copy import copy
from urllib.parse import urljoin, urlparse
import openpyxl
from bs4 import BeautifulSoup
from openpyxl.styles import Alignment, Font
warnings.filterwarnings('ignore')
TEMPLATE = r'D:\01.프로젝트\DB수집\자료_취합_예시.xlsx'
OUTDIR = r'D:\01.프로젝트\DB수집\작업파일\공공기관2'
PROBE = r'D:\01.프로젝트\DB수집\_스크립트\_공공기관2_probe.json'
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
# 기관별 오버라이드: sitemap URL + 컨테이너 셀렉터(선택) + 렌더 대기(wait). probe json을 기본값으로.
OVERRIDES = {
'국립생태원': {'sitemap': 'https://www.nie.re.kr/nie/main/main.do', 'wait': 5000},
'국토안전관리원': {'use_home': True, 'sel': '.all-menu', 'wait': 4500},
'국립중앙의료원': {'use_home': True, 'sel': '#siteMap', 'wait': 4500},
'건설근로자공제회': {'use_home': True, 'wait': 4500},
'국민연금공단': {'sitemap': 'https://www.nps.or.kr/main.do', 'sel': '.allmenu-container', 'wait': 4000},
'국민건강보험공단': {'sitemap': 'https://www.nhis.or.kr/nhis/index.do', 'sel': 'nav.head-gnb', 'wait': 4000},
}
JS_PATS = [
# goMenuPage('MENUID','/realpath',...) 형: 2번째 인자가 실제 URL
re.compile(r"""go\w*(?:Menu|Page|Move)\w*\s*\(\s*['"][^'"]*['"]\s*,\s*['"]([^'"]+)['"]""", re.I),
re.compile(r"""go(?:Menu|SubMenu|Page|Link|Url|View|Move)\s*\(\s*['"]([^'"]+)['"]""", re.I),
re.compile(r"""(?:location\.href|window\.location(?:\.href)?)\s*=\s*(?:encodeURI\()?\s*['"]([^'"]+)['"]"""),
re.compile(r"""window\.open\s*\(\s*['"]([^'"]+)['"]"""),
re.compile(r"""(?:fn_?\w*|aLink|movePage|menuMove)\s*\(\s*['"]([^'"]+)['"]""", re.I),
]
def js_href(a):
s = (a.get('href', '') or '') + ' ' + (a.get('onclick', '') or '')
for p in JS_PATS:
mm = p.search(s)
if mm:
u = mm.group(1).strip()
if u.startswith(('/', 'http', './', '?', '../')):
return u
return ''
def clean(s):
return re.sub(r'\s+', ' ', (s or '')).strip().replace('\xa0', '')
def render(url, wait=2500):
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
b = p.chromium.launch()
pg = b.new_page(user_agent=UA, viewport={'width': 1440, 'height': 2400})
try:
pg.goto(url, timeout=30000, wait_until='domcontentloaded')
except Exception:
pass
pg.wait_for_timeout(wait)
html = pg.content()
final = pg.url
b.close()
return final, html
SITEMAP_HINTS = ('sitemap', 'site_map', 'sitemapwrap', 'site-map', 'allmenu', 'all_menu', 'all-menu',
'menu_all', 'menuall', 'totalmenu', 'total_menu', 'totmenu', 'full_menu', 'fullmenu',
'gnb_all', 'gnball')
def pick_container(soup, prefer=None):
if prefer:
el = soup.select_one(prefer)
if el and len(el.find_all('a')) >= 8:
return el
cands = []
for d in soup.find_all(['div', 'section', 'main', 'nav', 'ul']):
cls = ' '.join(d.get('class', [])).lower() + ' ' + (d.get('id', '') or '').lower()
clsn = cls.replace(' ', '').replace('-', '').replace('_', '')
if any(x in cls for x in ['footer', 'aside']):
continue
ac = len(d.find_all('a'))
ul = len(d.find_all('ul'))
if ac < 12:
continue
score = ac + ul * 2
# 사이트맵/전체메뉴 컨테이너 강력 우대
if any(h in clsn for h in SITEMAP_HINTS):
score += 1000
# 전역 헤더/상단 gnb는 (사이트맵 페이지에선) 감점
if 'header' in cls or ('gnb' in clsn and not any(h in clsn for h in SITEMAP_HINTS)):
score -= 200
cands.append((score, ac, d))
if not cands:
return None
cands.sort(key=lambda x: x[0], reverse=True)
return cands[0][2]
NOISE_LABELS = {'메뉴없음', '메뉴 없음', '', 'home', 'home으로', '홈으로', 'eng', 'english', '로그인',
'login', '회원가입', '검색', 'search', '바로가기', '본문바로가기', '닫기', 'close', '전체메뉴',
'사이트맵', 'sitemap'}
def node_depth(el, container):
"""container 내부에서 el 위의 ul/ol/dl 조상 개수 = 중첩 깊이."""
d = 0
p = el.parent
while p is not None and p is not container:
if getattr(p, 'name', None) in ('ul', 'ol', 'dl'):
d += 1
p = p.parent
return d
def parse_table_sitemap(table):
"""테이블형 사이트맵: tr > th(대분류, 빈칸=이어받기) + th(중분류 a) + td(소분류 a들)."""
rows = []
curD = ''
for tr in table.find_all('tr'):
ths = tr.find_all('th', recursive=False)
if ths:
d_txt = clean(ths[0].get_text())
# 첫 th에 a가 없고 span/텍스트면 대분류 라벨(빈칸이면 이어받기)
if d_txt and not (len(ths) == 1 and ths[0].find('a')):
# 첫 th가 곧 중분류 a 단독인 경우는 제외 위해 a 유무 확인
if not ths[0].find('a') or len(ths) >= 2:
if not ths[0].find('a'):
curD = d_txt
E = ''; Eh = ''
if len(ths) >= 2:
ea = ths[1].find('a')
if ea:
E = clean(ea.get_text()); Eh = js_href(ea) or (ea.get('href') or '')
elif len(ths) == 1 and ths[0].find('a'):
ea = ths[0].find('a'); E = clean(ea.get_text()); Eh = js_href(ea) or (ea.get('href') or '')
td = tr.find('td')
leaves = td.find_all('a') if td else []
if leaves:
for la in leaves:
lab = clean(la.get_text())
if not lab:
continue
rows.append({'path': [curD, E, lab], 'href': js_href(la) or (la.get('href') or '')})
elif E:
rows.append({'path': [curD, E], 'href': Eh})
return [r for r in rows if r['path'] and any(r['path'])]
def extract_rows(container):
"""앵커 중심 깊이추출: 각 a/heading의 ul조상 수로 컬럼 결정(정규화). 범용. 테이블형 자동전환."""
tbls = container.find_all('table')
if tbls and sum(len(t.find_all('th')) for t in tbls) >= 4 and len(container.find_all('ul')) <= 1:
rows = []
for t in tbls:
rows += parse_table_sitemap(t)
if rows:
return rows
nodes = [] # [depth, label, href]
for el in container.find_all(['a', 'button', 'h2', 'h3', 'h4', 'h5', 'h6', 'dt']):
name = el.name
if name == 'a':
href = (el.get('href') or '').strip()
if href.startswith('#') or href.lower().startswith('javascript:') or not href:
href = js_href(el)
label = clean(el.get_text())
elif name == 'button':
cls = ' '.join(el.get('class', [])).lower()
if not (el.get('aria-expanded') is not None or el.get('aria-haspopup')
or any(x in cls for x in ('depth', 'trigger', 'gnb', '1d', '2d', '3d', 'menu'))):
continue
href = ''
label = clean(el.get('data-dir') or el.get('title') or el.get_text())
else:
if el.find('a'): # heading 안의 a는 따로 잡힘 → 중복방지
continue
href = ''
label = clean(el.get_text())
if not label or label.lower() in NOISE_LABELS or len(label) > 60:
continue
nodes.append([node_depth(el, container), label, href])
# 연속 중복 제거
dedup = []
for n in nodes:
if dedup and dedup[-1] == n:
continue
dedup.append(n)
nodes = dedup
if not nodes:
return []
mind = min(n[0] for n in nodes)
cols = [min(n[0] - mind, 6) for n in nodes]
out = []
path = [''] * 7
for i, (depth, label, href) in enumerate(nodes):
c = cols[i]
path[c] = label
for k in range(c + 1, 7):
path[k] = ''
is_leaf = (i == len(nodes) - 1) or (cols[i + 1] <= c)
if href or is_leaf:
out.append({'path': path[:c + 1], 'href': href})
return out
def rows_to_dicts(raw):
out = []
for it in raw:
p = it['path']
href = it.get('href', '')
if href.startswith('#') or href.lower().startswith('javascript:'):
href = ''
row = {'D': '', 'E': '', 'F': '', 'G': '', 'H': '', 'I': '', 'J': '', 'href': href}
for i, lab in enumerate(p[:7]):
row['DEFGHIJ'[i]] = lab
out.append(row)
return out
def write_excel(name, num, base, raw_rows):
domain = urlparse(base).netloc
output = os.path.join(OUTDIR, f'{num}.{name}', f'{name}.xlsx')
def abs_url(href):
if not href:
return ''
href = href.strip()
if href.startswith(('javascript:', '#')):
return ''
if href.startswith(('http://', 'https://')):
return href
return urljoin(base + '/', href)
def is_external(url):
return url.startswith(('http://', 'https://')) and domain not in url
# 부모-자식 URL 중복 제거 (전 깊이) — 부모행 D~lc 동일 + lc+1 채워짐 + K동일
def colval(r, c):
return r.get(c, '') or ''
final = []
i = 0
removed = 0
cols = 'DEFGHIJ'
while i < len(raw_rows):
row = raw_rows[i]
# 마지막 채워진 컬럼 찾기
lc = -1
for ci, c in enumerate(cols):
if colval(row, c) != '':
lc = ci
dup = False
if i + 1 < len(raw_rows) and lc >= 0 and lc < 6:
nxt = raw_rows[i + 1]
same = all(colval(row, cols[k]) == colval(nxt, cols[k]) for k in range(lc + 1))
if (same and colval(nxt, cols[lc + 1]) != '' and colval(row, cols[lc + 1]) == ''
and row.get('href', '') == nxt.get('href', '')):
dup = True
if dup:
removed += 1
i += 1
continue
final.append(row)
i += 1
if not final:
print(f' [{name}] 행 0개 — 스킵')
return 0
shutil.copy(TEMPLATE, output)
wb = openpyxl.load_workbook(output)
ws = wb.active
ws.title = f'{num:02d}_{name}'
HEADER = {'B1:R1', 'S1:W1', 'Y1:AA1'}
for rng in [str(m) for m in ws.merged_cells.ranges if str(m) not in HEADER]:
ws.unmerge_cells(rng)
for row in ws.iter_rows(min_row=3, max_row=ws.max_row, min_col=1, max_col=ws.max_column):
for cell in row:
cell.value = None
START = 3
template_r = 3
cur_max = ws.max_row
for idx, item in enumerate(final, start=START):
if idx > cur_max:
for c in range(1, ws.max_column + 1):
srcc = ws.cell(template_r, c)
tgt = ws.cell(idx, c)
if srcc.has_style:
tgt.font = copy(srcc.font); tgt.fill = copy(srcc.fill)
tgt.border = copy(srcc.border); tgt.alignment = copy(srcc.alignment)
tgt.number_format = srcc.number_format; tgt.protection = copy(srcc.protection)
url = abs_url(item.get('href', ''))
ws.cell(idx, 2).value = idx - 2
ws.cell(idx, 3).value = name
for ci, c in enumerate('DEFGHIJ'):
ws.cell(idx, 4 + ci).value = item.get(c, '')
ws.cell(idx, 11).value = url
if is_external(url):
ws.cell(idx, 19).value = '외부링크'
END = START + len(final) - 1
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
def merge_runs(col_letter, col_idx, group_cols=()):
runs = []
cur_val = ws.cell(START, col_idx).value
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
run_start = START
for r in range(START + 1, END + 1):
v = ws.cell(r, col_idx).value
g = tuple(ws.cell(r, gg).value for gg in group_cols)
if v == cur_val and g == cur_grp:
continue
if cur_val not in (None, '') and r - 1 > run_start:
runs.append((run_start, r - 1))
cur_val, cur_grp, run_start = v, g, r
if cur_val not in (None, '') and END > run_start:
runs.append((run_start, END))
for s, e in runs:
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
ws.cell(s, col_idx).alignment = center
return len(runs)
n_f = merge_runs('F', 6, group_cols=(4, 5))
n_e = merge_runs('E', 5, group_cols=(4,))
n_d = merge_runs('D', 4)
for r in range(START, END + 1):
for c in (4, 5, 6, 7):
if ws.cell(r, c).value is not None:
ws.cell(r, c).alignment = center
for r in range(1, END + 1):
ws.row_dimensions[r].height = 15
link_n = 0
for r in range(START, END + 1):
cell = ws.cell(r, 11)
u = cell.value
if u and isinstance(u, str) and u.startswith(('http://', 'https://')):
cell.hyperlink = u
old = cell.font
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
bold=old.bold, italic=old.italic, color='0000FF', underline='single')
cell.alignment = left
link_n += 1
wb.save(output)
ext_n = sum(1 for r in range(START, END + 1) if ws.cell(r, 19).value == '외부링크')
print(f' [{name}] 원본{len(raw_rows)}→중복{removed}{len(final)}행 | 병합D{n_d}E{n_e}F{n_f} | 외부{ext_n} K링크{link_n}{output}')
return len(final)
def load_probe():
with open(PROBE, encoding='utf-8') as f:
return {r['name']: r for r in json.load(f)}
def main():
probe = load_probe()
targets = sys.argv[1:] if len(sys.argv) > 1 else list(probe.keys())
for name in targets:
p = probe.get(name)
if not p:
print(f'[{name}] probe 정보 없음 — 스킵'); continue
ov = OVERRIDES.get(name, {})
num = int(p['num'])
base = p['base']
if ov.get('use_home'):
sm = p['home']
else:
sm = ov.get('sitemap') or p.get('sitemap') or p['home']
wait = ov.get('wait', 3000)
print(f"\n=== {num}.{name} === render {sm}")
try:
final, html = render(sm, wait=wait)
soup = BeautifulSoup(html, 'html.parser')
cont = pick_container(soup, ov.get('sel'))
if not cont:
print(f' [{name}] 컨테이너 못찾음 (a없음)'); continue
raw = extract_rows(cont)
dicts = rows_to_dicts(raw)
write_excel(name, num, base, dicts)
except Exception as e:
import traceback
print(f' [{name}] 실패: {e}')
traceback.print_exc()
if __name__ == '__main__':
main()

View File

@ -0,0 +1,321 @@
# -*- coding: utf-8 -*-
"""공공기관 Phase 2~4: L(게시판형태)·M(수량)·N(저작물유형)·O/P/Q(공공누리).
·군용 _chungnam_phase234_all.py 로직 재사용 + 범용화(넓은 본문셀렉터·다양한 상세패턴·오디오·이미지노이즈필터).
출력: 공공기관\{기관}.xlsx L~Q 채움 (역순 실행 권장).
사용: python _공공기관_phase234.py [기관명 ...] (없으면 번호 역순 전체)
"""
import sys, os, re, json, time, warnings
from urllib.parse import urljoin, urlparse
from concurrent.futures import ThreadPoolExecutor, as_completed
import requests
from bs4 import BeautifulSoup
import openpyxl
warnings.filterwarnings('ignore')
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
H = {'User-Agent': UA}
OUTDIR = r'D:\01.프로젝트\DB수집\작업파일\공공기관2'
PROBE = r'D:\01.프로젝트\DB수집\_스크립트\_공공기관2_probe.json'
TOTAL_PAT = re.compile(r'\s*(?:게시물|건수)?\s*[:\-]?\s*(\d[\d,]*)\s*(?:건|개|page|페이지|item)', re.I)
TOTAL_PAT2 = re.compile(r'(?:전체|총|total)\s*[:\-]?\s*(\d[\d,]*)', re.I)
KOGL_IMG_PAT = re.compile(r'(?:new_)?img_open(?:type|code)(\d{1,2})\.(?:png|jpe?g|gif)', re.I)
KOGL_LINK_PAT = re.compile(r'kogl\.or\.kr/info/licenseType(\d)', re.I)
YOUTUBE_PAT = re.compile(r'(?:youtube\.com|youtu\.be|vimeo\.com)', re.I)
VIDEO_EXT = re.compile(r'\.(mp4|webm|mov|avi|m3u8)(?:\?|$)', re.I)
AUDIO_EXT = re.compile(r'\.(mp3|wav|m4a|ogg|flac)(?:\?|$|["\'&])', re.I)
IMG_NOISE = re.compile(r'(ico[_\-/]|/icon|logo|btn|bul[_\-]|bg[_\-]|banner|sns|blank|spacer|loading|arrow|/dot|line[_\.]|top_|foot|header|common|btn_|_icon|symbol|copyright|qr_|movie_ico|no_img|noimage|share|facebook|insta|youtube_ic|blog|twitter|naver|kakao)', re.I)
DETAIL_PAT = re.compile(r'(mode=V|view\.do|/view|read\.do|/read|detail\.do|/detail|bbsView|nttId=|articleNo=|boardSeq=|bbtSn=|idx=|seq=|bIdx=|board_no=|wr_id=|dataSid=|ntceSn=|brdId=|bcIdx=|page_idx=|=view)', re.I)
BODY_SEL = [
'#content', '#contents', '#contentsArea', '.contentsArea', '.content', '.contents',
'#sub_content', '.sub_content', '.sub_contents', '#subContent', '.subContent',
'#container .content', '.board_view', '.bbs_view', '.view_con', '.view_cont',
'.board', '.bbs', '#bbs', '.sub_cont', '#cont', '.cont_area', '#content_area',
'main', '#main', 'article', '.board_list', '.bbs_list',
]
def make_session(weak=False):
s = requests.Session()
s.headers.update(H)
return s
def fetch(session, url, timeout=12):
try:
r = session.get(url, timeout=timeout, verify=False, allow_redirects=True)
meta = re.search(rb'charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
r.encoding = meta.group(1).decode(errors='ignore') if meta else r.apparent_encoding
if r.status_code == 200:
return BeautifulSoup(r.text, 'html.parser'), r.url
except Exception:
pass
return None, None
def get_body(soup, selectors):
for sel in selectors:
try:
el = soup.select_one(sel)
if el and len(el.get_text(strip=True)) > 20:
return el
except Exception:
pass
return soup
def detect_form(body):
has_paging = bool(body.select('.pagination, .paging, nav.paging, .page_nav, .paginate, .pgwrap, .board_paging, .num_wrap'))
text_inputs = [i for i in body.find_all('input')
if (i.get('type') or 'text').lower() in ('text', 'search')]
has_search = len(text_inputs) >= 1
has_listtable = bool(body.select('table.board_list, table.bbs_list, ul.board_list, .board_list tbody tr, .bbs_list li'))
txt = body.get_text(' ', strip=True)
m = TOTAL_PAT.search(txt) or TOTAL_PAT2.search(txt)
total = None
if m:
digits = m.group(1).replace(',', '')
if digits.isdigit():
total = int(digits)
is_board = has_paging or has_listtable or (total is not None and (has_search or has_paging or has_listtable))
if is_board:
return '게시판', total if total is not None else 0
if has_search and (has_paging or has_listtable):
return '게시판', total if total is not None else 0
return '페이지', 1
def extract_detail_urls(body, base_url, limit=6):
urls, seen = [], set()
for a in body.find_all('a', href=True):
h = a['href']
if not h or h.startswith('#') or h.lower().startswith('javascript:'):
continue
if DETAIL_PAT.search(h):
full = urljoin(base_url, h)
if full not in seen:
seen.add(full); urls.append(full)
if len(urls) >= limit:
break
return urls
def detect_media(body):
has_text = len(body.get_text(strip=True)) > 30
has_image = False
for img in body.find_all('img'):
src = img.get('src') or img.get('data-src') or ''
if not src or KOGL_IMG_PAT.search(src) or IMG_NOISE.search(src):
continue
w = img.get('width', '')
try:
if w and int(re.sub(r'\D', '', w) or 0) and int(re.sub(r'\D', '', w)) < 60:
continue
except Exception:
pass
has_image = True
break
has_video = bool(body.find_all('video'))
if not has_video:
for ifr in body.find_all('iframe'):
if YOUTUBE_PAT.search(ifr.get('src', '')):
has_video = True; break
if not has_video:
for a in body.find_all('a', href=True):
if YOUTUBE_PAT.search(a['href']):
has_video = True; break
if not has_video and VIDEO_EXT.search(str(body)):
has_video = True
has_audio = bool(body.find_all('audio')) or bool(AUDIO_EXT.search(str(body)))
return has_image, has_video, has_audio, has_text
def n_string(has_text, has_image, has_video, has_audio):
parts = []
if has_text:
parts.append('어문')
if has_image:
parts.append('이미지')
if has_video:
parts.append('영상')
if has_audio:
parts.append('오디오')
return ','.join(parts) if parts else '없음'
def img_has_valid_anchor(img):
p = img.parent
while p is not None:
if p.name == 'a':
href = p.get('href', '')
return bool(href and not href.startswith('#') and not href.lower().startswith('javascript:'))
p = p.parent
return False
def detect_kogl(body):
types, q_y = set(), False
for a in body.find_all('a', href=True):
m = KOGL_LINK_PAT.search(a['href'])
if m:
types.add(int(m.group(1))); q_y = True
for img in body.find_all('img'):
m = KOGL_IMG_PAT.search(img.get('src', ''))
if m and 1 <= int(m.group(1)) <= 4:
types.add(int(m.group(1)))
if img_has_valid_anchor(img):
q_y = True
for el in body.find_all(style=True):
m = KOGL_IMG_PAT.search(el.get('style', ''))
if m and 1 <= int(m.group(1)) <= 4:
types.add(int(m.group(1)))
# 1·2·3·4 전부 = 범례페이지 = 미부착
if types == {1, 2, 3, 4}:
return set(), None
if not types:
return set(), None
return types, ('Y' if q_y else 'N')
def process_row(session, url, body_selectors):
out = {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
soup, final = fetch(session, url)
if soup is None:
out['note'] = '접근 실패'
return out
body = get_body(soup, body_selectors)
form, count = detect_form(body)
out['L'] = form
out['M'] = count if form == '게시판' else 1
has_img, has_vid, has_aud, has_txt = detect_media(body)
types_main, q_main = detect_kogl(body)
P = '게시판' if types_main else ''
types_all = set(types_main)
q_flags = [q_main] if q_main else []
if form == '게시판':
for du in extract_detail_urls(body, final or url, limit=6):
d_soup, _ = fetch(session, du, timeout=10)
if not d_soup:
continue
d_body = get_body(d_soup, body_selectors)
di, dv, da, dt = detect_media(d_body)
has_img |= di; has_vid |= dv; has_aud |= da; has_txt |= dt
dt_types, dt_q = detect_kogl(d_body)
if dt_types and not types_main and not P:
P = '게시물'
types_all |= dt_types
if dt_q:
q_flags.append(dt_q)
out['N'] = n_string(has_txt, has_img, has_vid, has_aud)
if not types_all:
out['O'] = '미부착'
else:
out['O'] = ','.join(f'{n}유형' for n in sorted(types_all))
out['P'] = P if P else '게시판'
out['Q'] = 'Y' if 'Y' in q_flags else 'N'
return out
def run_site(name, num, body_selectors, workers=8):
xlsx = os.path.join(OUTDIR, f'{num}.{name}', f'{name}.xlsx')
if not os.path.exists(xlsx):
print(f'[{name}] 파일 없음 — 스킵'); return None
wb = openpyxl.load_workbook(xlsx)
ws = wb.active
START = 3
END = START - 1
for r in range(START, ws.max_row + 1):
if ws.cell(r, 2).value is None:
break
END = r
if END < START:
print(f'[{name}] 데이터행 없음'); return None
# 이미 처리됨(이어하기): L열 채워진 비율 ≥90%면 스킵
filled = sum(1 for r in range(START, END + 1) if ws.cell(r, 12).value)
if '--force' not in sys.argv and filled >= (END - START + 1) * 0.9:
print(f'[{num}.{name}] 이미 처리됨({filled}/{END-START+1}) — 스킵')
return {'name': name, 'forms': {'(skip)': filled}, 'attach': 0, 'rows': END - START + 1}
tasks = []
for r in range(START, END + 1):
url = ws.cell(r, 11).value
is_ext = (ws.cell(r, 19).value == '외부링크')
tasks.append((r, url, is_ext))
n_ext = sum(1 for t in tasks if t[2])
print(f'\n[{num}.{name}] {len(tasks)}행 (외부 {n_ext}) 처리…')
t0 = time.time()
session = make_session()
results = {}
def worker(task):
row, url, is_ext = task
if is_ext:
return row, {'L': '사이트', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
if not url or not isinstance(url, str) or not url.startswith('http'):
return row, {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': 'URL 없음'}
return row, process_row(session, url, body_selectors)
with ThreadPoolExecutor(max_workers=workers) as ex:
futs = [ex.submit(worker, t) for t in tasks]
done = 0
for fut in as_completed(futs):
row, res = fut.result()
results[row] = res
done += 1
if done % 50 == 0 or done == len(tasks):
print(f' {done}/{len(tasks)} ({time.time()-t0:.0f}s)')
for r in range(START, END + 1):
res = results.get(r)
if not res:
continue
if res.get('L'):
ws.cell(r, 12).value = res['L']
if res.get('M') != '':
ws.cell(r, 13).value = res['M']
if res.get('N'):
ws.cell(r, 14).value = res['N']
if res.get('O'):
ws.cell(r, 15).value = res['O']
if res.get('P'):
ws.cell(r, 16).value = res['P']
if res.get('Q'):
ws.cell(r, 17).value = res['Q']
if res.get('note') and not ws.cell(r, 19).value:
ws.cell(r, 19).value = res['note']
wb.save(xlsx)
forms, attach = {}, 0
for res in results.values():
forms[res.get('L', '')] = forms.get(res.get('L', ''), 0) + 1
if res.get('O') and res.get('O') != '미부착':
attach += 1
fs = ' '.join(f'{k}{v}' for k, v in forms.items() if k)
print(f' ✓ [{name}] {fs} | 공공누리부착 {attach} ({time.time()-t0:.0f}s)')
return {'name': name, 'forms': forms, 'attach': attach, 'rows': len(tasks)}
def main():
probe = {r['name']: r for r in json.load(open(PROBE, encoding='utf-8'))}
order = sorted(probe.values(), key=lambda x: -int(x['num'])) # 번호 역순
only = sys.argv[1:]
if only:
order = [p for p in order if p['name'] in only or str(p['num']) in only]
summ = []
for p in order:
try:
r = run_site(p['name'], int(p['num']), BODY_SEL)
if r:
summ.append(r)
except Exception as e:
import traceback
print(f"[{p['name']}] 실패: {e}")
traceback.print_exc()
print('\n=== 요약 ===')
for s in summ:
fs = ' '.join(f'{k}{v}' for k, v in s['forms'].items() if k)
print(f" {s['name']}: {fs} | 부착{s['attach']} /{s['rows']}")
if __name__ == '__main__':
main()

Some files were not shown because too many files have changed in this diff Show More